Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff GenAI Inference Engineer: Optimize LLM Serving Latency

$190.9k - $232.8k

Jobleads-US

A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software engineering background and a proven ability to collaborate with researchers and drive architectural decisions. Competitive compensation is offered, with a salary range of $190,900 to $232,800 USD. #J-18808-Ljbffr Jobleads-US

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff GenAI Inference Engineer: Optimize LLM Serving Latency in San Francisco, CA vacancy
  •  ...Our team analyzes inference stack performance across...  ...understanding into performance optimizations and models that...  ...You will build cost-to-serve estimates from...  ...functional teams reason about latency, capacity, utilization...  ...collaborating with engineering and research teams to... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build...  ...infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability. This role... 
    Suggested

    Inception

    San Francisco, CA
    3 days ago
  •  ...the Team OpenAI’s Inference team powers the...  ...fast-moving team of engineers focused on delivering...  ...needed to serve models that handle...  ...You'll build and optimize the systems that let...  ...high-throughput, low-latency delivery of image...  ...like vLLM, TensorRT-LLM, or custom model parallel... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...fast-growing team of engineers in San Francisco...  ...models Low-latency OCR models for targeted...  ..., high-throughput inference for OCR and...  ...batching and caching Optimize kernels,...  ...Evaluate vLLM, TensorRT LLM, and Triton tradeoffs...  ...profiling and model serving Nice to... 
    Suggested
    Visa sponsorship
    Relocation package

    PULSE

    San Francisco, CA
    23 hours ago
  • $190k - $265k

     ...business. Founded by engineers — and customer-...  ...model training, model serving, and Vector Search...  ...Foundation Model Inference team is the...  ...serve, scale, and optimize frontier models with...  ...you will have:Build LLM infrastructure powering...  ...reliability, latency, and efficiency of... 
    Suggested
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $215k - $260k

     ...means owning the inference stack end to end:...  ...bringing modern optimization techniques into real...  ...deep into the serving code when the defaults...  ...patterns, latency targets, and cost...  ...directly with customer engineering teams to tailor...  ...with modern LLM serving frameworks... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  •  ...a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams...  ...systems responsible for serving models reliably, efficiently...  ...launches, inference optimizations, cloud provider integrations...  ...in correctness, latency, or reliability meaningfully... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...Develop world-class model serving stack for state-of-the...  ...end-to-end and tail latency (p95/p99), increase throughput...  ..., and server-level optimizations. Build large-scale,... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  • $186k - $269k

    Build GenAI PoCs (LLMs, RAG, agentic frameworks) to...  ...software solutions.Serve as engineering Lead delivering complex...  ...safety, compliance, and latency.Minimum qualifications...  ...experience building LLM evaluation (evals) pipelines...  ..., observability, and optimizing LLM-native metrics at... 

    Google

    San Bruno, CA
    3 days ago
  • $170k - $216k

     ...of customers Software Engineers, Product, Data Science...  ...Build and evolve ML inference infrastructure for simulations...  ...for the reliability, latency, and user experience...  ...model deployment and serving.   You have: ~...  ...frameworks, TPUs and optimizing models for serving.... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    23 hours ago
  • $190.9k - $232.8k

     ...1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture...  ..., development, and optimization of the inference engine...  ...high throughput, low latency, and robust scaling. Your...  ...and collaborate on model-serving stack optimized for large... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $220k

    We build and run the inference engine behind every Perplexity...  ...architectures at scale with tight latency and cost budgets. Our...  .... Rust-native serving runtime. Develop our...  ...You understand modern LLM architectures and are...  ...and inference optimization techniques (e.g. quantization... 

    Perplexity

    San Francisco, CA
    3 days ago
  • LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for...  ...while integrating with model-serving infrastructure. The role emphasizes...  ...frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup... 

    Leoforce

    San Francisco, CA
    3 days ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...that powers large-scale LLM inference across our platform...  ...to make new inference optimizations broadly available to...  ...where reliability, latency, and scale are first-... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology...  .... About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure...  .... Integrate internal model-serving infrastructure (e.g., vLLM, Triton)... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building...  ...distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates will have... 

    Jobleads-US

    San Francisco, CA
    1 day ago
  •  ...ll doAs a Software Engineer on the AI Platform...  ...processing, model serving, and data...  ...and automated model optimization.This position is an...  ...large-scale model inference and data processingDesign...  ...strict uptime and latency...  ...level services for LLM orchestration, RAG... 
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    San Francisco, CA
    4 days ago
  • $137.1k - $201.6k

     ...DoorDash’s GenAI Platform team sits...  ...real-time GPU serving, high-throughput batch inference, and fine-tuning...  ...large cost and latency wins (for example...  ...including the LLM Gateway, Agent...  ...and inference engines, fine-tuning and...  ...RLHF/RLVR), agent optimization, and other post... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    14 hours ago
  •  ...model runtime within the inference engine that executes complex, frontier...  ...layers of the cluster serving software stack,...  ...efficient execution while optimizing for throughput, latency, utilization, and reliability...  ...Design and implement the LLM inference runtime for frontier... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  • $188k - $275k

     ...the team: The Inference team is responsible...  ...performance model serving capabilities that...  ...improve throughput, latency, reliability, and...  ...for an Applied AI Engineer to help us understand...  ...driving targeted optimizations for both platform-...  ...Familiarity with LLM inference systems... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    5 days ago
  •  ...powers mission-critical inference for the world's most...  ...build the platform engineers turn to to ship AI products...  ...systems, model serving, and developer experience...  ...Profile and optimize TensorRT-LLM kernels, analyze CUDA...  ...record of owning low‑latency, reliable backend services... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  • $197.3k - $225.1k

     ...Overview Lead AI Engineer (AI Foundations, LLM Core and Agentic AI)...  ...to reimagine how we serve our customers and businesses...  ...language model inference, similarity search,...  ...-of-the-art LLM optimization techniques to improve...  ...scalability, cost, latency, throughput — of large... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    more than 2 months ago
  •  ...TeamDoorDash’s GenAI Platform team sits...  ...real-time GPU serving, high-throughput batch inference, and fine-...  ...large cost and latency wins (for example...  ...including the LLM Gateway, Agent...  ...and inference engines, fine-tuning and...  ...RLHF/RLVR), agent optimization, and other post... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    3 days ago
  • $100k - $300k

     ...team, founded by the engineers who scaled...  ...Implement and iterate on LLM, VLM, and VLA...  ...input/tokenization, inference runners, and output...  ...production-grade serving and inference tooling...  ..., low-latency operation on bench...  ...Performance & optimization: Practical work improving... 
    Full time

    Humble Robotics

    San Francisco, CA
    14 hours ago
  • $286.2k - $326.7k

    Sr. Distinguished AI Engineer (Remote Eligible)...  ...capabilities to reimagine how we serve our customers and...  ...large language model inference, similarity search, guardrails...  ...state-of-the-art LLM optimization techniques to improve...  ...- scalability, cost, latency, throughput - of large... 
    Full time
    Part time
    Local area
    Remote work

    Capital One Financial Corporation

    San Francisco, CA
    23 hours ago
  • $150k - $170k

     ...oriented Applied AI Engineer to help develop,...  ...model integration, inference services,...  ...acquisition, model serving, monitoring, lifecycle...  ..., and performance optimization. Learn quickly...  ...learning models, LLM applications, computer...  ...edge deployment, low-latency inference,... 
    Work experience placement
    Casual work
    Work at office
    Relocation package
    2 days per week

    Chaos Industries

    San Francisco, CA
    2 days ago
  •  ...build, train, and serve AI models...  ...them to a working inference call fastest, and...  ...are a working engineer: you ship code,...  ...the Senior or Staff level and will...  ...ve worked with LLM APIs in production...  ...what inference latency, throughput,...  ...Depth in inference optimization, serving... 
    Permanent employment
    Flexible hours

    Fireworks AI

    San Francisco, CA
    23 hours ago
  •  ..., high-efficiency serving platform. Backed by...  ...support from AMD engineers the team is scaling...  ...AI models are optimized and deployed at scale...  ...Design and build a low-latency, chat-like...  ...engineering, open source inference engine like vLLM, Sglang, or TRT-LLM Streaming... 
    Work at office
    Flexible hours

    Sciforium

    San Francisco, CA
    4 days ago
  • $260k - $340k

     ...Principal Systems Software Engineer, you will serve as the visionary...  ...via zero-latency InfiniBand/RDMA fabrics...  ...IaaS: Design highly optimized, thin virtualization...  ...: Work alongside Staff and Senior engineers...  ...Large Language Model (LLM) training and inference at scale.Peer-... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $192k - $260k

     ...Foundation Model Serving is the API...  ...frontier AI model inference for open source...  ...re looking for engineers who have owned...  ...deep building LLM APIs and runtimes...  ...at scale.As a Staff Engineer, you’ll...  ...throughput, low-latency inference on...  ...trade-offs to optimize performance, throughput... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff GenAI Inference Engineer: Optimize LLM Serving Latency. Be the first to apply!