Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff GenAI Inference Engineer: Optimize LLM Serving Latency

$190.9k - $232.8k

Jobleads-US

A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software engineering background and a proven ability to collaborate with researchers and drive architectural decisions. Competitive compensation is offered, with a salary range of $190,900 to $232,800 USD. #J-18808-Ljbffr Jobleads-US

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff GenAI Inference Engineer: Optimize LLM Serving Latency in San Francisco, CA vacancy
  •  ...Our team analyzes inference stack performance across...  ...understanding into performance optimizations and models that...  ...You will build cost-to-serve estimates from...  ...functional teams reason about latency, capacity, utilization...  ...collaborating with engineering and research teams to... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    18 hours ago
  •  ...the Team OpenAI’s Inference team powers the...  ...fast-moving team of engineers focused on delivering...  ...needed to serve models that handle...  ...You'll build and optimize the systems that let...  ...high-throughput, low-latency delivery of image...  ...like vLLM, TensorRT-LLM, or custom model parallel... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    18 hours ago
  •  ...fast-growing team of engineers in San Francisco...  ...models Low-latency OCR models for targeted...  ..., high-throughput inference for OCR and...  ...batching and caching Optimize kernels,...  ...Evaluate vLLM, TensorRT LLM, and Triton tradeoffs...  ...profiling and model serving Nice to have... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Pulse

    San Francisco, CA
    18 hours ago
  •  ...performance model inference and accelerating...  ...drive the design, optimization, and scaling of...  ...role, you’ll lead engineering efforts to ensure...  ...-throughput, low-latency environments. You...  ...infrastructure for serving frontier AI...  ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    18 hours ago
  • $142.2k - $204.6k

     ...About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers...  ...our large language model (LLM) serving systems are fast, scalable, and...  ...into the engineOptimize for latency, throughput, memory... 
    Suggested
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $190k - $265k

     ...business. Founded by engineers — and customer-...  ...model training, model serving, and Vector Search...  ...Foundation Model Inference team is the...  ...serve, scale, and optimize frontier models with...  ...you will have:Build LLM infrastructure powering...  ...reliability, latency, and efficiency of... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • Crusoe in San Francisco is seeking a Staff Technical Program Manager to lead the Managed Inference platform team. You will ensure end-to-end program delivery for LLM workloads, driving innovation in...  ...a strong focus on AI and model serving. Join us at Crusoe and make a... 

    Crusoe

    San Francisco, CA
    2 days ago
  • $250k - $300k

     ...means owning the inference stack end to end:...  ...bringing modern optimization techniques into real...  ...deep into the serving code when the defaults...  ...patterns, latency targets, and cost...  ...directly with customer engineering teams to tailor...  ...with modern LLM serving frameworks... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $188k - $275k

     ...the team: The Inference team is responsible...  ...performance model serving capabilities that...  ...improve throughput, latency, reliability, and...  ...for an Applied AI Engineer to help us understand...  ...driving targeted optimizations for both platform-...  ...Familiarity with LLM inference systems... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    13 days ago
  • $300k

     ...committed researchers, engineers, policy experts, and...  ...the role Our Inference team is responsible for...  ...systems that serve Claude to millions of...  ...management systems LLM inference optimization, batching, and caching...  ...Currently, we expect all staff to be in one of our offices... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    18 hours ago
  •  ...performance infrastructure to serve OpenAI’s frontier...  ...scale. As part of the inference team, you’ll be...  ...tuning memory layouts, and optimizing model execution at the...  ...for a kernel-focused engineer to lead efforts in writing...  ...for throughput and latency. Contribute to and... 
    Full time

    OpenAI

    San Francisco, CA
    18 hours ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...Develop world-class model serving stack for state-of-the...  ...end-to-end and tail latency (p95/p99), increase throughput...  ..., and server-level optimizations. Build large-scale,... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    18 hours ago
  •  ...a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams...  ...systems responsible for serving models reliably, efficiently...  ...launches, inference optimizations, cloud provider integrations...  ...in correctness, latency, or reliability meaningfully... 
    Full time

    OpenAI

    San Francisco, CA
    18 hours ago
  • $186k - $269k

    Build GenAI PoCs (LLMs, RAG, agentic frameworks) to...  ...software solutions.Serve as engineering Lead delivering complex...  ...safety, compliance, and latency.Minimum qualifications...  ...experience building LLM evaluation (evals) pipelines...  ..., observability, and optimizing LLM-native metrics at... 

    Google

    San Bruno, CA
    2 days ago
  • $170k - $216k

     ...of customers Software Engineers, Product, Data Science...  ...Build and evolve ML inference infrastructure for simulations...  ...for the reliability, latency, and user experience...  ...model deployment and serving.   You have: ~...  ...frameworks, TPUs and optimizing models for serving.... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    18 hours ago
  •  ...scale intelligence to serve humanity. We’re...  ...of researchers, engineers, designers, and more...  ...of Technical Staff to join the Model...  ...many teams to deploy optimized NLP models to production in low latency, high throughput,...  ...and throughput of inference. ~ Strong understanding... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    18 hours ago
  • AI Backend Engineer, Artificial Intelligence (AI) Required...  ...real users, where latency, correctness,...  ...backend systems that serve AI‑powered insurance...  ...business logic. Build inference pipelines for LLM‑based and AI‑assisted...  ...automation systems. Optimize latency, throughput,... 
    Remote job
    Work from home

    Gina’s Tech Jobs - IT Recruiting Agency

    San Francisco, CA
    1 day ago
  • $220k

    We build and run the inference engine behind every Perplexity...  ...architectures at scale with tight latency and cost budgets. Our...  .... Rust-native serving runtime. Develop our...  ...You understand modern LLM architectures and are...  ...and inference optimization techniques (e.g. quantization... 

    Perplexity

    San Francisco, CA
    3 days ago
  • $300k

     ...committed researchers, engineers, policy experts, and business...  ...the Role The Cloud Inference team scales and optimizes Claude to serve the massive audiences of...  ...Strong familiarity with LLM inference optimization,...  ..., we expect all staff to be in one of our offices... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    18 hours ago
  •  ...About the Team OpenAI’s Inference team ensures that our most advanced...  ..., and at scale. We build and optimize the systems that power our...  ...communication libraries, and serving infrastructure - to...  ...About the Role We’re hiring engineers to scale and optimize OpenAI’... 
    Full time

    OpenAI

    San Francisco, CA
    18 hours ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...that powers large-scale LLM inference across our platform...  ...to make new inference optimizations broadly available to...  ...where reliability, latency, and scale are first-... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    18 hours ago
  • $320k

     ...committed researchers, engineers, policy experts, and...  ...Our mandate is to make inference deployment boring and...  ...unattended. Anthropic serves Claude to millions of...  ...resource-constrained optimization problem at its core: validation...  ..., we expect all staff to be in one of our... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    18 hours ago
  •  ...ll doAs a Software Engineer on the AI Platform...  ...processing, model serving, and data...  ...and automated model optimization.This position is an...  ...large-scale model inference and data processingDesign...  ...strict uptime and latency...  ...level services for LLM orchestration, RAG... 
    Contract work
    Work at office
    Local area
    Remote work
    2 days per week

    DocuSign

    San Francisco, CA
    4 days ago
  •  ...powers mission-critical inference for the world's most...  ...build the platform engineers turn to to ship AI products...  ...systems, model serving, and developer experience...  ...Profile and optimize TensorRT-LLM kernels, analyze CUDA...  ...record of owning low‑latency, reliable backend services... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    18 hours ago
  •  ...TeamDoorDash’s GenAI Platform team sits...  ...real-time GPU serving, high-throughput batch inference, and fine-...  ...large cost and latency wins (for example...  ...including the LLM Gateway, Agent...  ...and inference engines, fine-tuning and...  ...RLHF/RLVR), agent optimization, and other post... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    3 days ago
  • $286.2k - $326.7k

    Sr. Distinguished AI Engineer (Remote Eligible)...  ...capabilities to reimagine how we serve our customers and...  ...large language model inference, similarity search, guardrails...  ...state-of-the-art LLM optimization techniques to improve...  ...- scalability, cost, latency, throughput - of large... 
    Full time
    Part time
    Local area
    Remote work

    Capital One Financial Corporation

    San Francisco, CA
    18 hours ago
  • $314.8k - $359.3k

     ...Senior Distinguished AI Engineer At Capital...  ...capabilities to reimagine how we serve our customers and...  ...large language model inference, similarity search,...  ...state-of-the-art LLM optimization techniques to improve...  ...- scalability, cost, latency, throughput - of large... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    18 hours ago
  •  ..., high-efficiency serving platform. Backed by...  ...support from AMD engineers the team is scaling...  ...AI models are optimized and deployed at scale...  ...Design and build a low-latency, chat-like...  ...engineering, open source inference engine like vLLM, Sglang, or TRT-LLM Streaming... 
    Work at office
    Flexible hours

    Sciforium

    San Francisco, CA
    4 days ago
  • $192k - $260k

     ...Foundation Model Serving is the API...  ...frontier AI model inference for open source...  ...re looking for engineers who have owned...  ...deep building LLM APIs and runtimes...  ...at scale.As a Staff Engineer, you’ll...  ...throughput, low-latency inference on...  ...trade-offs to optimize performance, throughput... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $260k - $340k

     ...Principal Systems Software Engineer, you will serve as the visionary...  ...via zero-latency InfiniBand/RDMA fabrics...  ...IaaS: Design highly optimized, thin virtualization...  ...: Work alongside Staff and Senior engineers...  ...Large Language Model (LLM) training and inference at scale.Peer-... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff GenAI Inference Engineer: Optimize LLM Serving Latency. Be the first to apply!