Staff GenAI Inference Engineer: Optimize LLM Serving Latency
$190.9k - $232.8kJobleads-US
A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software engineering background and a proven ability to collaborate with researchers and drive architectural decisions. Competitive compensation is offered, with a salary range of $190,900 to $232,800 USD. #J-18808-Ljbffr Jobleads-US
- ...Our team analyzes inference stack performance across... ...understanding into performance optimizations and models that... ...You will build cost-to-serve estimates from... ...functional teams reason about latency, capacity, utilization... ...collaborating with engineering and research teams to...SuggestedFull time
- ...the Team OpenAI’s Inference team powers the... ...fast-moving team of engineers focused on delivering... ...needed to serve models that handle... ...You'll build and optimize the systems that let... ...high-throughput, low-latency delivery of image... ...like vLLM, TensorRT-LLM, or custom model parallel...SuggestedFull time
- ...fast-growing team of engineers in San Francisco... ...models Low-latency OCR models for targeted... ..., high-throughput inference for OCR and... ...batching and caching Optimize kernels,... ...Evaluate vLLM, TensorRT LLM, and Triton tradeoffs... ...profiling and model serving Nice to have...SuggestedFull timeWork at officeVisa sponsorshipRelocation package
- ...performance model inference and accelerating... ...drive the design, optimization, and scaling of... ...role, you’ll lead engineering efforts to ensure... ...-throughput, low-latency environments. You... ...infrastructure for serving frontier AI... ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron...SuggestedFull time
$142.2k - $204.6k
...About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers... ...our large language model (LLM) serving systems are fast, scalable, and... ...into the engineOptimize for latency, throughput, memory...SuggestedLocal areaWorldwide$190k - $265k
...business. Founded by engineers — and customer-... ...model training, model serving, and Vector Search... ...Foundation Model Inference team is the... ...serve, scale, and optimize frontier models with... ...you will have:Build LLM infrastructure powering... ...reliability, latency, and efficiency of...Local areaWorldwide- Crusoe in San Francisco is seeking a Staff Technical Program Manager to lead the Managed Inference platform team. You will ensure end-to-end program delivery for LLM workloads, driving innovation in... ...a strong focus on AI and model serving. Join us at Crusoe and make a...
$250k - $300k
...means owning the inference stack end to end:... ...bringing modern optimization techniques into real... ...deep into the serving code when the defaults... ...patterns, latency targets, and cost... ...directly with customer engineering teams to tailor... ...with modern LLM serving frameworks...Temporary work$188k - $275k
...the team: The Inference team is responsible... ...performance model serving capabilities that... ...improve throughput, latency, reliability, and... ...for an Applied AI Engineer to help us understand... ...driving targeted optimizations for both platform-... ...Familiarity with LLM inference systems...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$300k
...committed researchers, engineers, policy experts, and... ...the role Our Inference team is responsible for... ...systems that serve Claude to millions of... ...management systems LLM inference optimization, batching, and caching... ...Currently, we expect all staff to be in one of our offices...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...performance infrastructure to serve OpenAI’s frontier... ...scale. As part of the inference team, you’ll be... ...tuning memory layouts, and optimizing model execution at the... ...for a kernel-focused engineer to lead efforts in writing... ...for throughput and latency. Contribute to and...Full time
- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI... ...Develop world-class model serving stack for state-of-the... ...end-to-end and tail latency (p95/p99), increase throughput... ..., and server-level optimizations. Build large-scale,...Full timeFlexible hours
- ...a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams... ...systems responsible for serving models reliably, efficiently... ...launches, inference optimizations, cloud provider integrations... ...in correctness, latency, or reliability meaningfully...Full time
$186k - $269k
Build GenAI PoCs (LLMs, RAG, agentic frameworks) to... ...software solutions.Serve as engineering Lead delivering complex... ...safety, compliance, and latency.Minimum qualifications... ...experience building LLM evaluation (evals) pipelines... ..., observability, and optimizing LLM-native metrics at...$170k - $216k
...of customers Software Engineers, Product, Data Science... ...Build and evolve ML inference infrastructure for simulations... ...for the reliability, latency, and user experience... ...model deployment and serving. You have: ~... ...frameworks, TPUs and optimizing models for serving....Full timeRemote work- ...scale intelligence to serve humanity. We’re... ...of researchers, engineers, designers, and more... ...of Technical Staff to join the Model... ...many teams to deploy optimized NLP models to production in low latency, high throughput,... ...and throughput of inference. ~ Strong understanding...Full timeWork experience placementWork at officeRemote workFlexible hours
- AI Backend Engineer, Artificial Intelligence (AI) Required... ...real users, where latency, correctness,... ...backend systems that serve AI‑powered insurance... ...business logic. Build inference pipelines for LLM‑based and AI‑assisted... ...automation systems. Optimize latency, throughput,...Remote jobWork from home
$220k
We build and run the inference engine behind every Perplexity... ...architectures at scale with tight latency and cost budgets. Our... .... Rust-native serving runtime. Develop our... ...You understand modern LLM architectures and are... ...and inference optimization techniques (e.g. quantization...$300k
...committed researchers, engineers, policy experts, and business... ...the Role The Cloud Inference team scales and optimizes Claude to serve the massive audiences of... ...Strong familiarity with LLM inference optimization,... ..., we expect all staff to be in one of our offices...Full timeWork at officeVisa sponsorshipFlexible hours- ...About the Team OpenAI’s Inference team ensures that our most advanced... ..., and at scale. We build and optimize the systems that power our... ...communication libraries, and serving infrastructure - to... ...About the Role We’re hiring engineers to scale and optimize OpenAI’...Full time
- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI... ...that powers large-scale LLM inference across our platform... ...to make new inference optimizations broadly available to... ...where reliability, latency, and scale are first-...Full timeFlexible hours
$320k
...committed researchers, engineers, policy experts, and... ...Our mandate is to make inference deployment boring and... ...unattended. Anthropic serves Claude to millions of... ...resource-constrained optimization problem at its core: validation... ..., we expect all staff to be in one of our...Full timeWork at officeVisa sponsorshipFlexible hoursShift work- ...ll doAs a Software Engineer on the AI Platform... ...processing, model serving, and data... ...and automated model optimization.This position is an... ...large-scale model inference and data processingDesign... ...strict uptime and latency... ...level services for LLM orchestration, RAG...Contract workWork at officeLocal areaRemote work2 days per week
- ...powers mission-critical inference for the world's most... ...build the platform engineers turn to to ship AI products... ...systems, model serving, and developer experience... ...Profile and optimize TensorRT-LLM kernels, analyze CUDA... ...record of owning low‑latency, reliable backend services...Full timeFlexible hours
- ...TeamDoorDash’s GenAI Platform team sits... ...real-time GPU serving, high-throughput batch inference, and fine-... ...large cost and latency wins (for example... ...including the LLM Gateway, Agent... ...and inference engines, fine-tuning and... ...RLHF/RLVR), agent optimization, and other post...Hourly payWork at officeLocal areaRemote workFlexible hours
$286.2k - $326.7k
Sr. Distinguished AI Engineer (Remote Eligible)... ...capabilities to reimagine how we serve our customers and... ...large language model inference, similarity search, guardrails... ...state-of-the-art LLM optimization techniques to improve... ...- scalability, cost, latency, throughput - of large...Full timePart timeLocal areaRemote work$314.8k - $359.3k
...Senior Distinguished AI Engineer At Capital... ...capabilities to reimagine how we serve our customers and... ...large language model inference, similarity search,... ...state-of-the-art LLM optimization techniques to improve... ...- scalability, cost, latency, throughput - of large...Full timePart timeLocal area- ..., high-efficiency serving platform. Backed by... ...support from AMD engineers the team is scaling... ...AI models are optimized and deployed at scale... ...Design and build a low-latency, chat-like... ...engineering, open source inference engine like vLLM, Sglang, or TRT-LLM Streaming...Work at officeFlexible hours
$192k - $260k
...Foundation Model Serving is the API... ...frontier AI model inference for open source... ...re looking for engineers who have owned... ...deep building LLM APIs and runtimes... ...at scale.As a Staff Engineer, you’ll... ...throughput, low-latency inference on... ...trade-offs to optimize performance, throughput...Local areaWorldwide$260k - $340k
...Principal Systems Software Engineer, you will serve as the visionary... ...via zero-latency InfiniBand/RDMA fabrics... ...IaaS: Design highly optimized, thin virtualization... ...: Work alongside Staff and Senior engineers... ...Large Language Model (LLM) training and inference at scale.Peer-...Full timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff GenAI Inference Engineer: Optimize LLM Serving Latency. Be the first to apply!
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- engineering aide San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineering manager San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- project engineer assistant project manager San Francisco, CA



