Staff GenAI Inference Engineer: Optimize LLM Serving Latency
$190.9k - $232.8kJobleads-US
A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong software engineering background and a proven ability to collaborate with researchers and drive architectural decisions. Competitive compensation is offered, with a salary range of $190,900 to $232,800 USD. #J-18808-Ljbffr Jobleads-US
- ...Our team analyzes inference stack performance across... ...understanding into performance optimizations and models that... ...You will build cost-to-serve estimates from... ...functional teams reason about latency, capacity, utilization... ...collaborating with engineering and research teams to...SuggestedFull time
- ...San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build... ...infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability. This role...Suggested
- ...the Team OpenAI’s Inference team powers the... ...fast-moving team of engineers focused on delivering... ...needed to serve models that handle... ...You'll build and optimize the systems that let... ...high-throughput, low-latency delivery of image... ...like vLLM, TensorRT-LLM, or custom model parallel...SuggestedFull time
- ...fast-growing team of engineers in San Francisco... ...models Low-latency OCR models for targeted... ..., high-throughput inference for OCR and... ...batching and caching Optimize kernels,... ...Evaluate vLLM, TensorRT LLM, and Triton tradeoffs... ...profiling and model serving Nice to...SuggestedVisa sponsorshipRelocation package
$190k - $265k
...business. Founded by engineers — and customer-... ...model training, model serving, and Vector Search... ...Foundation Model Inference team is the... ...serve, scale, and optimize frontier models with... ...you will have:Build LLM infrastructure powering... ...reliability, latency, and efficiency of...SuggestedLocal areaWorldwide$215k - $260k
...means owning the inference stack end to end:... ...bringing modern optimization techniques into real... ...deep into the serving code when the defaults... ...patterns, latency targets, and cost... ...directly with customer engineering teams to tailor... ...with modern LLM serving frameworks...Temporary work- ...a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams... ...systems responsible for serving models reliably, efficiently... ...launches, inference optimizations, cloud provider integrations... ...in correctness, latency, or reliability meaningfully...Full time
- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI... ...Develop world-class model serving stack for state-of-the... ...end-to-end and tail latency (p95/p99), increase throughput... ..., and server-level optimizations. Build large-scale,...Full timeFlexible hours
$186k - $269k
Build GenAI PoCs (LLMs, RAG, agentic frameworks) to... ...software solutions.Serve as engineering Lead delivering complex... ...safety, compliance, and latency.Minimum qualifications... ...experience building LLM evaluation (evals) pipelines... ..., observability, and optimizing LLM-native metrics at...$170k - $216k
...of customers Software Engineers, Product, Data Science... ...Build and evolve ML inference infrastructure for simulations... ...for the reliability, latency, and user experience... ...model deployment and serving. You have: ~... ...frameworks, TPUs and optimizing models for serving....Full timeRemote work$190.9k - $232.8k
...1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture... ..., development, and optimization of the inference engine... ...high throughput, low latency, and robust scaling. Your... ...and collaborate on model-serving stack optimized for large...Local areaWorldwide$220k
We build and run the inference engine behind every Perplexity... ...architectures at scale with tight latency and cost budgets. Our... .... Rust-native serving runtime. Develop our... ...You understand modern LLM architectures and are... ...and inference optimization techniques (e.g. quantization...- LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for... ...while integrating with model-serving infrastructure. The role emphasizes... ...frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup...
- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI... ...that powers large-scale LLM inference across our platform... ...to make new inference optimizations broadly available to... ...where reliability, latency, and scale are first-...Full timeFlexible hours
- ...About the Team Our Inference team brings OpenAI’s most capable research and technology... .... About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure... .... Integrate internal model-serving infrastructure (e.g., vLLM, Triton)...Full time
- ...data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building... ...distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates will have...
- ...ll doAs a Software Engineer on the AI Platform... ...processing, model serving, and data... ...and automated model optimization.This position is an... ...large-scale model inference and data processingDesign... ...strict uptime and latency... ...level services for LLM orchestration, RAG...Contract workWork at officeLocal areaRemote work2 days per week
$137.1k - $201.6k
...DoorDash’s GenAI Platform team sits... ...real-time GPU serving, high-throughput batch inference, and fine-tuning... ...large cost and latency wins (for example... ...including the LLM Gateway, Agent... ...and inference engines, fine-tuning and... ...RLHF/RLVR), agent optimization, and other post...Hourly payWork at officeLocal areaRemote workFlexible hours- ...model runtime within the inference engine that executes complex, frontier... ...layers of the cluster serving software stack,... ...efficient execution while optimizing for throughput, latency, utilization, and reliability... ...Design and implement the LLM inference runtime for frontier...Full time
$188k - $275k
...the team: The Inference team is responsible... ...performance model serving capabilities that... ...improve throughput, latency, reliability, and... ...for an Applied AI Engineer to help us understand... ...driving targeted optimizations for both platform-... ...Familiarity with LLM inference systems...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours- ...powers mission-critical inference for the world's most... ...build the platform engineers turn to to ship AI products... ...systems, model serving, and developer experience... ...Profile and optimize TensorRT-LLM kernels, analyze CUDA... ...record of owning low‑latency, reliable backend services...Full timeFlexible hours
$197.3k - $225.1k
...Overview Lead AI Engineer (AI Foundations, LLM Core and Agentic AI)... ...to reimagine how we serve our customers and businesses... ...language model inference, similarity search,... ...-of-the-art LLM optimization techniques to improve... ...scalability, cost, latency, throughput — of large...Full timePart timeLocal area- ...TeamDoorDash’s GenAI Platform team sits... ...real-time GPU serving, high-throughput batch inference, and fine-... ...large cost and latency wins (for example... ...including the LLM Gateway, Agent... ...and inference engines, fine-tuning and... ...RLHF/RLVR), agent optimization, and other post...Hourly payWork at officeLocal areaRemote workFlexible hours
$100k - $300k
...team, founded by the engineers who scaled... ...Implement and iterate on LLM, VLM, and VLA... ...input/tokenization, inference runners, and output... ...production-grade serving and inference tooling... ..., low-latency operation on bench... ...Performance & optimization: Practical work improving...Full time$286.2k - $326.7k
Sr. Distinguished AI Engineer (Remote Eligible)... ...capabilities to reimagine how we serve our customers and... ...large language model inference, similarity search, guardrails... ...state-of-the-art LLM optimization techniques to improve... ...- scalability, cost, latency, throughput - of large...Full timePart timeLocal areaRemote work$150k - $170k
...oriented Applied AI Engineer to help develop,... ...model integration, inference services,... ...acquisition, model serving, monitoring, lifecycle... ..., and performance optimization. Learn quickly... ...learning models, LLM applications, computer... ...edge deployment, low-latency inference,...Work experience placementCasual workWork at officeRelocation package2 days per week- ...build, train, and serve AI models... ...them to a working inference call fastest, and... ...are a working engineer: you ship code,... ...the Senior or Staff level and will... ...ve worked with LLM APIs in production... ...what inference latency, throughput,... ...Depth in inference optimization, serving...Permanent employmentFlexible hours
- ..., high-efficiency serving platform. Backed by... ...support from AMD engineers the team is scaling... ...AI models are optimized and deployed at scale... ...Design and build a low-latency, chat-like... ...engineering, open source inference engine like vLLM, Sglang, or TRT-LLM Streaming...Work at officeFlexible hours
$260k - $340k
...Principal Systems Software Engineer, you will serve as the visionary... ...via zero-latency InfiniBand/RDMA fabrics... ...IaaS: Design highly optimized, thin virtualization... ...: Work alongside Staff and Senior engineers... ...Large Language Model (LLM) training and inference at scale.Peer-...Full timeTemporary work$192k - $260k
...Foundation Model Serving is the API... ...frontier AI model inference for open source... ...re looking for engineers who have owned... ...deep building LLM APIs and runtimes... ...at scale.As a Staff Engineer, you’ll... ...throughput, low-latency inference on... ...trade-offs to optimize performance, throughput...Local areaWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff GenAI Inference Engineer: Optimize LLM Serving Latency. Be the first to apply!
- project engineer assistant project manager San Francisco, CA
- senior staff systems engineer San Francisco, CA
- staff data engineer San Francisco, CA
- assistant chief engineer San Francisco, CA
- assistant engineer San Francisco, CA
- assistant electrical engineer San Francisco, CA
- engineering aide San Francisco, CA
- software engineer staff San Francisco, CA
- staff design engineer San Francisco, CA
- research assistant engineering San Francisco, CA


