LLM Inference Engineer
Hippocratic AI
About Us Hippocratic AI is the leading generative AI company in healthcare. We have the only system that can have safe, autonomous, clinical conversations with patients. We have trained our own LLMs as part of our Polaris constellation, resulting in a system with over 99.9% accuracy. About the Role We're seeking an experienced LLM Inference Engineer to optimize our large language model (LLM) serving infrastructure. The ideal candidate has: Extensive hands‑on experience with state‑of‑the‑art inference optimization techniques A track record of deploying efficient, scalable LLM systems in production environments What You'll Do Design and implement multi-node serving architectures for distributed LLM inference Optimize multi-LoRA serving systems Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality Implement speculative decoding and other latency optimization strategies Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases Continuously benchmark and improve system performance across various deployment scenarios and GPU types What You Bring Must-Have: Experience optimizing LLM inference systems at scale Proven expertise with distributed serving architectures for large language models Hands‑on experience implementing quantization techniques for transformer models Strong understanding of modern inference optimization methods, including: Speculative decoding techniques with draft models Eagle speculative decoding approaches Proficiency in Python and C++ Experience with CUDA programming and GPU optimization Nice‑to‑Have: Contributions to open‑source inference frameworks such as vLLM , SGLang , or TensorRT-LLM Experience with custom CUDA kernels Track record of deploying inference systems in production environments Deep understanding of performance optimization systems _Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters – we want to hear your story._ _Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value._ Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale! #J-18808-Ljbffr
- ...Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack:... ...unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly...Suggested
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work...SuggestedFull time$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why...Suggested
$124k - $195.5k
NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position... ...and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms...SuggestedFull time$184k - $287.5k
...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...:Implement language and multimodal model inference as part of NVIDIA Inference Microservices... ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference serving library...Full time- ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based... ...RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an expert... ...inference stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization...Contract workShift work
$152k - $241.5k
...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...problems for AI workloads (both inference and training) and successfully transition... ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language...Full time$195.2k - $361.2k
...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments... ...Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling...Full timeInternshipLocal areaImmediate startShift work$152k - $241.5k
...technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise will help...Full time$135k - $160k
...human life on Mars.APPLICATION SOFTWARE ENGINEER, INFERENCEThe application software team is... ...Our team maintains a high-performance AI inference platform that serves the best models internally... ...engines (e.g., SGLang, vLLM, TensorRT-LLM) Develop custom tools for tracing,...Permanent employmentTemporary workRemote workWorldwideWeekend work$198k - $326k
...platforms that power AI across LinkedIn. The LLM Serving team builds the critical... ...are looking for a Senior Staff Software Engineer with deep expertise at the intersection of... ...learning, GPU infrastructure, and large-scale inference. This is a highly technical, high-leverage...For contractorsWork at officeFlexible hours$142.8k - $274.8k
...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s... ...enable industry-leading AI training and inference. The Platform Systems Engineering (PSE)... ...workloads including: HPL/HPC benchmarks LLM training workloads Transformer-based...Ongoing contractWork at officeLocal areaWorldwide3 days per week- ...workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we... ...infrastructure for building and serving LLM’s at Moveworks. This role will be critical... ...including distributed training and inference pipeline for large language models(LLM),...Work at officeRemote workFlexible hours
$104k - $151k
...of that effort, from service desk support to process/platform engineering to the automation that connects it all.We're searching for an Automation... ...queues for triggering automated workflowsExperience building LLM-driven automations that make decisions, call tools/APIs, or...Permanent employmentFor contractorsWork at officeLocal area3 days per week- ...the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,... ...Neural Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in Electrical...3 days per week
$272k - $431.25k
As a Senior Engineering Manager for Agentic Systems & Platform Architecture, you will lead the... ..., and GPU-optimized training and inference workflows.Lead integration of the AI Data... ...open-source libraries; deep expertise in LLM/agent architectures—leading POCs and integrating...Full time$148.5k - $223.9k
...Job DetailsWe are infusing AI into every aspect of our software engineering. While AI is writing the code, validation is what needs to be... ...into actionable automation plans, focusing heavily on validating LLM integrations, prompt accuracy, and deterministic system...Full time$250k - $350k
...alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM,... ...powering our large-scale training, inference, and reinforcement learning. You'll own critical... .... Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems is a plus...Visa sponsorship$87.95k - $203.95k
...want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid to join our team in Santa Clara, California (US-CA), United States (US).This API LLM Integration Engineer...Temporary workWork at officeRemote workFlexible hours$184k - $287.5k
...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing...Full timeWork experience placement$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full... ...engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across...Full time$120k - $250k
...custom silicon for large-language-model inference and training, with HW/SW co-design across... ...integration path for downstream consumersBuild the LLM inference serving stack — paged KV cache,... ..., traces, and the Python surfaces ML engineers actually use — and hit measurable...Full timeContract workWork experience placementLocal areaRemote workMonday to FridayFlexible hours$224k - $356.5k
...platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud providers, and... ...scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo organization. Your responsibility...Full timeLocal areaRemote workWorldwide$224k - $356.5k
NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You... ..., and optimization of large-scale models for LLM, multimodal, and generative AI applications.Guide engineers...Full timeRemote workWorldwide$207k - $301k
...a high-performing team of systems and ML engineers. Drive a culture of excellence, psychological... ...goal and strategy for enhancing the LLM serving stack, focusing on performance, scalability... ....The Distributed Cloud (DSC) AI Inference Platform team operates at the critical intersection...$224k - $356.5k
...technical software manager to lead production AI inference for NVIDIA Inference Microservices (NIM),... ...stack, combining optimized inference engines, model profiles/recipes, validated... ...responsible for shipping production‑ready LLM NIMs, including planning, new model onboarding...$168k - $270.25k
...including NeMo microservices and NVIDIA Inference Microservices (NIM), enabling scalable, production... ..., technically strong test development engineer to drive quality, automation, and... ...architecture and improve testabilityValidate LLM and AI inference workflows, including...Full time$152k - $241.5k
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment... ...-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe...Full time$152k - $241.5k
...pushing the limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for... ...experience in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant software...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Engineer. Be the first to apply!


