Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Engineer

Hippocratic AI

About Us Hippocratic AI is the leading generative AI company in healthcare. We have the only system that can have safe, autonomous, clinical conversations with patients. We have trained our own LLMs as part of our Polaris constellation, resulting in a system with over 99.9% accuracy. About the Role We're seeking an experienced LLM Inference Engineer to optimize our large language model (LLM) serving infrastructure. The ideal candidate has: Extensive hands‑on experience with state‑of‑the‑art inference optimization techniques A track record of deploying efficient, scalable LLM systems in production environments What You'll Do Design and implement multi-node serving architectures for distributed LLM inference Optimize multi-LoRA serving systems Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality Implement speculative decoding and other latency optimization strategies Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases Continuously benchmark and improve system performance across various deployment scenarios and GPU types What You Bring Must-Have: Experience optimizing LLM inference systems at scale Proven expertise with distributed serving architectures for large language models Hands‑on experience implementing quantization techniques for transformer models Strong understanding of modern inference optimization methods, including: Speculative decoding techniques with draft models Eagle speculative decoding approaches Proficiency in Python and C++ Experience with CUDA programming and GPU optimization Nice‑to‑Have: Contributions to open‑source inference frameworks such as vLLM , SGLang , or TensorRT-LLM Experience with custom CUDA kernels Track record of deploying inference systems in production environments Deep understanding of performance optimization systems _Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters – we want to hear your story._ _Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value._ Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale! #J-18808-Ljbffr

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the LLM Inference Engineer in Palo Alto, CA vacancy
  •  ...Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack:...  ...unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly... 
    Suggested

    d-Matrix

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  •  ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated...  ...head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why... 
    Suggested

    AMD

    Santa Clara, CA
    4 days ago
  • $124k - $195.5k

    NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position...  ...and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $184k - $287.5k

     ...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who...  ...:Implement language and multimodal model inference as part of NVIDIA Inference Microservices...  ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference serving library... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based...  ...RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an expert...  ...inference stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization... 
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $152k - $241.5k

     ...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class...  ...problems for AI workloads (both inference and training) and successfully transition...  ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195.2k - $361.2k

     ...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments...  ...Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise will help... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $135k - $160k

     ...human life on Mars.APPLICATION SOFTWARE ENGINEER, INFERENCEThe application software team is...  ...Our team maintains a high-performance AI inference platform that serves the best models internally...  ...engines (e.g., SGLang, vLLM, TensorRT-LLM) Develop custom tools for tracing,... 
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    9 hours ago
  • $198k - $326k

     ...platforms that power AI across LinkedIn. The LLM Serving team builds the critical...  ...are looking for a Senior Staff Software Engineer with deep expertise at the intersection of...  ...learning, GPU infrastructure, and large-scale inference. This is a highly technical, high-leverage... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    4 days ago
  • $142.8k - $274.8k

     ...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s...  ...enable industry-leading AI training and inference. The Platform Systems Engineering (PSE)...  ...workloads including: HPL/HPC benchmarks LLM training workloads Transformer-based... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    1 day ago
  •  ...workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we...  ...infrastructure for building and serving LLM’s at Moveworks. This role will be critical...  ...including distributed training and inference pipeline for large language models(LLM),... 
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    4 days ago
  • $104k - $151k

     ...of that effort, from service desk support to process/platform engineering to the automation that connects it all.We're searching for an Automation...  ...queues for triggering automated workflowsExperience building LLM-driven automations that make decisions, call tools/APIs, or... 
    Permanent employment
    For contractors
    Work at office
    Local area
    3 days per week

    Aurora Innovation

    Mountain View, CA
    1 day ago
  •  ...the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,...  ...Neural Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in Electrical... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

    As a Senior Engineering Manager for Agentic Systems & Platform Architecture, you will lead the...  ..., and GPU-optimized training and inference workflows.Lead integration of the AI Data...  ...open-source libraries; deep expertise in LLM/agent architectures—leading POCs and integrating... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $148.5k - $223.9k

     ...Job DetailsWe are infusing AI into every aspect of our software engineering. While AI is writing the code, validation is what needs to be...  ...into actionable automation plans, focusing heavily on validating LLM integrations, prompt accuracy, and deterministic system... 
    Full time

    Salesforce

    Palo Alto, CA
    1 day ago
  • $250k - $350k

     ...alongside some of the world's leading ML systems engineers, including leaders behind Megatron-LM,...  ...powering our large-scale training, inference, and reinforcement learning. You'll own critical...  .... Familiarity with vLLM, TensorRT-LLM, or production LLM serving systems is a plus... 
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    4 days ago
  • $87.95k - $203.95k

     ...want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid to join our team in Santa Clara, California (US-CA), United States (US).This API LLM Integration Engineer... 
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full...  ...engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $120k - $250k

     ...custom silicon for large-language-model inference and training, with HW/SW co-design across...  ...integration path for downstream consumersBuild the LLM inference serving stack — paged KV cache,...  ..., traces, and the Python surfaces ML engineers actually use — and hit measurable... 
    Full time
    Contract work
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    3 days ago
  • $224k - $356.5k

     ...platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud providers, and...  ...scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo organization. Your responsibility... 
    Full time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

    NVIDIA is seeking an exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI model deployment. You...  ..., and optimization of large-scale models for LLM, multimodal, and generative AI applications.Guide engineers... 
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • $207k - $301k

     ...a high-performing team of systems and ML engineers. Drive a culture of excellence, psychological...  ...goal and strategy for enhancing the LLM serving stack, focusing on performance, scalability...  ....The Distributed Cloud (DSC) AI Inference Platform team operates at the critical intersection... 

    Google

    Sunnyvale, CA
    4 days ago
  • $224k - $356.5k

     ...technical software manager to lead production AI inference for NVIDIA Inference Microservices (NIM),...  ...stack, combining optimized inference engines, model profiles/recipes, validated...  ...responsible for shipping production‑ready LLM NIMs, including planning, new model onboarding... 

    Socket.dev

    Santa Clara, CA
    13 hours ago
  • $168k - $270.25k

     ...including NeMo microservices and NVIDIA Inference Microservices (NIM), enabling scalable, production...  ..., technically strong test development engineer to drive quality, automation, and...  ...architecture and improve testabilityValidate LLM and AI inference workflows, including... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment...  ...-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe... 
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $152k - $241.5k

     ...pushing the limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team and help shape the next generation of edge AI for...  ...experience in Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of relevant software... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Engineer. Be the first to apply!