Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Engineer (Mid, Sr, Staff)

Hippocratic AI Inc.

Role Mission

As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses-makings difference between conversational experiences that feel natural and those that feel broken. This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems.

What You Will Accomplish

Own your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduce latency, improve throughput, or optimize cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment.

Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel optimization techniques that become part of our core infrastructure, and made our serving stack a durable competitive advantage in healthcare AI deployment.

The Team

You’ll work alongside systems engineers, ML researchers, and infrastructure experts who are obsessed with making AI systems fast, reliable, and cost-effective. This is a team that values deep technical rigor, continuous benchmarking, and solving hard systems problems that have real impact on patient experience and business unit economics.

What You’ll Do

  • Design and implement multi-node serving architectures for distributed LLM inference

  • Optimize multi-LoRA serving systems

  • Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality

  • Implement speculative decoding and other latency optimization strategies

  • Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases

  • Continuously benchmark and improve system performance across various deployment scenarios and GPU types

Location Requirement

We believe the best ideas happen together. This role is based in our Menlo Park, California office, expected to be in office five days a week.

Compensation

Compensation is based on experience, expertise, and level of responsibility. We offer competitive packages that reflect the seniority and scope of the role, along with equity, health insurance, and other benefits.

What You Bring

Must-Have:

  • Experience optimizing LLM inference systems at scale

  • Proven expertise with distributed serving architectures for large language models

  • Hands-on experience implementing quantization techniques for transformer models

  • Strong understanding of modern inference optimization methods, including:

    • Speculative decoding techniques with draft models

    • Eagle speculative decoding approaches

  • Proficiency in Python and C++

  • Experience with CUDA programming and GPU optimization

Nice-to-Have:

  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM

  • Experience with custom CUDA kernels

  • Track record of deploying inference systems in production environments

  • Deep understanding of performance optimization systems

Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters – we want to hear your story.

Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value.

Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale!

References

  • 1. Polaris: A Safety-focused LLM Constellation Architecture for Healthcare,

  • 2. Polaris 2:

  • 3. Personalized Interactions:

  • 4. Human Touch in AI:

  • 5. Empathetic Intelligence:

Why Join Hippocratic AI

Reinvent healthcare with AI that puts safety first. We're building the world's first healthcare-only, safety-focused LLM – a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co-founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children's, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies – ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact View email address on click.appcast.io.

#J-18808-Ljbffr

Vacancy posted 3 hours ago
Similar jobs that could be interesting for youBased on the LLM Inference Engineer (Mid, Sr, Staff) in Menlo Park, CA vacancy
  •  ...companies. As a Senior Principal Software Engineer at JPMorganChase within the Commercial &...  ...deployment and optimization using model inference servers such as Triton Inference Server...  ...Demonstrated success architecting and deploying LLM & GNN solutions on AWS (e.g., SageMaker,... 
    Senior

    JPMorgan Chase & Co.

    Palo Alto, CA
    a month ago
  •  ...architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud...  ...-speed inference.About The RoleWe’re hiring a Senior Frontend Engineer to own and scale critical parts of the Cerebras Developer... 
    Senior
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    3 days ago
  • $195.2k - $262.2k

     ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across...  ...regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput,... 
    Senior
    Full time
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    3 days ago
  • $195.2k - $262.2k

     ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across...  ...Define and execute research programs in efficient LLM and VLM inference with measurable production... 
    Senior
    Full time
    Temporary work
    Immediate start
    Remote work

    Nebius

    Palo Alto, CA
    4 days ago
  •  ...Sales Engineer The Sales Engineer is the Technical Lead for the Proofpoint Pre-Sales process...  ...closely with customer/prospect security staff for technical discovery, as well as...  ...Pacific Northwest. This role will focus on Mid-Market accounts in the Bay Area and Pacific... 
    Senior
    Local area
    Flexible hours
    Night shift

    Proofpoint

    Sunnyvale, CA
    17 hours ago
  • $272k - $336k

     ...autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Systems Engineering team works together to blend software and hardware systems in groundbreaking new ways. We set the high performance standards that... 
    Senior
    Odd job
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $86.25k - $145k

    We are hiring a highly technical, business-savvy Senior GTM Engineer to help build the next generation of Navan’s GTM systems. This hands...  ....At least 2 years of hands-on experience building and deploying LLM-powered agents, copilots, or automations used by business teams... 
    Senior

    TripActions

    Palo Alto, CA
    3 days ago
  • $45.67 - $55.29 per hour

     ...or 40/week ~ Performance bonuses paid mid-year and year-end ~401(k) profit-sharing program ZFA Structural Engineers is a mid-sized structural engineering firm...  ...sustainably over the last decade and now employ a staff of 100 people. We work hard to provide a... 
    Senior
    Full time
    Work at office

    ZFA Structural Engineers

    Redwood City, CA
    5 hours ago
  •  ...Senior Forward Deployed Engineer (FDE)Location: Hybrid, Palo Alto, CAAs a Senior Forward Deployed Engineer (Sr. FDE) at Uniphore, you will take technical ownership of strategic...  ...auditable AI agents.Stay current on RAG, LLM/SLM advancements, Agentic AI tooling, vector... 
    Senior

    Tranzeal

    Palo Alto, CA
    3 days ago
  • $175k - $240k

     ...Business AI technology company, is seeking an experienced Senior Sales Engineer to join its high-performance Sales Engineering team to support...  ...presentations, demos showcasing GenAI, Agentic AI, RAG, and LLM-based solutions.  Build and deliver custom demos and POCs to... 
    Senior
    Full time

    Uniphore

    Palo Alto, CA
    2 hours ago
  •  ...San Carlos, California. Relocation benefits available. Summary Delta Star Inc. is seeking a skilled and innovative Sr. Electrical Design Engineer to design and engineer medium voltage, mobile power transformers and mobile substations. If you're passionate about... 
    Senior
    Relocation
    Relocation package

    Clear Destination

    San Carlos, CA
    3 days ago
  •  ...for everyone. ABOUT THE ROLE As a Senior Forward Deployed Engineer, you’ll deliver Retell’s voice AI solutions to enterprise clients...  ...Become go-to experts on building voice AI solutions, especially LLM prompting tips & tricks. Build full-stack integrations (... 
    Senior
    H1b
    Work at office

    Retell AI

    Redwood City, CA
    2 days ago
  •  ...great investors. If you want to disrupt the landscape of healthcare, join our team. We're looking for a Senior Quality Engineer to own quality on a mid-field MRI system as we drive toward a 510(k) submission. This is a foundational hire: you'll stand up and run the... 
    Senior
    Relocation package

    Adialante LLC

    Menlo Park, CA
    2 days ago
  • $119.8k - $234.7k

     ...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft...  ...-leading AI training and inference. The Platform Systems Engineering (PSE) team is seeking a Sr. AI Accelerator Tools Development...  ...including: HPL/HPC benchmarks LLM training workloads Transformer... 
    Senior
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    1 day ago
  •  ...Responsibilities About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model...  ...GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).Strong background in system optimizations: batching... 
    Senior
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    2 hours ago
  • $132.1k - $165.1k

     ...with manufacturing, validation, service, suppliers, and program teams to investigate field and production issues, develop robust engineering changes, and verify solutions through analysis and test. The ideal candidate brings strong engineering judgment, hands‑on problem... 
    Senior
    Full time
    Contract work
    Temporary work
    Part time
    Local area
    Shift work

    Rivian

    Palo Alto, CA
    2 days ago
  •  ...is delivered for millions of patients worldwide.Were a team of engineers, clinicians, and innovators united by one purpose: to make surgery...  ...compensation ranges are listed.SummaryType: Full-timeFunction: EngineeringExperience level: Mid-Senior LevelIndustry: Medical Device
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    5 days ago
  •  ...and implementation of technical safety requirements at the hardware level. Work with cross functional teams and hardware design engineers to implement hardware safety requirements. Work with various functional safety analysis methods (FTA, FMEDA) supporting... 
    Senior
    Full time
    Contract work

    System Safety Inc

    Palo Alto, CA
    3 days ago
  • $120k - $150k

     ...ll Do Join DES as a Senior Structural Engineer and play an important role in delivering...  ...technical guidance to engineers and developing staff Contribute to the professional...  ...supporting the development of junior and mid-level engineers OSHPD and DSA experience... 
    Senior
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Flexible hours

    AIA San Mateo County

    Redwood City, CA
    4 hours ago
  • $210k

     ...Sr. Reliability Engineer Location: Mountain View/San Jose Bay Area Team: System Engineering The Impact You’ll Make: We are seeking...  ...his interest in Avantus’ development business to KKR in mid-2024, it had delivered over $1 billion in profit and secured... 
    Senior
    Work at office
    Local area
    Night shift

    1st Avenue Power

    Mountain View, CA
    4 hours ago
  •  ...rigor to a rapidly evolving AI inference space. Our mission is to make...  ..., and speak fluently to both engineering and procurement — someone who...  ...— credible with a skeptical staff engineer, clear with a CFO,...  ...internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math,... 

    Deepinfra

    Palo Alto, CA
    5 days ago
  • $156k - $183k

     ...companies, and the planet. Role Description As a Senior Value Engineer specializing in the Public Sector, you are pushing the envelope...  .... This includes architecting and delivering secure, scalable LLM/agent systems with RAG, tools, and guardrails, while seamlessly... 
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Celonis

    Redwood City, CA
    4 days ago
  •  ...Salesforce, merchants' websites, and other platforms.Job RequirementsWe are looking for curious, energetic rock stars to join our engineering team.You will use scenario-driven testing methodology to develop, execute and maintain test cases according to the business... 
    Senior

    Central Business Solutions

    Palo Alto, CA
    2 days ago
  • $260k

     ...Mid/Senior-Level Transactional Tax Associate Baker McKenzie’s Tax Practice is a leader within our Firm and one of the most highly regarded...  ...structuring. Top academic credentials are a prerequisite; an LLM in taxation is not required, but may be favorably considered. Requirements... 
    Senior
    Hourly pay
    Full time
    Work at office
    Local area
    Worldwide

    Baker McKenzie

    Palo Alto, CA
    2 days ago
  • $160k - $200k

     ...infrastructure to build training and serving pipelines is also a key aspect to this role.  At Dexterity, our use cases require low latency inference that approaches real time and therefore you must be able to build highly performant models and serving architectures. While an... 
    Senior
    Work experience placement

    Dexterity

    Redwood City, CA
    11 days ago
  •  ...we want you on our team! Role Summary The Field Service Engineer (FSE) , along with the Field Application Scientists (FAS) are the...  ...Customer Training: Provide hands-on training to customer lab staff on proper instrument operation, routine care, and... 
    Senior
    Full time
    Contract work
    Remote work
    Home office
    Night shift

    Countable Labs

    Palo Alto, CA
    more than 2 months ago
  • $180k - $240k

    Sr. Computer Vision Engineer (Deep Learning) About Harbinger Harbinger is an American commercial electric vehicle (EV) company on a mission to...  ...vision applications. In‑depth understanding of training and inference pipelines, including data loading, augmentation, and loss... 
    Senior
    Local area

    Harbinger Motors

    Mountain View, CA
    2 days ago
  •  ...Vision-Language Model (VLM) & Visual Foundation Model (VFM) Field Engineer About Matroid Matroid helps enterprises build and...  ...can operate in cloud, edge, and hybrid environments, optimizing inference latency, GPU utilization, memory consumption, and throughput... 
    Senior
    Full time
    Work at office

    Matroid

    Palo Alto, CA
    18 days ago
  •  ...Work closely with the Electrical Hardware Design team to create engineering and delivery schedules, closely track engineering progress, and...  ...issues ~ Strong communication experience with executive staff members ~ Proven experience tracking and driving the progress... 
    Full time
    Contract work
    Local area

    Rivian VW Group

    Palo Alto, CA
    3 days ago
  • $180k - $240k

     ...Foundation. We are seeking a highly skilled Senior Deep Learning Engineer to drive the development and deployment of advanced perception...  ...applications. ~ In-depth understanding of training and inference pipelines, including data loading, augmentation, and loss... 
    Senior
    Full time
    Contract work
    Local area

    Harbinger Motors Inc.

    Mountain View, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Engineer (Mid, Sr, Staff). Be the first to apply!