Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Engineer (Mid, Sr, Staff)

Jobleads-US

Role Mission

As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses-makings difference between conversational experiences that feel natural and those that feel broken. This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems.

What You Will Accomplish

Own your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduce latency, improve throughput, or optimize cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment.

Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel optimization techniques that become part of our core infrastructure, and made our serving stack a durable competitive advantage in healthcare AI deployment.

The Team

You’ll work alongside systems engineers, ML researchers, and infrastructure experts who are obsessed with making AI systems fast, reliable, and cost-effective. This is a team that values deep technical rigor, continuous benchmarking, and solving hard systems problems that have real impact on patient experience and business unit economics.

What You’ll Do

  • Design and implement multi-node serving architectures for distributed LLM inference

  • Optimize multi-LoRA serving systems

  • Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality

  • Implement speculative decoding and other latency optimization strategies

  • Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases

  • Continuously benchmark and improve system performance across various deployment scenarios and GPU types

Location Requirement

We believe the best ideas happen together. This role is based in our Menlo Park, California office, expected to be in office five days a week.

Compensation

Compensation is based on experience, expertise, and level of responsibility. We offer competitive packages that reflect the seniority and scope of the role, along with equity, health insurance, and other benefits.

What You Bring

Must-Have:

  • Experience optimizing LLM inference systems at scale

  • Proven expertise with distributed serving architectures for large language models

  • Hands-on experience implementing quantization techniques for transformer models

  • Strong understanding of modern inference optimization methods, including:

    • Speculative decoding techniques with draft models

    • Eagle speculative decoding approaches

  • Proficiency in Python and C++

  • Experience with CUDA programming and GPU optimization

Nice-to-Have:

  • Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM

  • Experience with custom CUDA kernels

  • Track record of deploying inference systems in production environments

  • Deep understanding of performance optimization systems

Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters – we want to hear your story.

Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value.

Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale!

References

  • 1. Polaris: A Safety-focused LLM Constellation Architecture for Healthcare,

  • 2. Polaris 2:

  • 3. Personalized Interactions:

  • 4. Human Touch in AI:

  • 5. Empathetic Intelligence:

Why Join Hippocratic AI

Reinvent healthcare with AI that puts safety first. We're building the world's first healthcare-only, safety-focused LLM – a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.

Work with the people shaping the future. Hippocratic AI was co-founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.

Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children's, WellSpan Health, John Doerr, Rick Klausner, and others.

Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies – ensuring our platform is powerful, trusted, and truly transformative.

Equal Opportunity

Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact View email address on click.appcast.io.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the LLM Inference Engineer (Mid, Sr, Staff) in Menlo Park, CA vacancy
  •  ...to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly... 
    Suggested
    Work at office

    Foundation Capital

    Palo Alto, CA
    3 days ago
  •  ...Hippocratic AI Inc. is seeking an experienced LLM Inference Engineer to own and optimize its inference serving stack. You will drive sub-100ms responses, cost efficiency, and reliable availability for millions of patient conversations across healthcare systems. You... 
    Senior

    Jobleads-US

    Menlo Park, CA
    1 day ago
  •  ...companies. As a Senior Principal Software Engineer at JPMorganChase within the Commercial &...  ...deployment and optimization using model inference servers such as Triton Inference Server...  ...Demonstrated success architecting and deploying LLM & GNN solutions on AWS (e.g., SageMaker,... 
    Senior

    JPMorgan Chase & Co.

    Palo Alto, CA
    a month ago
  •  ...architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud...  ...-speed inference.About The RoleWe’re hiring a Senior Frontend Engineer to own and scale critical parts of the Cerebras Developer... 
    Senior
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $167.2k - $316.6k

     ...pursue their dreams. Ford is seeking a Sr. Staff IVI Product Strategy Lead with a strong...  ...(IVI) product roadmap across near-, mid-, and long-term priorities; guide internal...  ...What you'll do... Partner with product, engineering, design, go-to-market and vehicle... 
    Senior
    Immediate start
    Visa sponsorship
    Flexible hours

    Jobleads-US

    Palo Alto, CA
    1 day ago
  • $272k - $336k

     ...autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Systems Engineering team works together to blend software and hardware systems in groundbreaking new ways. We set the high performance standards that... 
    Senior
    Odd job
    Full time
    Remote work

    Waymo

    Mountain View, CA
    3 days ago
  • $86.25k - $145k

    We are hiring a highly technical, business-savvy Senior GTM Engineer to help build the next generation of Navan’s GTM systems. This hands...  ....At least 2 years of hands-on experience building and deploying LLM-powered agents, copilots, or automations used by business teams... 
    Senior

    TripActions

    Palo Alto, CA
    4 days ago
  • $45.67 - $55.29 per hour

     ...or 40/week ~ Performance bonuses paid mid-year and year-end ~401(k) profit-sharing program ZFA Structural Engineers is a mid-sized structural engineering firm...  ...sustainably over the last decade and now employ a staff of 100 people. We work hard to provide a... 
    Senior
    Full time
    Work at office

    ZFA Structural Engineers

    Redwood City, CA
    1 day ago
  •  ...Senior Forward Deployed Engineer (FDE)Location: Hybrid, Palo Alto, CAAs a Senior Forward Deployed Engineer (Sr. FDE) at Uniphore, you will take technical ownership of strategic...  ...auditable AI agents.Stay current on RAG, LLM/SLM advancements, Agentic AI tooling, vector... 
    Senior

    Tranzeal

    Palo Alto, CA
    4 days ago
  • $175k - $240k

     ...Business AI technology company, is seeking an experienced Senior Sales Engineer to join its high-performance Sales Engineering team to support...  ...presentations, demos showcasing GenAI, Agentic AI, RAG, and LLM-based solutions.  Build and deliver custom demos and POCs to... 
    Senior
    Full time

    Uniphore

    Palo Alto, CA
    1 day ago
  •  ...San Carlos, California. Relocation benefits available. Summary Delta Star Inc. is seeking a skilled and innovative Sr. Electrical Design Engineer to design and engineer medium voltage, mobile power transformers and mobile substations. If you're passionate about... 
    Senior
    Relocation
    Relocation package

    Clear Destination

    San Carlos, CA
    14 hours ago
  •  ...for everyone. ABOUT THE ROLE As a Senior Forward Deployed Engineer, you’ll deliver Retell’s voice AI solutions to enterprise clients...  ...Become go-to experts on building voice AI solutions, especially LLM prompting tips & tricks. Build full-stack integrations (... 
    Senior
    H1b
    Work at office

    Retell AI

    Redwood City, CA
    3 days ago
  •  ...Overview The Senior Project Engineer is responsible for ensuring administrative, contractual, financial and technical aspects of the...  ...to work in the United States. Job Details Seniority level: Mid-Senior level Employment type: Full-time Job function: Engineering... 
    Senior
    Full time
    Contract work
    For subcontractor
    Work at office

    Level 10 Construction

    Sunnyvale, CA
    1 day ago
  •  ...great investors. If you want to disrupt the landscape of healthcare, join our team. We're looking for a Senior Quality Engineer to own quality on a mid-field MRI system as we drive toward a 510(k) submission. This is a foundational hire: you'll stand up and run the... 
    Senior
    Relocation package

    Adialante LLC

    Menlo Park, CA
    3 days ago
  • $119.8k - $234.7k

     ...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft...  ...-leading AI training and inference. The Platform Systems Engineering (PSE) team is seeking a Sr. AI Accelerator Tools Development...  ...including: HPL/HPC benchmarks LLM training workloads Transformer... 
    Senior
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    2 days ago
  •  ...Responsibilities About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model...  ...GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).Strong background in system optimizations: batching... 
    Senior
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    1 day ago
  • $132.1k - $165.1k

     ...with manufacturing, validation, service, suppliers, and program teams to investigate field and production issues, develop robust engineering changes, and verify solutions through analysis and test. The ideal candidate brings strong engineering judgment, hands‑on problem... 
    Senior
    Full time
    Contract work
    Temporary work
    Part time
    Local area
    Shift work

    Rivian

    Palo Alto, CA
    3 days ago
  •  ...is delivered for millions of patients worldwide.Were a team of engineers, clinicians, and innovators united by one purpose: to make surgery...  ...compensation ranges are listed.SummaryType: Full-timeFunction: EngineeringExperience level: Mid-Senior LevelIndustry: Medical Device
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    Sunnyvale, CA
    1 day ago
  •  ...and implementation of technical safety requirements at the hardware level. Work with cross functional teams and hardware design engineers to implement hardware safety requirements. Work with various functional safety analysis methods (FTA, FMEDA) supporting... 
    Senior
    Full time
    Contract work

    System Safety Inc

    Palo Alto, CA
    14 hours ago
  • $120k - $150k

     ...ll Do Join DES as a Senior Structural Engineer and play an important role in delivering...  ...technical guidance to engineers and developing staff Contribute to the professional...  ...supporting the development of junior and mid-level engineers OSHPD and DSA experience... 
    Senior
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Flexible hours

    AIA San Mateo County

    Redwood City, CA
    1 day ago
  • $210k

     ...Sr. Reliability Engineer Location: Mountain View/San Jose Bay Area Team: System Engineering The Impact You’ll Make: We are seeking...  ...his interest in Avantus’ development business to KKR in mid-2024, it had delivered over $1 billion in profit and secured... 
    Senior
    Work at office
    Local area
    Night shift

    1st Avenue Power

    Mountain View, CA
    1 day ago
  •  ...rigor to a rapidly evolving AI inference space. Our mission is to make...  ..., and speak fluently to both engineering and procurement — someone who...  ...— credible with a skeptical staff engineer, clear with a CFO,...  ...internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math,... 

    Deepinfra

    Palo Alto, CA
    4 days ago
  • $260k

     ...Mid/Senior-Level Transactional Tax Associate Baker McKenzie’s Tax Practice is a leader within our Firm and one of the most highly regarded...  ...structuring. Top academic credentials are a prerequisite; an LLM in taxation is not required, but may be favorably considered. Requirements... 
    Senior
    Hourly pay
    Full time
    Work at office
    Local area
    Worldwide

    Baker McKenzie

    Palo Alto, CA
    3 days ago
  •  ...Salesforce, merchants' websites, and other platforms.Job RequirementsWe are looking for curious, energetic rock stars to join our engineering team.You will use scenario-driven testing methodology to develop, execute and maintain test cases according to the business... 
    Senior

    Central Business Solutions

    Palo Alto, CA
    3 days ago
  • We are seeking a Sr. RTL Verification Engineer to join our client's team in Redwood City who is inspired to bring the power of generative AI to enhance and speed the design and manufacture of complex semiconductors. Employment Type: Full-Time Permanent Location: Redwood... 
    Senior
    Permanent employment
    Full time
    Work visa

    Talencore

    Redwood City, CA
    5 days ago
  • $160k - $200k

     ...infrastructure to build training and serving pipelines is also a key aspect to this role.  At Dexterity, our use cases require low latency inference that approaches real time and therefore you must be able to build highly performant models and serving architectures. While an... 
    Senior
    Work experience placement

    Dexterity

    Redwood City, CA
    12 days ago
  • $180k - $240k

    Sr. Computer Vision Engineer (Deep Learning) About Harbinger Harbinger is an American commercial electric vehicle (EV) company on a mission to...  ...vision applications. In‑depth understanding of training and inference pipelines, including data loading, augmentation, and loss... 
    Senior
    Local area

    Harbinger Motors

    Mountain View, CA
    3 days ago
  •  ...we want you on our team! Role Summary The Field Service Engineer (FSE) , along with the Field Application Scientists (FAS) are the...  ...Customer Training: Provide hands-on training to customer lab staff on proper instrument operation, routine care, and... 
    Senior
    Full time
    Contract work
    Remote work
    Home office
    Night shift

    Countable Labs

    Palo Alto, CA
    more than 2 months ago
  •  ...Vision-Language Model (VLM) & Visual Foundation Model (VFM) Field Engineer About Matroid Matroid helps enterprises build and...  ...can operate in cloud, edge, and hybrid environments, optimizing inference latency, GPU utilization, memory consumption, and throughput... 
    Senior
    Full time
    Work at office

    Matroid

    Palo Alto, CA
    19 days ago
  • $173k - $225k

     ...AI We’re a fast-moving team of aviators, engineers, and operators building an AI platform to...  ...together in aviation. You will ship LLM-powered product features end-to-end. That...  ...series or video. Familiarity with GPU inference, Triton, or TensorRT-LLM. Aviation or... 
    Senior
    Permanent employment
    Full time
    Local area
    Remote work
    3 days per week

    Beacon AI

    San Carlos, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Engineer (Mid, Sr, Staff). Be the first to apply!