LLM Inference Engineer (Mid, Sr, Staff)
Hippocratic AI Inc.
Role Mission
As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses-makings difference between conversational experiences that feel natural and those that feel broken. This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems.
What You Will Accomplish
Own your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduce latency, improve throughput, or optimize cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment.
Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel optimization techniques that become part of our core infrastructure, and made our serving stack a durable competitive advantage in healthcare AI deployment.
The Team
You’ll work alongside systems engineers, ML researchers, and infrastructure experts who are obsessed with making AI systems fast, reliable, and cost-effective. This is a team that values deep technical rigor, continuous benchmarking, and solving hard systems problems that have real impact on patient experience and business unit economics.
What You’ll Do
Design and implement multi-node serving architectures for distributed LLM inference
Optimize multi-LoRA serving systems
Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality
Implement speculative decoding and other latency optimization strategies
Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases
Continuously benchmark and improve system performance across various deployment scenarios and GPU types
Location Requirement
We believe the best ideas happen together. This role is based in our Menlo Park, California office, expected to be in office five days a week.
Compensation
Compensation is based on experience, expertise, and level of responsibility. We offer competitive packages that reflect the seniority and scope of the role, along with equity, health insurance, and other benefits.
What You Bring
Must-Have:
Experience optimizing LLM inference systems at scale
Proven expertise with distributed serving architectures for large language models
Hands-on experience implementing quantization techniques for transformer models
Strong understanding of modern inference optimization methods, including:
Speculative decoding techniques with draft models
Eagle speculative decoding approaches
Proficiency in Python and C++
Experience with CUDA programming and GPU optimization
Nice-to-Have:
Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM
Experience with custom CUDA kernels
Track record of deploying inference systems in production environments
Deep understanding of performance optimization systems
Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters – we want to hear your story.
Open source contributor? Even better! If you've contributed to projects like vllm, sglang, lmdeploy or similar LLM optimization frameworks, we'd love to see your PRs. Your contributions to these communities demonstrate exactly the kind of collaborative innovation we value.
Join a team where your expertise won't just be appreciated—it will be celebrated and amplified. Help us shape the future of AI deployment at scale!
References
1. Polaris: A Safety-focused LLM Constellation Architecture for Healthcare,
2. Polaris 2:
3. Personalized Interactions:
4. Human Touch in AI:
5. Empathetic Intelligence:
Why Join Hippocratic AI
Reinvent healthcare with AI that puts safety first. We're building the world's first healthcare-only, safety-focused LLM – a breakthrough platform designed to transform patient outcomes at a global scale. This is category creation.
Work with the people shaping the future. Hippocratic AI was co-founded by CEO Munjal Shah and a team of physicians, hospital leaders, AI pioneers, and researchers from institutions like El Camino Health, Johns Hopkins, Washington University in St. Louis, Stanford, Google, Meta, Microsoft, and NVIDIA.
Backed by the world’s leading healthcare and AI investors. We recently raised a $126M Series C at a $3.5B valuation, led by Avenir Growth, bringing total funding to $404M with participation from CapitalG, General Catalyst, a16z, Kleiner Perkins, Premji Invest, UHS, Cincinnati Children's, WellSpan Health, John Doerr, Rick Klausner, and others.
Build alongside the best in healthcare and AI. Join experts who’ve spent their careers improving care, advancing science, and building world‑changing technologies – ensuring our platform is powerful, trusted, and truly transformative.
Equal Opportunity
Hippocratic AI is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, national origin, sex, age, disability, sexual orientation, gender identity or expression, genetic information, military or veteran status, or any other characteristic protected by applicable law. We are committed to building a team that reflects the patients we serve. We actively encourage applications from candidates of all backgrounds. If you require accommodations during the hiring process, please contact View email address on click.appcast.io.
#J-18808-Ljbffr- ...companies. As a Senior Principal Software Engineer at JPMorganChase within the Commercial &... ...deployment and optimization using model inference servers such as Triton Inference Server... ...Demonstrated success architecting and deploying LLM & GNN solutions on AWS (e.g., SageMaker,...Senior
- ...architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...-speed inference.About The RoleWe’re hiring a Senior Frontend Engineer to own and scale critical parts of the Cerebras Developer...SeniorShift work
$195.2k - $262.2k
...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across... ...regressions during production rollouts. Optimize LLM and VLM endpoints for latency, throughput,...SeniorFull timeTemporary workImmediate startRemote work$195.2k - $262.2k
...large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across... ...Define and execute research programs in efficient LLM and VLM inference with measurable production...SeniorFull timeTemporary workImmediate startRemote work- ...Sales Engineer The Sales Engineer is the Technical Lead for the Proofpoint Pre-Sales process... ...closely with customer/prospect security staff for technical discovery, as well as... ...Pacific Northwest. This role will focus on Mid-Market accounts in the Bay Area and Pacific...SeniorLocal areaFlexible hoursNight shift
$272k - $336k
...autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Systems Engineering team works together to blend software and hardware systems in groundbreaking new ways. We set the high performance standards that...SeniorOdd jobFull timeRemote work$86.25k - $145k
We are hiring a highly technical, business-savvy Senior GTM Engineer to help build the next generation of Navan’s GTM systems. This hands... ....At least 2 years of hands-on experience building and deploying LLM-powered agents, copilots, or automations used by business teams...Senior$45.67 - $55.29 per hour
...or 40/week ~ Performance bonuses paid mid-year and year-end ~401(k) profit-sharing program ZFA Structural Engineers is a mid-sized structural engineering firm... ...sustainably over the last decade and now employ a staff of 100 people. We work hard to provide a...SeniorFull timeWork at office- ...Senior Forward Deployed Engineer (FDE)Location: Hybrid, Palo Alto, CAAs a Senior Forward Deployed Engineer (Sr. FDE) at Uniphore, you will take technical ownership of strategic... ...auditable AI agents.Stay current on RAG, LLM/SLM advancements, Agentic AI tooling, vector...Senior
$175k - $240k
...Business AI technology company, is seeking an experienced Senior Sales Engineer to join its high-performance Sales Engineering team to support... ...presentations, demos showcasing GenAI, Agentic AI, RAG, and LLM-based solutions. Build and deliver custom demos and POCs to...SeniorFull time- ...San Carlos, California. Relocation benefits available. Summary Delta Star Inc. is seeking a skilled and innovative Sr. Electrical Design Engineer to design and engineer medium voltage, mobile power transformers and mobile substations. If you're passionate about...SeniorRelocationRelocation package
- ...for everyone. ABOUT THE ROLE As a Senior Forward Deployed Engineer, you’ll deliver Retell’s voice AI solutions to enterprise clients... ...Become go-to experts on building voice AI solutions, especially LLM prompting tips & tricks. Build full-stack integrations (...SeniorH1bWork at office
- ...great investors. If you want to disrupt the landscape of healthcare, join our team. We're looking for a Senior Quality Engineer to own quality on a mid-field MRI system as we drive toward a 510(k) submission. This is a foundational hire: you'll stand up and run the...SeniorRelocation package
$119.8k - $234.7k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft... ...-leading AI training and inference. The Platform Systems Engineering (PSE) team is seeking a Sr. AI Accelerator Tools Development... ...including: HPL/HPC benchmarks LLM training workloads Transformer...SeniorOngoing contractWork at officeLocal areaWorldwide3 days per week- ...Responsibilities About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design and optimize large-scale model... ...GPU inference engines (vLLM, SGLang, Triton, TensorRT-LLM, etc.).Strong background in system optimizations: batching...SeniorWork at officeLocal area
$132.1k - $165.1k
...with manufacturing, validation, service, suppliers, and program teams to investigate field and production issues, develop robust engineering changes, and verify solutions through analysis and test. The ideal candidate brings strong engineering judgment, hands‑on problem...SeniorFull timeContract workTemporary workPart timeLocal areaShift work- ...is delivered for millions of patients worldwide.Were a team of engineers, clinicians, and innovators united by one purpose: to make surgery... ...compensation ranges are listed.SummaryType: Full-timeFunction: EngineeringExperience level: Mid-Senior LevelIndustry: Medical DeviceSeniorLocal areaWorldwideFlexible hours
- ...and implementation of technical safety requirements at the hardware level. Work with cross functional teams and hardware design engineers to implement hardware safety requirements. Work with various functional safety analysis methods (FTA, FMEDA) supporting...SeniorFull timeContract work
$120k - $150k
...ll Do Join DES as a Senior Structural Engineer and play an important role in delivering... ...technical guidance to engineers and developing staff Contribute to the professional... ...supporting the development of junior and mid-level engineers OSHPD and DSA experience...SeniorFull timeTemporary workFor contractorsWork at officeLocal areaFlexible hours$210k
...Sr. Reliability Engineer Location: Mountain View/San Jose Bay Area Team: System Engineering The Impact You’ll Make: We are seeking... ...his interest in Avantus’ development business to KKR in mid-2024, it had delivered over $1 billion in profit and secured...SeniorWork at officeLocal areaNight shift- ...rigor to a rapidly evolving AI inference space. Our mission is to make... ..., and speak fluently to both engineering and procurement — someone who... ...— credible with a skeptical staff engineer, clear with a CFO,... ...internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math,...
$156k - $183k
...companies, and the planet. Role Description As a Senior Value Engineer specializing in the Public Sector, you are pushing the envelope... .... This includes architecting and delivering secure, scalable LLM/agent systems with RAG, tools, and guardrails, while seamlessly...SeniorFull timeWork experience placementWork at officeLocal areaRemote workWorldwideFlexible hoursShift work- ...Salesforce, merchants' websites, and other platforms.Job RequirementsWe are looking for curious, energetic rock stars to join our engineering team.You will use scenario-driven testing methodology to develop, execute and maintain test cases according to the business...Senior
$260k
...Mid/Senior-Level Transactional Tax Associate Baker McKenzie’s Tax Practice is a leader within our Firm and one of the most highly regarded... ...structuring. Top academic credentials are a prerequisite; an LLM in taxation is not required, but may be favorably considered. Requirements...SeniorHourly payFull timeWork at officeLocal areaWorldwide$160k - $200k
...infrastructure to build training and serving pipelines is also a key aspect to this role. At Dexterity, our use cases require low latency inference that approaches real time and therefore you must be able to build highly performant models and serving architectures. While an...SeniorWork experience placement- ...we want you on our team! Role Summary The Field Service Engineer (FSE) , along with the Field Application Scientists (FAS) are the... ...Customer Training: Provide hands-on training to customer lab staff on proper instrument operation, routine care, and...SeniorFull timeContract workRemote workHome officeNight shift
$180k - $240k
Sr. Computer Vision Engineer (Deep Learning) About Harbinger Harbinger is an American commercial electric vehicle (EV) company on a mission to... ...vision applications. In‑depth understanding of training and inference pipelines, including data loading, augmentation, and loss...SeniorLocal area- ...Vision-Language Model (VLM) & Visual Foundation Model (VFM) Field Engineer About Matroid Matroid helps enterprises build and... ...can operate in cloud, edge, and hybrid environments, optimizing inference latency, GPU utilization, memory consumption, and throughput...SeniorFull timeWork at office
- ...Work closely with the Electrical Hardware Design team to create engineering and delivery schedules, closely track engineering progress, and... ...issues ~ Strong communication experience with executive staff members ~ Proven experience tracking and driving the progress...Full timeContract workLocal area
$180k - $240k
...Foundation. We are seeking a highly skilled Senior Deep Learning Engineer to drive the development and deployment of advanced perception... ...applications. ~ In-depth understanding of training and inference pipelines, including data loading, augmentation, and loss...SeniorFull timeContract workLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Engineer (Mid, Sr, Staff). Be the first to apply!
- technology administrator Menlo Park, CA
- assistant engineer Menlo Park, CA
- staff engineer Menlo Park, CA
- senior staff systems engineer Menlo Park, CA
- engineering aide Menlo Park, CA
- senior software engineer ruby on rails Menlo Park, CA
- senior compensation manager Menlo Park, CA
- senior manager Menlo Park, CA
- senior living Menlo Park, CA
- senior vmware engineer Menlo Park, CA



