Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Systems Engineer, Inference

$150k - $220k

Runpod

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production. Responsibilities

Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.

Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.

Work closely with product and infrastructure teams to shape how inference is offered on Runpod.

Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.

Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.

Requirements

5+ years of professional system engineering experience.

Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.

Strong software engineering skills in

Python

. You're comfortable working in large, performance-critical codebases.

A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.

Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.

Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.

The ability to explain your results clearly in writing and turn them into decisions.

Preferred

Experience writing or tuning GPU kernels in CUDA or Triton.

Contributions to inference or ML systems projects.

Experience with multi-node GPU systems and high-speed networking.

Experience at a company where inference cost and latency were core business metrics.

What You’ll Receive: The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location

Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.

Generous medical, dental & vision plans

Flexible PTO- take the time you need to recharge

Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication

Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

$1,200 Home Office & Equipment Stipend-We set you up for success from day one with gear and support to create your ideal workspace

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Senior ML Systems Engineer, Inference in Remote vacancy
  • $150k - $220k

     ...one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at...  ...in our CEO's funding announcement: We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world... 
    Senior
    Remote work
    Home office
    Visa sponsorship
    Work visa
    Flexible hours

    Runpod

    Oregon State
    16 hours ago
  •  ...Job Responsibilities: Engineer, design, implement, and improve highly scalable machine learning systems and tools for enabling research Apply knowledge of relevant research...  ...experience ~0-2 years of Distributed ML Training (FSDP/DDP) experience ~5+ years of... 
    Senior
    Work experience placement

    SGS Consulting

    Remote
    more than 2 months ago
  • $174.9k - $261.3k

     ...understand the world! The Data Labeling Engineering team designs, builds, and operates...  ...engineering , data engineering , and AI/ML , defining the strategies, tooling, and...  ...leadership, and direct impact on systems that unblock the next generation of AV capabilities... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Atlanta, GA
    1 day ago
  • General Motors is seeking a Senior Engineer for the Embodied AI Scaling Foundations team to measure and visualize AV model performance at scale. You will design and implement evaluation and introspection tools used by GM AV teams, collaborating across Data, Infra, and validation... 
    Senior
    Remote job

    General Motors

    Seattle, WA
    2 days ago
  • General Motors, through Embodied AI, seeks a Senior Engineer to measure and visualize AV model performance. You will design and implement...  ...influence safety and scalability of next‑generation autonomous systems and to contribute to a cohesive evaluation flywheel. #J-1880... 
    Senior
    Remote job

    General Motors

    Austin, TX
    4 days ago
  • $144.7k - $261.3k

     ...autonomous vehicle development. We engineer high-performance tools that identify...  ...models and partner with data-intensive ML teams to drive rapid innovation....  ...scalability of next-generation autonomous systems. About The Role As a Senior Engineer in the Embodied AI Scaling... 
    Senior
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Seattle, WA
    2 days ago
  •  ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From...  ...scale GPU orchestration to inference optimization, we own the...  ...for frontier models. As a Senior Machine Learning Engineer...  ...NVIDIA Dynamo, or similar systems. - Build and productionize... 
    Senior
    Full time

    Nebius

    Remote
    5 days ago
  •  ...converse with all of their business systems through natural language to...  ...with Moveworks’ Reasoning Engine and natural language capabilities...  ...to help build cutting edge ML infrastructure for building and...  ...including distributed training and inference pipeline for large language... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours

    Servicenow

    Remote
    a month ago
  •  ...About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical...  ...model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.... 
    Full time
    Work at office
    3 days per week

    Pika

    Remote
    more than 2 months ago
  • $213k - $263k

     ...learning, learning from demonstration, generative modeling, Bayesian inference, hierarchical learning, and robust evaluation. This role...  ...have: Experience in building large-scale deep learning systems.  Experience in VLM/LLM models. We prefer: Experience... 
    Senior
    Full time
    Remote work

    Waymo

    Mountain View, CA
    more than 2 months ago
  •  ...scientific data and machine learning systems. They are looking for a Senior Machine Learning Engineer with deep industry domain...  ...project, taking ownership of ML pipelines, model integration, and...  ...and build reliable training, inference, evaluation, and deployment pipelines... 
    Senior
    Full time
    For contractors
    Remote work
    Flexible hours
    3 days per week

    Workana

    Indianapolis, IN
    1 day ago
  •  ...possible with today’s search systems and this role is intended to...  .... You'll lead the search and ML architecture that powers search...  ..., retrieval architectures, inference pipelines, and serving infrastructureDrive...  ...and elevate senior engineers across Search Serving; raise... 
    Senior
    Work at office
    Local area

    Atlassian

    Austin, TX
    13 days ago
  • $199.2k - $298.8k

     ...latest developments in AI and ML for autonomous driving. Independently...  ....  Mentors and guides engineers within the group.   Bachelor...  ...).  Development Tools & Eco-System (at scale) - Proficiency in...  ...Hyperpods, Anyscale, Etc. Model Inference Orchestration.   Perks of... 
    Senior
    Full time
    Immediate start
    Relocation

    Company

    Remote
    more than 2 months ago
  •  ...Product to determine where new ML capabilities can meaningfully...  ...high quality agentic search systems, addressing tail latency,...  ...serving, query processing, and ML inference. Turn advances in information...  ...fragmented systems, mentor senior engineers, and align technical and product... 
    Senior
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    13 days ago
  • $175k - $200k

     ...Senior Machine Learning Engineer Truveta is the world’s first health provider led data...  ...adaptive agentic systems and fine-tuned foundation models...  ...engineering. You’ll blend ML craftsmanship with platform...  ...efficiency, interpretability, and inference performance. Think and... 
    Senior
    Full time
    For contractors
    Visa sponsorship
    Work visa
    Flexible hours

    Truveta

    Remote
    26 days ago
  • $170.6k - $261.3k

     ...breakthrough hardware and battery systems to intuitive design,...  ...transportation on a global scale.As a Senior Machine Learning Engineer on the State Estimation and...  ...develop and improve the ML perception model that...  ...efficient training and inference pipelines, including model... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    a month ago
  • $220k - $247k

     ...marketing.   What You’ll Do As a  Senior Staff Machine Learning Engineer , you will operate at the company...  ...lead the design of large-scale ML systems and shared platforms that power all...  ...platforms (training, evaluation, inference, safety) used across multiple teams... 
    Senior
    Full time
    Work at office
    Immediate start
    Flexible hours
    3 days per week

    Typeface

    Remote
    more than 2 months ago
  •  ...spans cutting-edge LLMs / ML, large-scale data systems, financial reasoning, and...  ...joining a team of exceptional engineers, analysts, and investors...  ...media company.   As a Senior ML Engineer, working in a...  ...pipelines and model training to inference, evaluation, and... 
    Senior
    Full time
    Work at office
    Local area

    Versant Limited

    Remote
    a month ago
  • $180k - $240k

     ...recruiting for one of its clients a Senior Machine Learning Engineer - this is a fully remote...  ...join our small but mighty ML team building production-...  ..., and can ship LLM-powered systems that handle real, high-...  ...AI problems — low latency inference, hallucination reduction, prompt... 
    Senior
    Remote work
    Flexible hours

    Career Renew

    Chicago, IL
    6 days ago
  • $145k - $165k

     ...offering tremendous career growth potential. Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full-time...  ...design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in... 
    Full time
    H1b
    Local area
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Pflugerville, TX
    18 hours ago
  •  ...complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems...  ...in developing large distributed systems or high-load web services. ~ Open-... 
    Senior
    Remote work

    Nebius

    Remote
    23 days ago
  • $160k - $220k

     ...style approaches Solid machine learning engineering foundations, including reproducible...  ...registries, GPU scheduling, or containerized inference Experience with transformer-based or...  ...-identifying information out of the systems you touch Technologies: AI Cloud... 
    Senior
    Full time
    Remote work

    Sonari Technologies LLC

    Cambridge, MA
    4 days ago
  •  ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From...  ...scale GPU orchestration to inference optimization, we own the...  ...approaches on real-world systems. The results will often lead...  ...currently looking for senior- and staff-level ML engineers... 
    Senior
    Full time
    Remote work

    Nebius

    Remote
    15 days ago
  • $173k - $253k

    Matterport - Senior ML Ops Engineer Job Description CoStar Group is a leading global provider...  ...to analyze model performance, optimize inference speed and resource utilization, and...  ...environments.Familiarity with version control systems (e.g., Git) and agile development... 
    Senior
    Full time
    Work at office
    Work from home

    Matterport

    Sunnyvale, CA
    15 days ago
  • $300k - $400k

     ...About the Role You will own the systems layer that makes our frontier model training and inference fast, efficient, and tightly...  ...Profiling and benchmarking distributed ML systems to identify and...  ...world’s best — the scientists, engineers, and problem-solvers who don’t... 
    Visa sponsorship
    Flexible hours
    Shift work

    Periodic Labs

    Menlo Park, CA
    3 days ago
  • $250k - $350k

     ...to develop reliable AI systems for the world’s most important...  ..., paired with applied ML research, design, and...  ...Learning Research Engineer, you will operate across...  ...training/fine-tuning, inference, memory and retrieval,...  ...technical direction, mentor senior and staff-track... 
    Senior
    Full time

    Scale AI

    New York, NY
    a month ago
  • $170.1k - $258.3k

     ...including Level 4-capable fully self-driving systems, to move us toward safer, more...  ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...sit at the heart of our on-vehicle ML inference for ADAS and autonomous driving . We own... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    4 days ago
  •  ...simulation, and modern software engineering to accelerate the design,...  ...models can train on, and the systems that move those models from research...  ...waiting on it. Build the ML data pipelines for training,...  ..., online, and asynchronous inference, with safe rollout, rollback,... 
    Senior

    Xora Innovation

    San Diego, CA
    a month ago
  •  ...usability and safety of automated driving systems. Our vision is to create autonomy that...  ...your career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus... 
    Full time
    Work at office
    Work from home

    Wayve

    United Kingdom
    19 days ago
  •  ...any real time in the healthcare system, you've felt it, and it only...  ...is the first Machine Learning Engineer role at Metriport. We have access...  ...honest. This is applied ML on messy, high-dimensional, real...  ...infrastructure: training and inference pipelines, experiment tracking... 
    Senior
    Work at office
    Work from home
    Relocation
    Flexible hours

    Metriport

    San Francisco, CA
    22 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Systems Engineer, Inference. Be the first to apply!