Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior ML Systems Engineer, Inference

$150k - $220k
Full-time

Runpod

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: runpod. io/blog/one-million-developers . We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production. Responsibilities - Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable. - Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

- Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments. - Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically. - Work closely with product and infrastructure teams to shape how inference is offered on Runpod. - Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back. - Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough. Requirements - 5+ years of professional system engineering experience. - Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale. - Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases. - A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput. - Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving. - Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools. - The ability to explain your results clearly in writing and turn them into decisions. Preferred - Experience writing or tuning GPU kernels in CUDA or Triton. - Contributions to inference or ML systems projects. - Experience with multi-node GPU systems and high-speed networking. - Experience at a company where inference cost and latency were core business metrics. What You’ll Receive: - The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location - Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.

Vacancy posted 13 days ago
Similar jobs that could be interesting for youBased on the Senior ML Systems Engineer, Inference in Remote vacancy
  • $295k

     ...enterprises who are building AI systems. We believe that our work is...  ...Cohere is a team of researchers, engineers, designers, and more, who are...  ...Overview:We’re looking for a senior engineer to help build,...  ...working across the full stack of ML systems, this role gives you the... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    11 hours ago
  • $144.7k - $261.3k

     ...for autonomous vehicle development. We engineer high-performance tools that identify...  ...and partner with data-intensive ML teams to drive rapid innovation.Why Join...  ...scalability of next-generation autonomous systems.About the RoleAs a Senior Engineer in the Embodied AI Scaling... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Warren, MI
    1 day ago
  •  ...Job Responsibilities: Engineer, design, implement, and improve highly scalable machine learning systems and tools for enabling research Apply knowledge of relevant research...  ...experience ~0-2 years of Distributed ML Training (FSDP/DDP) experience ~5+ years of... 
    Senior
    Work experience placement

    SGS Consulting

    Remote
    more than 2 months ago
  •  ...Product to determine where new ML capabilities can meaningfully...  ...high quality agentic search systems, addressing tail latency,...  ...serving, query processing, and ML inference. Turn advances in information...  ...fragmented systems, mentor senior engineers, and align technical and product... 
    Senior
    Work at office
    Local area

    Atlassian

    San Francisco, CA
    3 days ago
  • $170.6k - $261.3k

     ...breakthrough hardware and battery systems to intuitive design,...  ...transportation on a global scale.As a Senior Machine Learning Engineer on the State Estimation and...  ...develop and improve the ML perception model that...  ...efficient training and inference pipelines, including model... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    4 days ago
  •  ...converse with all of their business systems through natural language to...  ...with Moveworks’ Reasoning Engine and natural language capabilities...  ...to help build cutting edge ML infrastructure for building and...  ...including distributed training and inference pipeline for large language... 
    Senior
    Permanent employment
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    2 days ago
  • $170.1k - $258.3k

     ...including Level 4-capable fully self-driving systems, to move us toward safer, more...  ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...sit at the heart of our on‑vehicle ML inference for ADAS and autonomous driving. We own... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Austin, TX
    1 day ago
  •  ...scientific data and machine learning systems. They are looking for a Senior Machine Learning Engineer with deep industry domain...  ...project, taking ownership of ML pipelines, model integration, and...  ...and build reliable training, inference, evaluation, and deployment pipelines... 
    Senior
    For contractors
    Remote work
    Flexible hours
    3 days per week

    Workana

    Oregon State
    1 day ago
  • $173k - $253k

    Matterport - Senior ML Ops Engineer Job Description CoStar Group is a leading global provider...  ...to analyze model performance, optimize inference speed and resource utilization, and...  ...environments.Familiarity with version control systems (e.g., Git) and agile development... 
    Senior
    Full time
    Work at office
    Work from home

    Matterport

    Sunnyvale, CA
    4 days ago
  • $145k - $165k

     ...offering tremendous career growth potential. Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full-time,...  ...design, build, and operate high-performance, highly reliable inference platforms for serving large machine learning models in production... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Austin, TX
    4 days ago
  • $250k - $350k

     ...to develop reliable AI systems for the world’s most important...  ..., paired with applied ML research, design, and...  ...Learning Research Engineer, you will operate across...  ...training/fine-tuning, inference, memory and retrieval,...  ...technical direction, mentor senior and staff-track... 
    Senior
    Full time

    Scale AI

    San Francisco, CA
    11 hours ago
  • $137.8k - $206.6k

     ...Department: HC-NA-IDA Analytics Engineering & BI ArchitectureRecruiter:...  ...to patients faster. As a Senior Data & MLOps Engineer, you own...  ...governance standards for data and ML solutions.Apply data privacy,...  ..., drift detection, batch inference, real-time inference, or automated... 
    Senior
    Work experience placement
    Work at office
    Local area
    Immediate start
    Work from home
    2 days per week
    3 days per week

    Merck

    Boston, MA
    3 hours ago
  • $141k - $249k

     ...Collaborate closely with autonomy and algorithm engineers to scale safe self-driving systems using an AI-first approach. Expand the model deployment...  ...truck. Create and benchmark new CUDA kernels for inference. Comprehensively profile model runtime and memory... 
    Senior
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    3 days ago
  • $180k - $240k

     ...recruiting for one of its clients a Senior Machine Learning Engineer - this is a fully remote...  ...join our small but mighty ML team building production-...  ..., and can ship LLM-powered systems that handle real, high-...  ...AI problems — low latency inference, hallucination reduction, prompt... 
    Senior
    Remote work
    Flexible hours

    Career Renew

    Chicago, IL
    15 days ago
  • $195k - $217k

     ...working closely with other engineering teams across the...  ...traditional machine learning (ML) and hybrid ML/...  ...engineering bar for ML systems company-wide. Key Responsibilities...  ...teams, guiding senior engineers and influencing...  ...model training and inference (including LLM serving)... 
    Remote work

    PointClickCare

    Eastern, KY
    2 days ago
  • $165k - $175k

     ...in the machine learning systems behind their...  ...you inherit a finished ML platform and simply maintain...  ...work directly with the engineers building the models that...  ...Description As a Senior MLOps Engineer, you will...  ...model training and online inference. Design processes for... 
    Senior
    Work at office
    3 days per week

    PhillyTech.Co

    King of Prussia, PA
    5 days ago
  •  ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From...  ...scale GPU orchestration to inference optimization, we own the...  ...approaches on real-world systems. The results will often lead...  ...currently looking for senior- and staff-level ML engineers... 
    Senior
    Full time
    Remote work

    Nebius

    Remote
    24 days ago
  • $160k - $220k

     ...style approaches Solid machine learning engineering foundations, including reproducible...  ...registries, GPU scheduling, or containerized inference Experience with transformer-based or...  ...-identifying information out of the systems you touch Technologies: AI Cloud... 
    Senior
    Full time
    Remote work

    Sonari Technologies LLC

    Cambridge, MA
    13 days ago
  • $262k - $364k

     ...user understanding, and budgeting predictions.Optimize ML inference and resolve system-level bottlenecks to improve serving efficiency for low-...  ...production.Preferred qualifications:Master’s degree or PhD in Engineering, Computer Science, or a related technical field.5 years... 
    Senior

    Google

    Mountain View, CA
    1 day ago
  •  ...usability and safety of automated driving systems. Our vision is to create autonomy that...  ...your career! The role As a Staff ML Performance Engineer, you’ll play a key role in high-impact projects, optimising ML inference for edge accelerators and GPUs. The focus... 
    Full time
    Work at office
    Work from home

    Wayve

    United Kingdom
    28 days ago
  •  ...the Role We are looking for a Senior Data Scientist - Experimentation & Causal Inference to serve as the statistical...  ...backend services — our platform engineering team handles that. Your job is...  ...and you will build the automated systems that answer that question before... 
    Senior
    Full time

    ByLabs

    Washington DC
    2 days ago
  •  ...T-W-Th) Who We AreWe are an engineering-focused IT organization responsible...  ...computing, data, and AI/ML integration solutions that...  ....The RoleWe are seeking a Senior Systems Engineer to lead the implementation...  ...(e.g., model endpoints, inference services, pipelines, RAG/... 
    Senior
    Full time
    H1b
    Local area
    Work from home
    Relocation package
    3 days per week

    General Motors

    Warren, MI
    3 days ago
  • $120.1k - $214.5k

     ...lives and help make the health system work better for everyone. From...  ...data scientists and software engineers through data extraction, research...  ...:Ship production ML systems end-to-end: problem framing...  ...architectures with low-latency inference and high availabilityBuild and... 
    Senior
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work

    UnitedHealth Group

    San Diego, CA
    1 day ago
  •  ...SUMMARY: We are seeking a Senior/Principal Machine Learning Engineer with deep expertise in...  ...deploying state-of-the-art AI systems for image, video, and...  ...scale AI models using modern ML infrastructure including...  ..., MLflow, distributed inference). Architect and oversee... 
    Senior
    Work at office
    Local area
    Remote work

    VALID8 Financial

    Eastern, KY
    3 days ago
  •  ...Labs is building intelligent systems that sit at the intersection...  ...chemistry, and large-scale engineering. Our goal is to translate complex...  .... The Role As a Senior AI/ML Engineer, you will lead the...  ...decisions around model serving, inference efficiency, and lifecycle... 
    Senior
    Remote work
    Flexible hours

    Absentia

    Eastern, KY
    3 days ago
  • $147.6k - $274k

     ...OpportunityWithin AI for Drug Discovery, the Software Engineering team builds and operates software...  ...engineering experience; and for the Senior Machine Learning Engineer level, you...  ..., automated testing, data modeling, and system design.You have experience independently... 
    Senior
    Full time
    Local area
    Worldwide
    Relocation package
    3 days per week

    Genentech

    New York, NY
    2 days ago
  • $119.8k - $234.7k

     ...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s...  ...that mission.Microsoft's Hardware Systems organization is developing AI-native silicon...  ...enable industry-leading AI training and inference.The Platform Systems Engineering (PSE) team... 
    Senior
    Ongoing contract
    Work at office
    Local area
    Flexible hours
    3 days per week

    Microsoft Corporation

    Mountain View, CA
    4 days ago
  • Job Title: Senior Machine Learning Engineer100% Remote...  ...owner of critical ML subsystems in production...  ...practical solutions, and ship systems that operate reliably...  ...training, evaluation, inference, and iteration.Turn...  ...research, product, and engineering to deliver real user impact... 
    Senior
    Full time
    Remote work

    Spectraforce Technologies

    Seattle, WA
    1 day ago
  • $144k - $192k

    Mission Summary:We are looking for a Machine Learning Systems Engineer to join our ML Acceleration team. In this role, you will be responsible for...  ...optimizing machine learning model execution during training and inference, alongside a strong understanding of fundamental machine... 
    Work at office
    Remote work

    Motional

    Boston, MA
    3 days ago
  • $209k - $313k

     ...Saturn, and other digital services.Snap Engineering teams build fun and technically sophisticated...  ...:Strong understanding of causal inference and modern approaches to estimating treatment...  ...(A/B tests) and leveraging causal ML in production systemsPreferred Qualifications... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Los Angeles, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior ML Systems Engineer, Inference. Be the first to apply!