Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote Senior ML Systems Engineer, Inference

$150k - $220k

Runpod

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

 

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

 

Learn more in our CEO's funding announcement: .

We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production.

Responsibilities
  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.

  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.

  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

  • Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.

  • Work closely with product and infrastructure teams to shape how inference is offered on Runpod.

  • Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.

  • Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.

Requirements
  • 5+ years of professional system engineering experience.

  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.

  • Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases.

  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.

  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.

  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.

  • The ability to explain your results clearly in writing and turn them into decisions.

Preferred
  • Experience writing or tuning GPU kernels in CUDA or Triton.

  • Contributions to inference or ML systems projects.

  • Experience with multi-node GPU systems and high-speed networking.

  • Experience at a company where inference cost and latency were core business metrics.

What You'll Receive:

  • The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate's experience, qualifications, and location

  • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.

  • Generous medical, dental & vision plans

  • Flexible PTO- take the time you need to recharge

  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication

  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.

  • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Remote Senior ML Systems Engineer, Inference in Oregon, OH vacancy
  • $150k - $220k

     ...platform has processed more than 20 billion inference requests. We closed a $100M Series A...  ...will depend on. We're a small, remote-first team. We take ownership seriously...  ...million-developers. We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best... 
    Remote work
    Senior
    Full time

    Runpod

    Remote
    7 days ago
  •  ...At Atlassian, the ML System Engineer will design and optimize large-scale model serving systems, spanning distributed infrastructure...  ...serving systems, benchmarking and tuning inference engines, and partnering with senior ML engineers to deploy open-source LLMs. #J-18808... 
    Remote job

    Jobleads-US

    Kentucky
    19 hours ago
  •  ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research,... 
    Remote work
    Senior
    Contract work

    Jobleads-US

    Tempe, AZ
    19 hours ago
  •  ...experimentation. We build a cohort of senior ML researchers to define correct and...  ...hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation...  ...production engineering. The role is remote, flexible, and asynchronous, with... 
    Remote job
    Senior
    Flexible hours

    Jobleads-US

    Ann Arbor, MI
    18 hours ago
  •  ...platform for the mobile app economy. As a Senior Staff Software Engineer on the Serving team, you will design,...  ...time, while optimizing GPU-powered inference pipelines. You will own changes...  ...capacity planning, collaborating with ML engineers to productionize #J-18808... 
    Remote job
    Senior

    Jobleads-US

    Kentucky
    18 hours ago
  •  ...from exceptional, human-generated data. We're backed by top investors, including Y Combinator and Box Group, and support all leading AI labs. This role offers fully remote, flexible, asynchronous work with competitive hourly compensation. #J-18808-Ljbffr Jobleads-US
    Remote work
    Senior
    Hourly pay
    Flexible hours

    Jobleads-US

    Hartford, CT
    18 hours ago
  • $174.9k - $261.3k

    Remote/Hybrid Sunnyvale, California, United States of America Remote...  ...the world! The Data Labeling Engineering team designs, builds, and...  ...engineering , data engineering , and AI/ML , defining the strategies,...  ..., and direct impact on systems that unblock the next generation... 
    Remote work
    Senior
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Lincoln, NE
    4 days ago
  • $195.2k - $262.2k

     ...large in-house AI/ML infrastructure. Built by engineers, for engineers. From...  ...GPU orchestration to inference optimization, we...  ...frontier models. As a Senior Machine Learning...  ...Dynamo, or similar systems. Build and productionize...  ...caregivers. Remote work reimbursement:... 
    Remote work
    Senior
    Full time
    Temporary work
    Immediate start

    Nebius

    Palo Alto, CA
    3 days ago
  • $180k - $240k

     ...for one of its clients a Senior Machine Learning Engineer - this is a fully remote role for US/Canada...  ...join our small but mighty ML team building production...  ...can ship LLM-powered systems that handle real, high-...  ...problems — low latency inference, hallucination reduction... 
    Remote work
    Senior
    Flexible hours

    Career Renew

    Chicago, IL
    11 days ago
  •  ...complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems...  ...in developing large distributed systems or high-load web services. ~ Open-... 
    Remote work
    Senior

    Nebius

    Remote
    28 days ago
  •  ...Qualcomm Incorporated in San Diego, CA seeks a Core ML Engineer to design, develop, and optimize machine learning systems powering next-gen AI platforms and applications...  ...You will build scalable ML pipelines, optimize inference, and deploy models across APIs and... 
    Senior

    Jobleads-US

    San Diego, CA
    18 hours ago
  •  ...NVIDIA Corporation in Austin, TX is seeking a Senior Software Engineer to build accelerated PyTorch-based solutions for large-scale ML models. You will work on GNNs, TFMs, and ensemble models, advancing training and inference on GPU infrastructure. You will... 
    Senior

    Jobleads-US

    Austin, TX
    19 hours ago
  •  ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From...  ...scale GPU orchestration to inference optimization, we own the...  ...approaches on real-world systems. The results will often lead...  ...currently looking for senior- and staff-level ML engineers... 
    Remote work
    Senior
    Full time

    Nebius

    Remote
    20 days ago
  • $195k - $217k

     ...closely with other engineering teams across the...  ...machine learning (ML) and hybrid ML/LLMsolutions...  ...bar for ML systems company-wide....  ...engineering teams, guiding senior engineers and...  ...model training and inference (including LLM...  ...cost. #LI-AJ1 #LI-remote $195,000 - $217,... 
    Remote work

    Pointclickcare

    Glassboro, NJ
    1 day ago
  •  ...Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances. You will lead customization of open-source... 
    Senior

    Jobleads-US

    Seattle, WA
    19 hours ago
  • $144.7k - $261.3k

     ...autonomous vehicle development. We engineer high-performance tools...  ...with data-intensive ML teams to drive rapid innovation...  ...-generation autonomous systems. About the Role As a Senior Engineer in the Embodied...  ...#GM-AV-1 This role is based remotely, but if the selected... 
    Remote work
    Senior
    Local area
    Work from home
    Flexible hours

    General Motors

    Austin, TX
    5 days ago
  • $120 per hour

     ...Dorsey . Position: MLOps Engineer, LLM Systems (Serving, GPU Kernels,...  ...0–$120/hour Location: Remote Commitment: 40 hours/week...  ...profiling , debugging , and inference serving . Write accurate,...  ...AI model performance on ML systems and training infrastructure... 
    Remote job
    Contract work
    Summer work

    Mercor

    New York, NY
    19 days ago
  •  ...all of their business systems through natural language...  ...with Moveworks’ Reasoning Engine and natural language...  ...help build cutting edge ML infrastructure for...  ...distributed training and inference pipeline for large language...  ...personas (flexible, remote, or required in office)... 
    Remote work
    Senior
    Permanent employment
    Full time
    Work at office
    Flexible hours

    ServiceNow

    Mountain View, CA
    a month ago
  • $145k - $165k

     ...offering tremendous career growth potential. Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary...  ...build, and operate high-performance, highly reliable inference platforms for serving large machine learning models... 
    Remote work
    Full time
    H1b
    Local area
    Immediate start
    Visa sponsorship

    Bright Vision Technologies

    Remote
    2 days ago
  • $120.1k - $214.5k

     ...and help make the health system work better for...  ...scientists and software engineers through data extraction...  ...the flexibility to work remotely * from anywhere within...  ...Responsibilities:Ship production ML systems end-to-end:...  ...with low-latency inference and high availabilityBuild... 
    Remote work
    Senior
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area

    UnitedHealth Group

    San Diego, CA
    2 days ago
  •  ...Office Location: Remote, US Based JOB SUMMARY...  ...We are seeking a Senior/Principal Machine Learning Engineer with deep expertise...  ...-of-the-art AI systems for image, video, and...  ...models using modern ML infrastructure including...  ..., distributed inference). Architect and... 
    Remote work
    Senior
    H1b
    Work at office
    Local area
    Work visa

    Jobleads-US

    Charleston, SC
    18 hours ago
  •  ...building intelligent systems that sit at the intersection...  ..., and large-scale engineering. Our goal is to...  .... The Role As a Senior AI/ML Engineer, you will lead...  ...around model serving, inference efficiency, and lifecycle...  ...technical depth. Flexible remote or hybrid work... 
    Remote work
    Senior
    Flexible hours

    Absentia

    Eastern, KY
    4 days ago
  • $144.7k - $261.3k

     ...infrastructure, and ML/AI GPU platforms for...  ...GM is looking for a Senior Performance Engineer to join the AV Capacity...  ...support GM’s long-term GPU system strategy and “...  ...scale ML training and inference environments. Your...  ...discretion. Ability to sit remote in Seattle, WA until... 
    Remote work
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours
    3 days per week

    General Motors LLC

    Sunnyvale, CA
    4 hours ago
  •  ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Jobleads-US

    Irving, TX
    18 hours ago
  • $87.12k - $151.25k

     ...apply now.We are currently seeking a Senior AI/ML Engineer - Hybrid to join our team in Irving, Texas...  ...training, evaluation, and real‑time inference.Automate model deployment, monitoring,...  .... The starting pay range for this remote role is $87,120-$151,250. This range reflects... 
    Remote work
    Senior
    Temporary work
    Work at office
    Flexible hours

    NTT DATA

    Irving, TX
    5 days ago
  •  ...Search, Q&A, and Conversational AI system that integrates seamlessly with...  ...Atlassian ecosystem.About the AI & ML Platform TeamOur team’s goal is...  ...Responsibilities About This RoleAs a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design and... 
    Senior
    Work at office
    Local area

    Atlassian

    Austin, TX
    4 hours ago
  • $81.2k - $101.5k

     ...to Bring the Future! Job Purpose The Battery Management System Engineer role is responsible for delivering defined engineering assignments...  ...skills Workstyle This is an onsite job. One remote workday per week may be possible with prior departmental approval... 
    Remote work
    Senior
    Full time
    Temporary work
    Work experience placement
    Relocation package

    Honda Dev. and Mfg. of Am.,LLC

    Raymond, OH
    27 days ago
  • $70 - $80 per hour

    A dynamic media agency in the United States is seeking a Senior Machine Learning Engineer/Data Scientist to develop and deliver machine-learning solutions...  ...Python, and experience with Bayesian modeling. This is a remote role requiring collaboration across various teams and... 
    Remote work
    Senior
    Full time
    Contract work

    AUSTIN WORKS

    United States
    5 days ago
  •  ...compatibility with existing systems and enterprise...  ...server, cloud, and platform engineering teams. Operationalize...  ...experience supporting AI/ML platforms, MLOps...  ...Experience with model serving, inference optimization, or AI...  ...equivalent experience REMOTE WORK NOTICE: This position... 
    Remote work
    Senior
    Work at office

    ARA

    Raleigh, NC
    22 hours ago
  •  ...Search Serving, the full-time Senior Principal Machine Learning Engineer will design and build high-quality search systems, set the long-term...  ...for online retrieval and ML inference, and establish strong production...  ..., all while working remotely or in various office locations... 
    Remote work
    Senior
    Full time
    Work at office

    Virtual Vocations Inc

    United States
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote Senior ML Systems Engineer, Inference. Be the first to apply!