Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote Senior ML Systems Engineer, Inference

$150k - $220k

grabjobs

Durham, NC
  • Remote job

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production. Responsibilities Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable. Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect. Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments. Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically. Work closely with product and infrastructure teams to shape how inference is offered on Runpod. Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back. Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough. Requirements 5+ years of professional system engineering experience. Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale. Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases. A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput. Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving. Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools. The ability to explain your results clearly in writing and turn them into decisions. Preferred Experience writing or tuning GPU kernels in CUDA or Triton. Contributions to inference or ML systems projects. Experience with multi-node GPU systems and high-speed networking. Experience at a company where inference cost and latency were core business metrics. What You’ll Receive: The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas. grabjobs

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Remote Senior ML Systems Engineer, Inference in Durham, NC vacancy
  • $150k - $220k

     ...platform has processed more than 20 billion inference requests. We closed a $100M Series A...  ...will depend on. We're a small, remote-first team. We take ownership seriously...  ...announcement: . We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best... 
    Remote job
    Senior
    Home office
    Visa sponsorship
    Work visa
    Flexible hours

    grabjobs

    Mango, FL
    2 days ago
  •  ...Amazon is seeking a Senior Software Development Engineer to design and implement inference data plane software for large models on custom hardware. You will own end-to-end features from kernel development to deployment, enable efficient data movement, and drive performance... 
    Senior

    Jobleads-US

    Kentucky
    1 day ago
  •  ...At Atlassian, the ML System Engineer will design and optimize large-scale model serving systems, spanning distributed infrastructure...  ...serving systems, benchmarking and tuning inference engines, and partnering with senior ML engineers to deploy open-source LLMs. #J-18808... 
    Remote job

    Jobleads-US

    Kentucky
    2 days ago
  •  ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research,... 
    Remote work
    Senior
    Contract work

    Jobleads-US

    Tempe, AZ
    2 days ago
  •  ...AfterQuery is building a research lab focused on ML inference and GPU kernel engineering, covering inference serving systems, kernel optimization, and deployment...  ...years of experience or equivalent. It’s fully remote, asynchronous, and offers competitive hourly compensation... 
    Remote job
    Senior
    Hourly pay

    Jobleads-US

    Laurel, MD
    1 day ago
  •  ...experimentation. We build a cohort of senior ML researchers to define correct and...  ...hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation...  ...production engineering. The role is remote, flexible, and asynchronous, with... 
    Remote job
    Senior
    Flexible hours

    Jobleads-US

    Ann Arbor, MI
    2 days ago
  •  ...datasets and experimentation. We are building a cohort of senior ML researchers focused on inference and GPU kernel engineering to define criteria for correctness and excellence on real-world problems. This fully remote, contract role invites experts to design scenarios,... 
    Remote work
    Senior
    Contract work
    Flexible hours

    Jobleads-US

    Irvine, CA
    1 day ago
  • $135 per hour

     ...datasets and experimentation. We believe great AI comes from exceptional, human-generated data. This remote, contract role focuses on ML inference and GPU kernel engineering. Competitive hourly pay ($135/hr) based on experience, with a broader range ($100–$170/hr) and a... 
    Remote job
    Senior
    Hourly pay
    Contract work
    Flexible hours

    Jobleads-US

    Rockville, MD
    1 day ago
  •  ...platform for the mobile app economy. As a Senior Staff Software Engineer on the Serving team, you will design,...  ...time, while optimizing GPU-powered inference pipelines. You will own changes...  ...capacity planning, collaborating with ML engineers to productionize #J-18808... 
    Remote job
    Senior

    Jobleads-US

    Kentucky
    2 days ago
  • $174.9k - $261.3k

    Remote/Hybrid Sunnyvale, California, United States of America Remote...  ...the world! The Data Labeling Engineering team designs, builds, and...  ...engineering , data engineering , and AI/ML , defining the strategies,...  ..., and direct impact on systems that unblock the next generation... 
    Remote work
    Senior
    Full time
    Local area
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Concord, NH
    5 days ago
  • $180k - $240k

     ...for one of its clients a Senior Machine Learning Engineer - this is a fully remote role for US/Canada...  ...join our small but mighty ML team building production...  ...can ship LLM-powered systems that handle real, high-...  ...problems — low latency inference, hallucination reduction... 
    Remote work
    Senior
    Flexible hours

    Career Renew

    Chicago, IL
    12 days ago
  •  ...NVIDIA Corporation in Austin, TX is seeking a Senior Software Engineer to build accelerated PyTorch-based solutions for large-scale ML models. You will work on GNNs, TFMs, and ensemble models, advancing training and inference on GPU infrastructure. You will... 
    Senior

    Jobleads-US

    Austin, TX
    2 days ago
  •  ...Qualcomm Incorporated in San Diego, CA seeks a Core ML Engineer to design, develop, and optimize machine learning systems powering next-gen AI platforms and applications...  ...You will build scalable ML pipelines, optimize inference, and deploy models across APIs and... 
    Senior

    Jobleads-US

    San Diego, CA
    2 days ago
  •  ...complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems...  ...in developing large distributed systems or high-load web services. ~ Open-... 
    Remote work
    Senior

    Nebius

    Remote
    29 days ago
  •  ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From...  ...scale GPU orchestration to inference optimization, we own the...  ...approaches on real-world systems. The results will often lead...  ...currently looking for senior- and staff-level ML engineers... 
    Remote work
    Senior
    Full time

    Nebius

    Remote
    21 days ago
  •  ...office, home, or hybrid arrangements. Atlassians can work from anywhere with a legal entity, and this ML System Engineer role sits on the AI & ML Platform's Inference team, designing scalable model-serving infra. The position requires 3+ years in software engineering... 
    Work at office
    Work from home
    Flexible hours

    Jobleads-US

    Seattle, WA
    1 day ago
  • $144.7k - $261.3k

     ...autonomous vehicle development. We engineer high-performance tools...  ...with data-intensive ML teams to drive rapid innovation...  ...-generation autonomous systems. About the Role As a Senior Engineer in the Embodied...  ...#GM-AV-1 This role is based remotely, but if the selected... 
    Remote work
    Senior
    Local area
    Work from home
    Flexible hours

    General Motors

    Austin, TX
    1 day ago
  •  ...machine learning and LLMs. You will work with data scientists and software engineers to build scalable production solutions, deploying and monitoring models for millions of medical charts daily. Remote work from anywhere in the U.S. is supported, with in-person... 
    Remote job
    Senior

    Jobleads-US

    San Diego, CA
    1 day ago
  •  ...NVIDIA Corporation is seeking a Senior Machine Learning Applications and Compiler Engineer to advance end-to-end inference optimization across NVIDIA platforms. You will work at the intersection of large-scale systems, compilers, and deep learning to map neural network... 
    Senior

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances. You will lead customization of open-source... 
    Senior

    Jobleads-US

    Seattle, WA
    2 days ago
  • $120 per hour

     ...Dorsey . Position: MLOps Engineer, LLM Systems (Serving, GPU Kernels,...  ...0–$120/hour Location: Remote Commitment: 40 hours/week...  ...profiling , debugging , and inference serving . Write accurate,...  ...AI model performance on ML systems and training infrastructure... 
    Remote job
    Contract work
    Summer work

    Mercor

    New York, NY
    20 days ago
  •  ...Samsara is hiring a Machine Learning Engineer for Safety AI who will design production ML APIs, build data flywheels, and optimize artifacts...  ...and ensure low latency, high-throughput inference across millions of devices. This remote role is open to candidates in the US or... 
    Remote job
    Senior

    Jobleads-US

    San Francisco, CA
    1 day ago
  • $195k - $217k

     ...closely with other engineering teams across the...  ...machine learning (ML) and hybrid ML/LLMsolutions...  ...bar for ML systems company-wide. Key...  ...teams, guiding senior engineers and influencing...  ...training and inference (including LLM serving...  .... #LI-AJ1 #LI-remote $195,000 - $217,0... 
    Remote job

    grabjobs

    Thousand Oaks, CA
    2 days ago
  • $145k - $165k

     ...offering tremendous career growth potential. Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary...  ...build, and operate high-performance, highly reliable inference platforms for serving large machine learning models... 
    Remote work
    Full time
    H1b
    Local area
    Immediate start
    Visa sponsorship

    Bright Vision Technologies

    Remote
    3 days ago
  • $160k - $220k

     ...approaches Solid machine learning engineering foundations, including...  ...scheduling, or containerized inference Experience with...  ...identifying information out of the systems you touch Technologies:...  ...engineers. The position is hybrid remote in Cambridge, MA 02142, and includes... 
    Remote work
    Senior
    Full time

    Sonari Technologies LLC

    Cambridge, MA
    10 days ago
  • $120.1k - $214.5k

     ...and help make the health system work better for...  ...scientists and software engineers through data extraction...  ...the flexibility to work remotely * from anywhere within...  ...Responsibilities:Ship production ML systems end-to-end:...  ...with low-latency inference and high availabilityBuild... 
    Remote work
    Senior
    Minimum wage
    Full time
    Work experience placement
    Work at office
    Local area

    UnitedHealth Group

    San Diego, CA
    3 days ago
  •  ...Office Location: Remote, US Based JOB SUMMARY...  ...We are seeking a Senior/Principal Machine Learning Engineer with deep expertise...  ...-of-the-art AI systems for image, video, and...  ...models using modern ML infrastructure including...  ..., distributed inference). Architect and... 
    Remote work
    Senior
    H1b
    Work at office
    Local area
    Work visa

    Jobleads-US

    Charleston, SC
    2 days ago
  •  ...building intelligent systems that sit at the intersection...  ..., and large-scale engineering. Our goal is to...  .... The Role As a Senior AI/ML Engineer, you will lead...  ...around model serving, inference efficiency, and lifecycle...  ...technical depth. Flexible remote or hybrid work... 
    Remote work
    Senior
    Flexible hours

    Absentia

    Eastern, KY
    11 hours ago
  •  ...AfterQuery is assembling a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems, kernel optimization, and deployment...  ...domain-specific problems. The role is fully remote, asynchronous, and contract-based, with emphasis... 
    Remote job
    Contract work
    Work at office

    Jobleads-US

    Boise, ID
    1 day ago
  •  ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production... 
    Remote job
    Hourly pay
    Contract work
    Flexible hours

    Jobleads-US

    Irving, TX
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote Senior ML Systems Engineer, Inference. Be the first to apply!