Remote Senior ML Systems Engineer, Inference
$150k - $220kRunpod
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.
We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.
Learn more in our CEO's funding announcement: .
We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production.
Responsibilities
Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.
Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.
Work closely with product and infrastructure teams to shape how inference is offered on Runpod.
Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back.
Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.
Requirements
5+ years of professional system engineering experience.
Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.
Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases.
A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.
Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.
Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
The ability to explain your results clearly in writing and turn them into decisions.
Preferred
Experience writing or tuning GPU kernels in CUDA or Triton.
Contributions to inference or ML systems projects.
Experience with multi-node GPU systems and high-speed networking.
Experience at a company where inference cost and latency were core business metrics.
What You’ll Receive:
The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location
Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
Generous medical, dental & vision plans
Flexible PTO- take the time you need to recharge
Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.
$1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace
Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.
$150k - $220k
...platform has processed more than 20 billion inference requests. We closed a $100M Series A... ...will depend on. We're a small, remote-first team. We take ownership seriously... ...million-developers. We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best...Remote workSeniorFull time- ...At Atlassian, the ML System Engineer will design and optimize large-scale model serving systems, spanning distributed infrastructure... ...serving systems, benchmarking and tuning inference engines, and partnering with senior ML engineers to deploy open-source LLMs. #J-18808...Remote job
- ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research,...Remote workSeniorContract work
- ...experimentation. We build a cohort of senior ML researchers to define correct and... ...hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation... ...production engineering. The role is remote, flexible, and asynchronous, with...Remote jobSeniorFlexible hours
- ...platform for the mobile app economy. As a Senior Staff Software Engineer on the Serving team, you will design,... ...time, while optimizing GPU-powered inference pipelines. You will own changes... ...capacity planning, collaborating with ML engineers to productionize #J-18808...Remote jobSenior
$174.9k - $261.3k
...Remote/Hybrid Sunnyvale, California, United States of America... ...the world! The Data Labeling Engineering team designs, builds, and operates... ...engineering , and AI/ML , defining the strategies, tooling... ..., and direct impact on systems that unblock the next generation...Remote workSeniorFull timeLocal areaWork from homeRelocation packageFlexible hours- ...from exceptional, human-generated data. We're backed by top investors, including Y Combinator and Box Group, and support all leading AI labs. This role offers fully remote, flexible, asynchronous work with competitive hourly compensation. #J-18808-Ljbffr Jobleads-USRemote workSeniorHourly payFlexible hours
$195.2k - $262.2k
...large in-house AI/ML infrastructure. Built by engineers, for engineers. From... ...GPU orchestration to inference optimization, we... ...frontier models. As a Senior Machine Learning... ...Dynamo, or similar systems. Build and productionize... ...caregivers. Remote work reimbursement:...Remote workSeniorFull timeTemporary workImmediate start$180k - $240k
...for one of its clients a Senior Machine Learning Engineer - this is a fully remote role for US/Canada... ...join our small but mighty ML team building production... ...can ship LLM-powered systems that handle real, high-... ...problems — low latency inference, hallucination reduction...Remote workSeniorFlexible hours- ...NVIDIA Corporation in Austin, TX is seeking a Senior Software Engineer to build accelerated PyTorch-based solutions for large-scale ML models. You will work on GNNs, TFMs, and ensemble models, advancing training and inference on GPU infrastructure. You will...Senior
- ...Qualcomm Incorporated in San Diego, CA seeks a Core ML Engineer to design, develop, and optimize machine learning systems powering next-gen AI platforms and applications... ...You will build scalable ML pipelines, optimize inference, and deploy models across APIs and...Senior
- ...complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems... ...in developing large distributed systems or high-load web services. ~ Open-...Remote workSenior
- ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From... ...scale GPU orchestration to inference optimization, we own the... ...approaches on real-world systems. The results will often lead... ...currently looking for senior- and staff-level ML engineers...Remote workSeniorFull time
$144.7k - $261.3k
...autonomous vehicle development. We engineer high-performance tools... ...with data-intensive ML teams to drive rapid innovation... ...-generation autonomous systems. About the Role As a Senior Engineer in the Embodied... ...#GM-AV-1 This role is based remotely, but if the selected...Remote workSeniorLocal areaWork from homeFlexible hours- ...Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances. You will lead customization of open-source...Senior
$120 per hour
...Dorsey . Position: MLOps Engineer, LLM Systems (Serving, GPU Kernels,... ...0–$120/hour Location: Remote Commitment: 40 hours/week... ...profiling , debugging , and inference serving . Write accurate,... ...AI model performance on ML systems and training infrastructure...Remote jobContract workSummer work- ...all of their business systems through natural language... ...with Moveworks’ Reasoning Engine and natural language... ...help build cutting edge ML infrastructure for... ...distributed training and inference pipeline for large language... ...personas (flexible, remote, or required in office)...Remote workSeniorPermanent employmentFull timeWork at officeFlexible hours
$145k - $165k
...offering tremendous career growth potential. Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary... ...build, and operate high-performance, highly reliable inference platforms for serving large machine learning models...Remote workFull timeH1bLocal areaImmediate startVisa sponsorship- ...Office Location: Remote, US Based JOB SUMMARY... ...We are seeking a Senior/Principal Machine Learning Engineer with deep expertise... ...-of-the-art AI systems for image, video, and... ...models using modern ML infrastructure including... ..., distributed inference). Architect and...Remote workSeniorH1bWork at officeLocal areaWork visa
$120.1k - $214.5k
...and help make the health system work better for... ...scientists and software engineers through data extraction... ...the flexibility to work remotely * from anywhere within... ...Responsibilities:Ship production ML systems end-to-end:... ...with low-latency inference and high availabilityBuild...Remote workSeniorMinimum wageFull timeWork experience placementWork at officeLocal area- ...building intelligent systems that sit at the intersection... ..., and large-scale engineering. Our goal is to... .... The Role As a Senior AI/ML Engineer, you will lead... ...around model serving, inference efficiency, and lifecycle... ...technical depth. Flexible remote or hybrid work...Remote workSeniorFlexible hours
- ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production...Remote jobHourly payContract workFlexible hours
$87.12k - $151.25k
...apply now.We are currently seeking a Senior AI/ML Engineer - Hybrid to join our team in Irving, Texas... ...training, evaluation, and real‑time inference.Automate model deployment, monitoring,... .... The starting pay range for this remote role is $87,120-$151,250. This range reflects...Remote workSeniorTemporary workWork at officeFlexible hours$81.2k - $101.5k
...to Bring the Future! Job Purpose The Battery Management System Engineer role is responsible for delivering defined engineering assignments... ...skills Workstyle This is an onsite job. One remote workday per week may be possible with prior departmental approval...Remote workSeniorFull timeTemporary workWork experience placementRelocation package- ...Amazon is hiring a Senior Software Development Engineer for AI/ML workloads on AWS Neuron to accelerate inference on Inferentia and Trainium accelerators. You will design, implement... ...with cross-functional teams, analyze system performance, and contribute to architecture...Senior
- ...compatibility with existing systems and enterprise... ...server, cloud, and platform engineering teams. Operationalize... ...experience supporting AI/ML platforms, MLOps... ...Experience with model serving, inference optimization, or AI... ...equivalent experience REMOTE WORK NOTICE: This position...Remote workSeniorWork at office
$70 - $80 per hour
A dynamic media agency in the United States is seeking a Senior Machine Learning Engineer/Data Scientist to develop and deliver machine-learning solutions... ...Python, and experience with Bayesian modeling. This is a remote role requiring collaboration across various teams and...Remote workSeniorFull timeContract work- ...Search Serving, the full-time Senior Principal Machine Learning Engineer will design and build high-quality search systems, set the long-term... ...for online retrieval and ML inference, and establish strong production... ..., all while working remotely or in various office locations...Remote workSeniorFull timeWork at office
- ...What You’ll Do Data Engineering & Pipeline Development... ...that enable downstream ML, analytics, and operational... ...manage training and inference workflows on GCP using... ...operational guardrails for ML systems Evaluate emerging... ...up to three flexible remote days per month....Remote workSeniorWork at officeVisa sponsorshipFlexible hoursShift work
- ...Job Responsibilities: Engineer, design, implement, and improve highly scalable machine learning systems and tools for enabling research Apply knowledge of relevant research... ...experience ~0-2 years of Distributed ML Training (FSDP/DDP) experience ~5+ years of...SeniorWork experience placement
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Senior ML Systems Engineer, Inference. Be the first to apply!
- immediate hire remote Wadsworth, OH
- remote scheduling Wadsworth, OH
- remote travel writer Wadsworth, OH
- remote work Wadsworth, OH
- remote internship accounting Wadsworth, OH
- online remote Wadsworth, OH
- remote accounting Wadsworth, OH
- part time remote medical coder Wadsworth, OH
- executive assistant - work from home: remote Wadsworth, OH
- remote Wadsworth, OH





