Remote Senior ML Systems Engineer, Inference
$150k - $220kgrabjobs
- Remote job
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on. We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale. Learn more in our CEO's funding announcement: . We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You'll lead that effort. You'll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production. Responsibilities Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable. Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect. Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments. Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically. Work closely with product and infrastructure teams to shape how inference is offered on Runpod. Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what's worth adopting, what's worth building, and what's worth contributing back. Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough. Requirements 5+ years of professional system engineering experience. Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale. Strong software engineering skills in Python . You're comfortable working in large, performance-critical codebases. A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput. Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving. Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools. The ability to explain your results clearly in writing and turn them into decisions. Preferred Experience writing or tuning GPU kernels in CUDA or Triton. Contributions to inference or ML systems projects. Experience with multi-node GPU systems and high-speed networking. Experience at a company where inference cost and latency were core business metrics. What You’ll Receive: The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. Generous medical, dental & vision plans Flexible PTO- take the time you need to recharge Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas. grabjobs
$150k - $220k
...platform has processed more than 20 billion inference requests. We closed a $100M Series A... ...will depend on. We're a small, remote-first team. We take ownership seriously... ...announcement: . We're looking for a ML Systems Engineer, Inference. We want Runpod to be the best...Remote jobSeniorHome officeVisa sponsorshipWork visaFlexible hours- ...Amazon is seeking a Senior Software Development Engineer to design and implement inference data plane software for large models on custom hardware. You will own end-to-end features from kernel development to deployment, enable efficient data movement, and drive performance...Senior
- ...At Atlassian, the ML System Engineer will design and optimize large-scale model serving systems, spanning distributed infrastructure... ...serving systems, benchmarking and tuning inference engines, and partnering with senior ML engineers to deploy open-source LLMs. #J-18808...Remote job
- ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research,...Remote workSeniorContract work
- ...AfterQuery is building a research lab focused on ML inference and GPU kernel engineering, covering inference serving systems, kernel optimization, and deployment... ...years of experience or equivalent. It’s fully remote, asynchronous, and offers competitive hourly compensation...Remote jobSeniorHourly pay
- ...datasets and experimentation. We are building a cohort of senior ML researchers focused on inference and GPU kernel engineering to define criteria for correctness and excellence on real-world problems. This fully remote, contract role invites experts to design scenarios,...Remote workSeniorContract workFlexible hours
- ...experimentation. We build a cohort of senior ML researchers to define correct and... ...hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation... ...production engineering. The role is remote, flexible, and asynchronous, with...Remote jobSeniorFlexible hours
$135 per hour
...datasets and experimentation. We believe great AI comes from exceptional, human-generated data. This remote, contract role focuses on ML inference and GPU kernel engineering. Competitive hourly pay ($135/hr) based on experience, with a broader range ($100–$170/hr) and a...Remote jobSeniorHourly payContract workFlexible hours- ...platform for the mobile app economy. As a Senior Staff Software Engineer on the Serving team, you will design,... ...time, while optimizing GPU-powered inference pipelines. You will own changes... ...capacity planning, collaborating with ML engineers to productionize #J-18808...Remote jobSenior
$174.9k - $261.3k
Remote/Hybrid Sunnyvale, California, United States of America Remote... ...the world! The Data Labeling Engineering team designs, builds, and... ...engineering , data engineering , and AI/ML , defining the strategies,... ..., and direct impact on systems that unblock the next generation...Remote workSeniorFull timeLocal areaWork from homeRelocation packageFlexible hours$180k - $240k
...for one of its clients a Senior Machine Learning Engineer - this is a fully remote role for US/Canada... ...join our small but mighty ML team building production... ...can ship LLM-powered systems that handle real, high-... ...problems — low latency inference, hallucination reduction...Remote workSeniorFlexible hours- ...NVIDIA Corporation in Austin, TX is seeking a Senior Software Engineer to build accelerated PyTorch-based solutions for large-scale ML models. You will work on GNNs, TFMs, and ensemble models, advancing training and inference on GPU infrastructure. You will...Senior
- ...Qualcomm Incorporated in San Diego, CA seeks a Core ML Engineer to design, develop, and optimize machine learning systems powering next-gen AI platforms and applications... ...You will build scalable ML pipelines, optimize inference, and deploy models across APIs and...Senior
- ...complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems... ...in developing large distributed systems or high-load web services. ~ Open-...Remote workSenior
- ...building large in-house AI/ML infrastructure. Built by engineers, for engineers. From... ...scale GPU orchestration to inference optimization, we own the... ...approaches on real-world systems. The results will often lead... ...currently looking for senior- and staff-level ML engineers...Remote workSeniorFull time
- ...office, home, or hybrid arrangements. Atlassians can work from anywhere with a legal entity, and this ML System Engineer role sits on the AI & ML Platform's Inference team, designing scalable model-serving infra. The position requires 3+ years in software engineering...Work at officeWork from homeFlexible hours
- ...machine learning and LLMs. You will work with data scientists and software engineers to build scalable production solutions, deploying and monitoring models for millions of medical charts daily. Remote work from anywhere in the U.S. is supported, with in-person...Remote jobSenior
$144.7k - $261.3k
...autonomous vehicle development. We engineer high-performance tools... ...with data-intensive ML teams to drive rapid innovation... ...-generation autonomous systems. About the Role As a Senior Engineer in the Embodied... ...#GM-AV-1 This role is based remotely, but if the selected...Remote workSeniorLocal areaWork from homeFlexible hours- ...NVIDIA Corporation is seeking a Senior Machine Learning Applications and Compiler Engineer to advance end-to-end inference optimization across NVIDIA platforms. You will work at the intersection of large-scale systems, compilers, and deep learning to map neural network...Senior
- ...Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances. You will lead customization of open-source...Senior
$120 per hour
...Dorsey . Position: MLOps Engineer, LLM Systems (Serving, GPU Kernels,... ...0–$120/hour Location: Remote Commitment: 40 hours/week... ...profiling , debugging , and inference serving . Write accurate,... ...AI model performance on ML systems and training infrastructure...Remote jobContract workSummer work- ...Samsara is hiring a Machine Learning Engineer for Safety AI who will design production ML APIs, build data flywheels, and optimize artifacts... ...and ensure low latency, high-throughput inference across millions of devices. This remote role is open to candidates in the US or...Remote jobSenior
$195k - $217k
...closely with other engineering teams across the... ...machine learning (ML) and hybrid ML/LLMsolutions... ...bar for ML systems company-wide. Key... ...teams, guiding senior engineers and influencing... ...training and inference (including LLM serving... .... #LI-AJ1 #LI-remote $195,000 - $217,0...Remote job$145k - $165k
...offering tremendous career growth potential. Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 Salary... ...build, and operate high-performance, highly reliable inference platforms for serving large machine learning models...Remote workFull timeH1bLocal areaImmediate startVisa sponsorship$160k - $220k
...approaches Solid machine learning engineering foundations, including... ...scheduling, or containerized inference Experience with... ...identifying information out of the systems you touch Technologies:... ...engineers. The position is hybrid remote in Cambridge, MA 02142, and includes...Remote workSeniorFull time$120.1k - $214.5k
...and help make the health system work better for... ...scientists and software engineers through data extraction... ...the flexibility to work remotely * from anywhere within... ...Responsibilities:Ship production ML systems end-to-end:... ...with low-latency inference and high availabilityBuild...Remote workSeniorMinimum wageFull timeWork experience placementWork at officeLocal area- ...Office Location: Remote, US Based JOB SUMMARY... ...We are seeking a Senior/Principal Machine Learning Engineer with deep expertise... ...-of-the-art AI systems for image, video, and... ...models using modern ML infrastructure including... ..., distributed inference). Architect and...Remote workSeniorH1bWork at officeLocal areaWork visa
- ...building intelligent systems that sit at the intersection... ..., and large-scale engineering. Our goal is to... .... The Role As a Senior AI/ML Engineer, you will lead... ...around model serving, inference efficiency, and lifecycle... ...technical depth. Flexible remote or hybrid work...Remote workSeniorFlexible hours
- ...AfterQuery is assembling a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems, kernel optimization, and deployment... ...domain-specific problems. The role is fully remote, asynchronous, and contract-based, with emphasis...Remote jobContract workWork at office
- ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production...Remote jobHourly payContract workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Senior ML Systems Engineer, Inference. Be the first to apply!
- immediate hire remote Perth Amboy, NJ
- remote scheduling Perth Amboy, NJ
- remote work Perth Amboy, NJ
- remote data entry no experience Perth Amboy, NJ
- part time remote medical coder Perth Amboy, NJ
- executive assistant - work from home: remote Perth Amboy, NJ
- remote Perth Amboy, NJ
- entry level finance remote Perth Amboy, NJ
- remote customer service chat Perth Amboy, NJ
- remote from home Perth Amboy, NJ




