MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)
$90 - $120 per hourMercor
Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced Large Language Models.
1\. Overview
Join a leading AI lab's cutting-edge GenAI team and help build foundational AI models from the ground up. We're seeking MLOps Engineers with hands-on experience in large language model infrastructure across any of four areas: GPU kernel programming, performance profiling and trace analysis, debugging accelerated and distributed workloads, and high-throughput inference serving. This role involves AI model training and evaluation work, including writing and assessing MLOps and ML systems tasks and solutions to generate high-quality training data for frontier AI systems.
This is a W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI Lab as part of their extended workforce. This is a 40-hour full-time engagement, with no conflicts/no other engagements.
2\. Key Responsibilities
- Design challenging, domain-relevant tasks across four areas, GPU kernels, performance profiling, debugging, and inference serving, and write accurate, well-structured solutions to them.
- Guide research and engineering teams to close knowledge gaps and improve AI model performance on ML systems, training infrastructure, and framework-level topics.
- Evaluate MLOps and ML systems tasks and solutions, and provide clear, written technical feedback that stands up to reviewer scrutiny.
- Develop guidelines and detailed rubrics or evaluation frameworks covering kernel-level optimization, profiler output interpretation, distributed systems reasoning, and serving throughput and latency trade-offs.
- Collaborate with other subject matter experts to keep training data consistent and accurate.
3\. Core Qualifications
- 2+ years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering. This is a hands-on systems role rather than an applied modelling or data science one.
- Practical experience in at least one of the following, with more than one a strong plus: writing or optimizing custom GPU kernels (CUDA, Triton, Pallas); performance profiling and trace analysis (Kineto, torch.profiler, Nsight, XLA or JAX profiler); debugging distributed or accelerator-bound workloads; serving large language models at scale (vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, continuous batching).
- Working production experience with JAX and/or PyTorch. Framework-level depth is a strong plus: custom operators, distributed training (FSDP, DDP, DeepSpeed, Megatron), or compiler and graph-level work.
- Familiarity with modern accelerators such as A100, H100, B200 or TPU, and the ability to reason about throughput, latency and memory trade-offs.
- Demonstrable career progression.
- Ability to engage reliably for at least 40 hours/week during weekdays.
- Strong written communication skills and the ability to explain complex technical decisions clearly.
About Cincinnatus LLC:
Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.
Equal Employment Opportunity:
Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.
- ...up. We're seeking MLOps Engineers with hands-on experience... ...of four areas: GPU kernel programming, performance profiling and trace analysis... ...inference serving. This role involves... ...assessing MLOps and ML systems tasks and... ...SGLang, TensorRT-LLM, Ray Serve, KV cache...SuggestedFull timeContract workTemporary workWeekday work
$90 - $120 per hour
Gridnaut Recruiting is hiring a remote MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) contractor (pay $90-$120/hr). Contribute to frontier AI research and evaluation work. Ideal candidates: Have 2+ years of hands-on professional experience in ML systems,...SuggestedTemporary workFor contractorsRemote workWeekday work$90 - $120 per hour
...General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) Type: Contract Compensation: $90–$120/hour Location: Remote Commitment: 40 hours/...SuggestedContract workSummer workRemote work- ...organization, apply now.We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC),... .../decode optimization and KV cache management. Inference Serving: Deploy and manage inference engines including vLLM and...SuggestedWork at officeRemote workFlexible hours
$200k - $400k
...Machine Learning Engineer (Reinforcement... .../Quantization/GPU), Multiple roles... ...advanced AI systems that can operate... ...high-performance kernels, compression... ...Language Models, LLM, Foundation Models... ..., Model Serving, Parallel Computing... ...Model Evaluation, MLOps Disclaimer...SuggestedRemote work$150k - $230k
...advanced AI, recommendation systems, and adtech.Recognized... ...-on Machine Learning Engineer to drive the post-... ...training on mid-to-large GPU clusters, applying distributed... ..., and ensure training-serving consistency.Stay... ...code.RequirementsHands-on LLM post-training...Full timeLocal areaWork from home$200k - $280k
...AnalyticsJob Number: 192105Eligible for remote: YesApply: JobThe GPU- MLOps Engineer will own the Azure platform layer end-to-end — deploying AI/... ..., partnering closely with data scientists, ML engineers, and LLM engineers to keep everything running smoothly. What You Will...Remote workFlexible hours$150k - $220k
...looking for a ML Systems Engineer, Inference. We want... ...the world to run LLM inference, meaning... ...effort. You'll own LLM serving performance end to... ...repeatable. Profile and diagnose... ...management down to kernels and interconnect.... ...node and multi-node GPU deployments. Turn...Full timeRemote workHome officeVisa sponsorshipWork visaFlexible hours$100k - $150k
...MLOps Engineer -Remote Bright Vision Technologies is a technology... ...platforms for serving large machine learning... ...The role focuses on the systems engineering side of AI... ...KV cache strategies for LLM serving workloads. Integrate... ...understanding of GPU architecture, memory hierarchies...Full timeH1bLocal areaRemote workVisa sponsorship$150k - $175k
...observability, and the AI serving and routing layer.... ...treat that as an engineering responsibility,... ...agentic delivery system (the loops that... ...you document ~ MLOps experience: deploying and operating LLM or ML systems in production... ...Jobot candidate profile, and any job...Temporary workLocal areaWork from homeHome officeFlexible hoursNight shift$100k - $150k
...ML Performance Engineer - Remote Bright... ...neural network systems. The role spans the... ...from low-level kernel optimization to distributed... ...understanding of GPU architecture,... ...Profile and optimize end-to... ...speculative decoding for LLM serving. Drive compiler...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$100k - $150k
...MLOps Engineer - Remote Bright Vision Technologies... ...inference platforms for serving large machine learning... ...The role focuses on the systems engineering side of AI... ...caching, autoscaling, GPU utilization, and end-to... ...KV cache strategies for LLM serving workloads. Integrate...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship$295k
...who are building AI systems. We believe that... ...team of researchers, engineers, designers, and... ...Implementing performance profiling across the ML... ...for large-scale LLM training.Design distributed... ..., or custom kernels/fused ops.... ...with evaluation and serving frameworks (vLLM,...Full timeWork at officeLocal areaRemote workHome office- ...frontiers through novel datasets and experimentation. We are building a cohort of senior ML researchers focused on inference and GPU kernel engineering to define criteria for correctness and excellence on real-world problems. This fully remote, contract role invites...Contract workRemote workFlexible hours
$135 per hour
...We believe great AI comes from exceptional, human-generated data. This remote, contract role focuses on ML inference and GPU kernel engineering. Competitive hourly pay ($135/hr) based on experience, with a broader range ($100–$170/hr) and a flexible, asynchronous...Remote jobHourly payContract workFlexible hours- ...AfterQuery is assembling a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems, kernel optimization, and deployment infrastructure. This is research-and-evaluation work, not production engineering, centered on defining...Remote jobContract workWork at office
- ...senior ML researchers to define correct and excellent performance on hard, real-world problems, focusing on inference and GPU kernel engineering. This is research-and-evaluation work, not production engineering. The role is remote, flexible, and asynchronous, with...Remote jobFlexible hours
- ...AfterQuery seeks an ML Research Engineer focusing on inference and GPU kernels for remote, contract work. You will architect challenging evaluation problems, craft reference solutions, and judge AI outputs with rigor and domain insight. The role emphasizes research...Contract workRemote work
- ...AfterQuery is building a research lab focused on ML inference and GPU kernel engineering, covering inference serving systems, kernel optimization, and deployment infrastructure for large models. This is research-and-evaluation work, not production engineering, defining...Remote jobHourly pay
- ...AfterQuery is seeking an ML Research Engineer focusing on inference and GPU kernels. This remote contract role involves designing research scenarios, creating reference solutions, and grading AI outputs. You’ll contribute to frontier AI evaluation rather than production...Remote jobHourly payContract workFlexible hours
$170.1k - $258.3k
...capable fully self-driving systems, to move us toward... ...accessible mobility. For the AI Kernels & Compilers team, that... ..., and performance engineering so that every cycle on... ...high‑performance GPU kernels and custom libraries... ...that make it easier to profile, debug, and validate CUDA...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$193.3k - $261.5k
...Trainium.The Acceleration Kernel Library team is at the... ...software boundary, our engineers craft high-performance... ..., and machine learning systems, you'll bring expertise... ...analysis using profiling tools to identify and resolve... ...architectures- Experience with GPU kernel optimization and...InternshipLocal areaWork from homeFlexible hours- ...AfterQuery is building a research cohort focused on ML inference and GPU kernel engineering, spanning inference serving systems to deployment infrastructure. This is research-and-evaluation work, not production engineering, defining what is considered correct and excellent...Remote job
$200k - $230k
SVP, Lead AI/MLOps Infrastructure Engineer - Full Time - HybridWe’re partnering... ...to model serving, reliability, and cost... ..., compute (including GPU), storage, and model... ...deploymentProductionize AI/ML and GenAI (LLM) workloads in... ...similar)Solid Linux, systems, and troubleshooting...Full timeRemote work- About the RoleAs Senior MLOps Engineer, you will focus on supporting cross... ..., including model serving (real-time and batch), training... ...environments, and orchestration systems, with a focus on performance,... ...orchestration of autonomous workflows or LLM-driven agentsJob SummaryJob...Remote work
$178.2k - $232.65k
...Engineering | Seattle, United States | Remote, Remote |... ...legal entity. As a ML System Engineer on the AI &... ...large-scale model serving systems end-to-end. You... ...-level optimizations (GPU kernels, quantization, speculative... ..., Triton, TensorRT-LLM, etc.). It would...Work at officeLocal areaRemote work$190.2k - $345.65k
...Staff Machine Learning Engineer to architect and... ...intelligence — the systems that turn massive... ...retrieval stack that serves it, and the tool... ...correctness, and cost profile enterprise scale... ...ANN index tuning, GPU-accelerated enrichment... ...building retrieval for LLM and agentic systems...Full timeTemporary workLocal areaWorldwide$300k - $400k
...You will own the systems layer that makes our... ...: scheduling, kernels, RDMA, weight synchronization... ...and offline profilers that surface... ...communication and GPU kernels to extract... ..., scheduling, and serving architecture at production... ...— the scientists, engineers, and problem-...Visa sponsorshipFlexible hoursShift work- ...implement robust model serving infrastructure using platforms... ...# Improve GPU usage, enable autoscaling... ...degree in Computer Science, Engineering, or a related field (or... ...years of experience as MLOps engineer or DevOps... ...~ Experience in AI/ML systems security, compliance, and...Work at officeRemote workRelocation package
- ...Role : MLOPS Engineer / Architect Location : Charlotte NC (Onsite)... ...Kubernetes ML pipelines Model serving CI/CD Monitoring... ...Conduct continuous performance profiling, load testing, and... ...monitoring, alerting, and logging systems to track model drift, data quality...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling). Be the first to apply!
- senior linux systems engineer Remote
- system engineer contract Remote
- data systems engineer Remote
- senior staff systems engineer Remote
- microsoft systems engineer Remote
- operations support system engineer Remote
- advanced systems engineer Remote
- software system engineer Remote
- space systems engineer Remote
- adas systems engineer Remote




