ML Systems Engineer — Accelerate GPU Inference & Training
Jobleads-US
Aionia Group in Menlo Park, CA is hiring a Member of Technical Staff, ML Systems to accelerate model training and inference across image, video, and world-model workloads. You will work with a founding team on kernels, runtimes, and distributed engines that power production-scale ML stacks.
You’ll optimize GPU performance, profile bottlenecks with Nsight, and implement low-level CUDA and Triton improvements.
#J-18808-Ljbffr Jobleads-US$250k - $350k
...seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's... ...edge inference acceleration, GPU parallelism, advanced... ...ensuring our creative AI systems deliver industry-leading user... ...models into production.Improve Training Efficiency: (Bonus)...TrainingWork at office3 days per week$250k - $350k
...state-of-the-art models to accelerate breakthroughs across... ...some of the world's leading ML systems engineers, including leaders behind... ...powering our large-scale training, inference, and reinforcement learning... ...environments Optimize memory, GPU kernels and communication...TrainingVisa sponsorship$300k - $400k
...the-art models to accelerate breakthroughs... ...You will own the systems layer that makes... ...our frontier model training and inference fast, efficient,... ...communication and GPU kernels to extract... ...benchmarking distributed ML systems to... ...the scientists, engineers, and problem-solvers...TrainingVisa sponsorshipFlexible hoursShift work- ...take your software engineering career to the next... ...hygiene and system architectureRequired... ...and skillsFormal training or certification on... ...with emphasis on ML systems.Hands-on experience... ...or operating GPU workloads in Kubernetes... ...Serving, Triton Inference Server)Familiarity...Training
- ...innovations that accelerate how teams work, discover... ...Machine Learning Systems Engineer (P60) to lead... ...Design and Build ML SystemsArchitect and... ...scalable systems for training, fine-tuning, and... ...model training, inference pipelines, or search... ...computing, or GPU optimization.Familiarity...TrainingWork at officeLocal area
$220k - $350k
...Job Description ML Infrastructure Engineer Company: Dyna... ...Engineer to own training infrastructure end... ...a multi-cloud GPU fleet into a world... ...low-latency inference pipelines for real... ...hire Multimodal systems (video, audio,... ...DeepSpeed or Accelerate; model serving optimization...TrainingFull timeH1bWork at officeVisa sponsorship$188.5k - $282.7k
...Semantic AI Governance Engine, which is the first system designed to... ....As an Applied ML Engineer on the... ...curating data, training small models, serving... ...Serving and Inference Infrastructure (... ...through shared GPU pools, KV-cache-... ...in Securing and Accelerating the World's AI TransformationRubrik...TrainingPermanent employment$171.7k - $303.9k
...world!The Data Labeling Engineering team designs, builds,... ...engineering, and AI/ML, defining the strategies... ...that create reliable training data at scale. Our tools... ...across teams and systems that unblock the next... ...operational triage, etc) to accelerate understanding,...TrainingFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...Machine Learning Engineer Chicago, IL; New... ...each month. Our systems process and move... ...own our production ML systems and build... ...systems, training/serving parity, retraining... ...computing and GPU-accelerated workloads (e.g.,... ...distributed training/inference), including...TrainingWork at officeImmediate startRemote work
- Cerebras Systems builds the world's largest AI chip,... ...deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...closely with hardware engineers to isolate and resolve... ...failures quickly and accelerate debugging.Reproduce,...Training
- ...JPMorganChase is seeking a Software Engineer III to design and operate an end-to-end ML training platform on AWS and other clouds. You will run GPU workloads, optimize performance, and enable Gen AI workflows within a governed, secure environment. You will collaborate...Training
- ...for a Senior MLOps engineer to work closely... ...build and deploy ML models on a modern... ...model serving systems, hyper-parameter... ...for distributed training on GPU-enabled clusters... ...real-time and batch inference systems, ensuring... ...quantize LLMs for accelerating inference on specific...Training
$180k
...mission is to create AI systems that can accurately... ..., and focused on engineering excellence. This organization... ...ROLE: As an ML Infrastructure... ...building, and scaling GPU compute infrastructure, training frameworks, and... ...data, training, and inference systems Collaborating...TrainingTemporary workWork experience placement$182k - $242k
...technical expertise to accelerate breakthroughs... ...of managing GPU infrastructure.... ...successfully training self-improving... ...fraction of the total inference market, which... ...for strong engineers with great... ...experience with modern ML frameworks such... ...model training systems Experience...TrainingPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$150k
...edge foundation model training, alongside world-class... ...data scientists, and engineers, tackling the most fundamental... ...The Role The GPU Kernel Engineer will... ...at training and inference, and support the team... ...new and cutting-edge systems. The ideal candidate will...TrainingFull timeVisa sponsorship- Cerebras Systems builds the world's largest... ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale... ...The RoleThe Core ML team develops novel... ...Wafer-Scale Engine. Our work spans efficient... ...-level assembly, accelerator programming, or a...Training
- ...a Senior MLOps engineer to work closely... ...build and deploy ML models on a... ...distributed model training on large... ...distributed training on GPU-enabled... ...throughput, real-time inference as well as... ...pipelines to ensure system health,... ...quantize LLMs for accelerating inference on specific...Training
$250k - $320k
Staff Infrastructure Engineer We are partnered with a Stealth... ...ll design and optimize the inference platform, GPU‑based training clusters, and data... ...play a key role in scaling systems for both research and production... ...experience in Software / ML Infrastructure Engineering...TrainingFull timeImmediate start- ...and Conversational AI system that integrates... ...ecosystem.About the AI & ML Platform TeamOur... ...a Senior ML System Engineer on the AI & ML Platform’s Inference team, you will design... ...level optimizations (GPU kernels, quantization... ...reliablyBenchmark, fine-tune, and accelerate inference engines....Work at officeLocal area
- ...life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)... ...the central nervous system of SpaceX - we create... ...throughout SpaceX to accelerate launch vehicle production... ...support for training workloads.RESPONSIBILITIES... ...including low-level GPU kernel work, quantization...TrainingPermanent employmentTemporary workRemote workWorldwideWeekend work
$150k - $230k
..., recommendation systems, and adtech.Recognized... ...Machine Learning Engineer to drive the post-training of our large... ...on mid-to-large GPU clusters, applying... ...engineering for ML. You can independently... ...Hugging Face TRL/Accelerate, DeepSpeed or FSDP, and inference engines like vLLM...TrainingFull timeLocal areaWork from home$173k - $253k
Matterport - Senior ML Ops Engineer Job Description... ...performance, optimize inference speed and resource utilization... ...model architectures, training procedures, and... ...optimization, hardware acceleration, and efficient AI.... ...with version control systems (e.g., Git) and agile...TrainingFull timeWork at officeWork from home$119.8k - $234.7k
...and Infrastructure Engineering (SCHIE) is the... ...Microsoft's Hardware Systems organization is... ...combines custom accelerators, advanced networking... ...-leading AI training and inference capabilities. The... ...systems. Translate AI/ML workload... ...:Experience with GPU, FPGA, TPU, or custom...TrainingOngoing contractWork at officeLocal areaWorldwide3 days per week$180k - $280k
...you passionate about accelerating the future of autonomous... ....As a Staff AI/ML Engineer within the Onboard Embodied... ...onboard ML systems powering fully autonomous... ...sophisticated neural networks trained from large-scale... ...capable of real-time inference and robust autonomous...TrainingFull timeLocal areaWork from homeRelocationRelocation package$197.5k - $272k
...trust Sonatus to accelerate this shift. Our... ...Learning Engineer to join our seasoned... ..., including system logs, traces, and... ...the end-to-end ML pipeline—from data... ...and model training to deployment on... ...execution on CPU/GPU-bound targets... ...(C++14/17 for inference).Deep proficiency...TrainingWork at officeWorldwideFlexible hoursShift work3 days per week$153.2k - $234.1k
...hardware and battery systems to intuitive design,... ...you passionate about accelerating the future of autonomous... .... As a Senior ML Infra Engineer, you will work on the... ...dataset generation, training, evaluation and iteration... ...training across large GPU/CPU clusters or specialized...TrainingFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...About the job ML Engineer Our Client Is a rapidly growing Tier 1 VC backed... ...the evolving world of intelligent systems. Location : New York, NY Work... ...acquisition, preprocessing, model training, deployment, inference, and monitoring in production environments...TrainingFull time
$180.9k - $265.32k
...sustainability, design and engineering, ambition and... ...driving systems. This role focuses... ...on productizing ML models, ensuring... ...Optimization Optimize inference pipelines using... ...1/2, OpenCV, and GPU acceleration. Hands-on experience... ..., education and training; certifications;...TrainingHourly payNight shift$150k - $230k
...About Clockwork Systems Clockwork.io – Software... ...to increase GPU cluster utilization... ...and veteran systems engineers who share a vision... ...and performance acceleration that dynamically... ...performance distributed GPU training. You'll work at... ..., InfiniBand) ML framework or...Training$295.25k - $345.04k
...Learning Infrastructure Engineer, you’ll build... ...that powers ML systems across our organization... ...boundaries of large-scale training and serving. Your... ...to low-latency inference and production... ...quantization, optimizing GPU utilization and... ...ML platforms that accelerate model development...TrainingFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Systems Engineer — Accelerate GPU Inference & Training. Be the first to apply!
- senior staff systems engineer Menlo Park, CA
- healthcare systems engineer Menlo Park, CA
- systems engineer Menlo Park, CA
- data engineer machine learning Menlo Park, CA
- artificial intelligence - machine learning intern Menlo Park, CA
- machine learning Menlo Park, CA
- machine learning research scientist Menlo Park, CA
- machine learning scientist Menlo Park, CA
- computer vision machine learning engineer
- ai ml developer



