Senior ML Training Systems Engineer - Distributed CUDA
Genesis AI
Genesis AI in San Francisco is seeking a senior ML infrastructure engineer to design and optimize distributed training systems and performance-critical components. You will profile bottlenecks, implement low‑level code (CUDA, Triton) and ensure efficient hardware utilization across multi‑node GPU clusters. Join a team focused on scalable AI foundations, monitoring tools, and robust performance improvements for large‑scale runs. #J-18808-Ljbffr Genesis AI
- ...Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of... ...models. You will design distributed training systems and optimize GPU utilization... ...5 years of experience in ML infrastructure and a strong...SeniorTraining
- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale... ...to serve humanity. We’re training and deploying frontier models... ...of large‑scale training, distributed systems, and HPC infrastructure... ...performance issues across CUDA/NCCL, networking, IO, and...SeniorTrainingFull timeWork at officeRemote workFlexible hours
- Magic AI, Inc. is seeking a Member of Technical Staff to design and operate distributed systems for serving models in production and driving large-scale post-training workflows. You will work where model execution meets distributed infrastructure, influencing latency,...SeniorTraining
- ...leading AI research firm located in San Francisco is seeking a Senior ML Systems Engineer to build and maintain the training framework for large-scale language models. The role involves designing distributed training solutions and improving training throughput across...SeniorTrainingFlexible hours
$166k - $225k
...improve their business. Founded by engineers — and customer obsessed — we leap... ...be building the next generation distributed data storage and processing systems that can outperform specialized... ...experience, relevant certifications and training, and specific work location....SeniorTrainingLocal areaWorldwide$200.8k - $251k
...optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 -...TrainingFull time- Dormont Manufacturing Co is looking for a Software Engineer for their Pre-training Systems team in San Francisco. Your primary role will be to design and maintain the distributed infrastructure that trains long-context models at scale, tackling challenges related to memory...Training
- Reflection is seeking a role to build and operate distributed training systems powering frontier models in SF. You will work with research teams to... ...training pipelines. This role emphasizes collaboration with ML researchers, debugging across GPU stacks, and improving...SeniorTraining
- ...for an exceptional Staff Software Engineer to help define and build the distributed backend that powers our AI agent... ...patterns, and build the foundational systems that will support millions of... ...systems. Experience with LLMOps or ML infrastructure is highly desirable...SeniorImmediate startRemote workFlexible hours
$229.9k - $286.2k
Senior Lead AI Engineer (Gen AI Platform Services: Distributed Systems) Overview: At Capital One, we are creating responsible and reliable... ...time, our applications of AI & ML are bringing humanity and... ...components including foundation model training, large language model inference...SeniorTrainingFull timePart timeLocal area$148.5k - $223.9k
...Salesforce is seeking a senior engineering candidate to join the... ...proactively design systems that prevent them,... ....Understanding of AI/ML concepts applied to operations... ....Strong knowledge of distributed systems and Linux/... ...promotion, benefits, training, assessment of job...SeniorTrainingFull timeWorldwideWeekend work$160k - $194k
...The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in... ...Demonstrated track record for Senior or Staff level software... ...network architectures and AI/ML landscape is a big plus... ...conditions, experience and training, licensure and...TrainingFull timeRemote workFlexible hours$117.2k - $313.7k
...opportunities for Lead software engineers who want their lines... .../frameworks in distributed filesystems in an ever... ...that improve system scalability, robustness... ...Experience with Big-Data/ML and S3 Hands-on experience... ...promotion, benefits, training, assessment of job performance...SeniorTrainingFull timeImmediate startRemote work$227.2k - $417k
...the Role:As a Software Engineer on the ML Infrastructure team,... ...ML model serving systems that support Deep Learning... ..., and a mentor to senior engineers, fostering... ...throughput, and low latency distributed systems using... ..., ElastiCache, model training orchestration, etc.Understanding...TrainingFull timeTemporary workLocal areaFlexible hours$250k
...platform designed for AI training, experimentation,... ...is looking for a Senior / Staff Site Reliability Engineer to support and... ...closely with platform, ML, and infrastructure... ...across distributed compute environments... ...available infrastructure systems Improve CI/CD pipelines...SeniorTrainingFull timeRemote work$165k - $200k
...with us at Crusoe.About This RoleAs a Senior Systems Engineer, you’ll play a key role in building... ...broader infrastructure needsMentoring and training Service Desk team members, developing... ...off-hours syncs to support distributed operationsWhat You’ll Bring to the Team...SeniorTrainingTemporary work$190k - $230k
...About This RoleWe’re seeking a Senior Systems Engineer to play a key role in... ...teams through code reviews, training, and technical enablement programsEvaluating... ..., including 3+ years in AI/ML or AI application... ...designing scalable, distributed systems in cloud environments...SeniorTrainingTemporary work- ...the da Vinci surgical system and Ion—have transformed... ...worldwide.We’re a team of engineers, clinicians, and... ...Function of PositionAs a Senior Systems GPU Engineer -... ...alongside research, SW/ HW/ ML engineering, regulatory... ...in GPU Compute API - CUDA, OpenCL• Proficiency in...SeniorLocal areaWorldwideFlexible hours
$205.9k - $407.5k
...looking to bring on a Senior Principal ML GPU Architect to lead... ...the Director of ML Engineering who is responsible for... ...function changes in training and inference speed/scale... ...backward passes in CUDA/CuTe. ~ Write... ...inference code for large, distributed training/inference...SeniorTrainingTemporary work- ...hiring a Customer Cluster Engineer to own three to five reserved... ...accounts, managing their training and inference performance end... .... You have 5+ years in distributed- or ML-systems engineering, strong PyTorch... ...FSDP/Megatron-LM debugging, CUDA and Triton/CUTLASS familiarity...SeniorTrainingRemote job
$229.9k - $262.4k
...Overview Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes) Do you love building and pioneering in the technology space?... ...committed to pioneering and responsibly implementing AI/ML across Capital One . We achieve this by building...SeniorFull timePart timeInternshipLocal area- ...Paradigm is seeking a Senior Software Engineer in San Francisco, California. This role requires 7+ years of... ..., particularly in managing and deploying distributed services. Candidates should be skilled in debugging complex systems and should possess familiarity with tools...Senior
$220k - $320k
...diving deep into CUDA kernels, and turning... ...into production systems, we'd love to meet... ...net Inference.net trains and hosts specialized... ...‑person team of engineers who work in‑person... ...Collaborate with applied ML engineers to... ...) Experience with distributed inference and...SeniorTrainingWork at office$15k
...frontier of applying AI/ML to investment... ...lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will... ...while also engineering systemic improvements and... ...or machine learning training systems (Kubeflow,... ...OpenTelemetry)Experience with distributed storage...SeniorTrainingWork at officeLocal areaRemote work$187k
...enable data scientists and ML engineers to develop, train, deploy, and monitor... ...design and build scalable systems that support model training... ...work at the intersection of distributed systems, cloud infrastructure... ...with GPU programming(CUDA) and GPU costs/optimizationStrong...SeniorTrainingFull timeWork at officeLocal areaRemote work$210k - $300k
...economical commerce. We’re training robot AGI to power... ...the world’s best engineers and operators. If... ...to join our ML Infrastructure... ...training and inference systems that power our... ...optimize low-level CUDA kernels. Design... ...and sampling for distributed training. Participate...SeniorTrainingFull timeLocal areaFlexible hours- ...problems. We’re training and deploying frontier... ...are building AI systems. We believe that... ...team of researchers, engineers, designers, and... ...capabilities. Knowledge of distributed training... ...GPU kernels using CUDA, optimising performance... ...into complex ML codebases to identify...SeniorTrainingFull timeWork at officeLocal areaRemote workHome office
- ...startup building production-grade ML infrastructure used by... ...customers. They are looking for a Senior AI/ML Engineer to own model training pipelines, evaluation systems, and inference serving at scale... ...~ Experience with distributed training, GPU optimization, or...SeniorTrainingFull time
$194k - $266k
...one API change. We are looking for a Senior Systems Engineer to help build that layer. This is a deeply... ...You will work across the stack — from distributed systems and high‑throughput APIs to... .... Bonus Points Experience building AI/ML infrastructure, inference platforms, model...SeniorTemporary workLocal areaFlexible hours$300k
...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability... ...AI workloads, automate systems at petascale, and be... ...workloads. Collaborate with ML, networking, and... ...reliability engineering, distributed systems, or hardware acceleration...SeniorTrainingPermanent employment
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior ML Training Systems Engineer - Distributed CUDA. Be the first to apply!
- ai ml engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- data scientist machine learning engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- senior windows systems engineer San Francisco, CA




