ML Infrastructure Engineer
$350kThinking Machines Lab
AI Infrastructure EngineerThe mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.What You'll DoOwn the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completionPartner directly with research teams during active model runs, embedding with them to unblock training and speed up iterationDebug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root causeBuild monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stallingImprove checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of computeBuild internal tools that reduce toil and improve cluster utilization across post-training and RL workloadsParticipate in an on-call rotation supporting production model runsWrite postmortems and turn recurring failure patterns into permanent infrastructure fixesMinimum Qualifications4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in productionTrack record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issuesStrong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a systemSolid grounding in Linux systems internals and networking fundamentalsComfortable owning production systems, including participating in on-call rotationsPreferred QualificationsExperience operating GPU or TPU training clusters at scaleFamiliarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloadsExperience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performanceExperience building observability tooling purpose-built for ML training, not just general infrastructureA track record of thriving in fast-changing, research-driven environments where priorities shift with the scienceLogisticsLocation: This role is based in San Francisco, CA.Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- ...oprecruiting.comPhone: (***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro Area (100% On-Site)... ...company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks...SuggestedFull timeWork at officeFlexible hours
- ...ML Infrastructure EngineerSpectral Labs is a spatial intelligence company building reasoning models for engineering physical systems. Our model SGS-1 is state-of-the-art for parametric geometry, and we are currently building the next generation of models to revolutionize...Suggested
- ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and...SuggestedWork at officeVisa sponsorship
$190k - $210k
...backed by climate-tech and Silicon Valley investors. For more information, please visit Role Description As a Senior ML Infrastructure Engineer, you will work directly in the Automation org with the core ML, Ops, and Analytics teams to help improve and build out...SuggestedLocal area- ...000s of developers and enterprise users.ML performance, quality, and systems acumen... ...Qualifications 7+ years of software engineering experience building and operating enterprise... ...building the next layer of enterprise AI infrastructure.Job SummaryCategory: EngineeringSuggestedWork at office
- Anthropic is seeking an engineer to own the infrastructure behind safeguards research. You will build the tooling researchers rely on to run experiments, train detection methods, and select detections for launch. This role sits between research and production, ensuring...
$292k - $417.2k
...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction...Full timeTemporary workLocal areaFlexible hours- ...Job Description Job Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to... ...video streams. As a Machine Learning Engineer in ML Runtime & Optimization , you will develop technologies...
$227.33k - $312.58k
We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization. In this role, you’ll be responsible for... ...Observability: Airflow, Dagster, data quality and lineage toolsCloud & Infrastructure: AWS or GCP, containerized data workloads, CI/CD,...Full timeWork at officeLocal areaImmediate start3 days per week$148.5k - $223.9k
...the right place! Agentforce is the future of AI, and you are the future of Salesforce.This role is for a Senior Machine Learning Engineer within the Trust Intelligence Platform team who will architect data-driven strategies for threat detection across the security organization...Full time$166k - $210.25k
RDQ127R59SummaryAs a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and... ...algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from cluster...Local areaWorldwide$198k - $230k
...innovation, and adapt to individual work styles. Senior MLOps Engineer (Applied AI Focus) As a Senior MLOps Engineer on our Product... ...engineering to integrate measurement loops into our broader infrastructure (AWS/GCP), ensuring our model lifecycle is automated and observable...Work at officeRemote workWork from homeWorldwideHome officeFlexible hours$170.1k - $258.3k
...export, kernel development, and performance engineering so that every cycle on our accelerators... ...that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous... ...workloads.Build and improve tooling and infrastructure that make it easier to profile, debug, and...Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours- A decentralized AI platform company in the United States is seeking an experienced ML Training Platform Engineer to design and build robust infrastructure for ML training. The ideal candidate has over 5 years in infrastructure and platform engineering, with expertise in...
- ...moonshot AI lab focused on mechanistic interpretability, new architectures, and pretraining science. As an ML Engineer, you will build and operate the infrastructure enabling cutting-edge research in training and evaluating large models. You will optimize inference and...
- ...Ensure that ML models can be effectively developed, deployed, managed, and monitored in Production environments. Productionize ML... ...training, validation, and deployment utilizing CI/CD practices. Infrastructure management – set up and manage infrastructure for ML workloads...Permanent employmentContract workLocal area
- ...Job Description We are currently looking for an exceptional engineer to work in the position of Machine Learning Engineer on the Personalization team. The Personalization team at Boomtrain are responsible for designing and building the models and systems that provide...
$130k - $240k
...researchers, designers, growth experts, and engineers rethinking human-computer interaction... ...well-known angels. About The Role As a ML engineer at Wispr, you’ll play a crucial... ...features of our voice interface, building infrastructure to handle What are we looking for?...- ...RaindropRaindrop Raindrop is the monitoring platform for AI agents. Engineering teams at Fortune 100s and the fastest-growing AI companies (... ...of requests a day + growing. Architect, implement, and scale ML pipelinesQuick iteration without compromising on qualityDeeply...Temporary work
$150k - $250k
Garuda Ventures is looking for a Senior Software Engineer (IC) to develop cloud and on-prem systems that will power future factories... ...software and a strong background in backend systems and infrastructure. The position offers a competitive salary range of $150,000—$...- ...that every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first... ...and designing new features of our voice interface, building infrastructure to handle What are we looking for? Previous founding or startup...H1bWork at officeRemote workRelocationVisa sponsorshipFlexible hours
$128.7k - $261.3k
...approaches to model export, kernel development, and performance engineering so that every cycle on our accelerators translates into better... ...tooling that makes that path fast, reliable, and effortless for ML engineers across the AV organization to compile their models....Full timeLocal areaRemote workWork from homeRelocation packageFlexible hours$170k - $200k
...advantage in the robotics revolution. The Instawork Robotics ML Engineer will help build and scale the technology powering physical... ...with expertise in distributed systems, cloud computing on AWS infrastructure, and scalable data processing for large-scale datasets.An...Hourly payInternshipLocal areaShift work$212k - $318k
...creator economy and are looking for a Senior Machine Learning Engineer, Infrastructure to support our mission.This role is based in San Francisco... ...TeamYou'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how...Work at officeLocal areaRemote workWorldwideFlexible hours2 days per week3 days per week- ...creating a genuinely open, collaborative path to frontier‑scale AI. We’re looking for an ML Training Platform Engineer to architect, build, and scale the foundational infrastructure powering our decentralised ML training platform. You will own core systems spanning...Work experience placement
$197.3k - $225.1k
...Lead AI/ML Engineer (Platform, kubeflow) Overview At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine...Full timePart timeLocal area- Faire is seeking a Staff Machine Learning Platform Engineer to design, improve, and operate a scalable ML platform that accelerates model training,... ...Spark, Delta Lake, MLflow, Python and SQL, cloud/infrastructure-as-code, and strong MLOps practices. #J-18808-Ljbffr...Remote jobLocal area
- Google Cloud is seeking a Customer Engineer in San Francisco to partner with technical sales as an AI/ML subject matter expert, helping customers and partners understand Google Cloud and develop creative cloud solutions. You will engage in proofs of concepts, present to...
$166k - $244k
A leading technology company is seeking a Research Software Engineer to develop next-generation technologies that enhance communication through advanced software and hardware. This role involves collaborating with researchers, optimizing machine learning performance, and...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Infrastructure Engineer. Be the first to apply!
- data scientist machine learning engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- computer vision machine learning engineer San Francisco, CA
- machine learning engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- graduate machine learning engineer San Francisco, CA
- machine learning software engineer San Francisco, CA
- junior machine learning research engineer San Francisco, CA
- senior ml engineer San Francisco, CA
- data infrastructure engineer San Francisco, CA


