Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

ML Infrastructure Engineer

$350k

Thinking Machines Lab

AI Infrastructure EngineerThe mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.What You'll DoOwn the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completionPartner directly with research teams during active model runs, embedding with them to unblock training and speed up iterationDebug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root causeBuild monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stallingImprove checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of computeBuild internal tools that reduce toil and improve cluster utilization across post-training and RL workloadsParticipate in an on-call rotation supporting production model runsWrite postmortems and turn recurring failure patterns into permanent infrastructure fixesMinimum Qualifications4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in productionTrack record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issuesStrong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a systemSolid grounding in Linux systems internals and networking fundamentalsComfortable owning production systems, including participating in on-call rotationsPreferred QualificationsExperience operating GPU or TPU training clusters at scaleFamiliarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloadsExperience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performanceExperience building observability tooling purpose-built for ML training, not just general infrastructureA track record of thriving in fast-changing, research-driven environments where priorities shift with the scienceLogisticsLocation: This role is based in San Francisco, CA.Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the ML Infrastructure Engineer in San Francisco, CA vacancy
  •  ...oprecruiting.comPhone: (***) ***-****Job Title: Machine Learning Infrastructure EngineerLocation: San Francisco, CA Metro Area (100% On-Site)...  ...company is seeking a Machine Learning Infrastructure Engineer to help architect the compute, training, and execution frameworks... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Objective Paradigm

    San Francisco, CA
    1 day ago
  •  ...ML Infrastructure EngineerSpectral Labs is a spatial intelligence company building reasoning models for engineering physical systems. Our model SGS-1 is state-of-the-art for parametric geometry, and we are currently building the next generation of models to revolutionize... 
    Suggested

    Spectral Labs

    San Francisco, CA
    7 hours ago
  •  ...Job Description Job Description We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and... 
    Suggested
    Work at office
    Visa sponsorship

    Spellbrush

    San Francisco, CA
    3 days ago
  • $190k - $210k

     ...backed by climate-tech and Silicon Valley investors. For more information, please visit  Role Description As a Senior ML Infrastructure Engineer, you will work directly in the Automation org with the core ML, Ops, and Analytics teams to help improve and build out... 
    Suggested
    Local area

    Gridware

    San Francisco, CA
    3 days ago
  •  ...000s of developers and enterprise users.ML performance, quality, and systems acumen...  ...Qualifications 7+ years of software engineering experience building and operating enterprise...  ...building the next layer of enterprise AI infrastructure.Job SummaryCategory: Engineering
    Suggested
    Work at office

    Atlassian

    San Francisco, CA
    3 days ago
  • Anthropic is seeking an engineer to own the infrastructure behind safeguards research. You will build the tooling researchers rely on to run experiments, train detection methods, and select detections for launch. This role sits between research and production, ensuring... 

    Jobzhr

    San Francisco, CA
    3 days ago
  • $292k - $417.2k

     ...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    7 hours ago
  •  ...Job Description Job Description The AI Infrastructure team at Zensors builds the engine that powers our visual sensing platform. We provide the tools to...  ...video streams. As a Machine Learning Engineer in ML Runtime & Optimization , you will develop technologies... 

    Zensors

    San Francisco, CA
    3 days ago
  • $227.33k - $312.58k

    We’re looking for a Staff ML Data Engineer to join Procore’s AI & Frontier Models organization. In this role, you’ll be responsible for...  ...Observability: Airflow, Dagster, data quality and lineage toolsCloud & Infrastructure: AWS or GCP, containerized data workloads, CI/CD,... 
    Full time
    Work at office
    Local area
    Immediate start
    3 days per week

    Procore Technologies

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...the right place! Agentforce is the future of AI, and you are the future of Salesforce.This role is for a Senior Machine Learning Engineer within the Trust Intelligence Platform team who will architect data-driven strategies for threat detection across the security organization... 
    Full time

    Salesforce

    San Francisco, CA
    7 hours ago
  • $166k - $210.25k

    RDQ127R59SummaryAs a Senior Applied ML Engineer on the Applied AI team at Databricks, you will use machine learning, scheduling, and...  ...algorithms to maximize the efficiency and performance of our infrastructure. Your work will span the entire stack—from cluster... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $198k - $230k

     ...innovation, and adapt to individual work styles. Senior MLOps Engineer (Applied AI Focus) As a Senior MLOps Engineer on our Product...  ...engineering to integrate measurement loops into our broader infrastructure (AWS/GCP), ensuring our model lifecycle is automated and observable... 
    Work at office
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    CreatorIQ

    San Francisco, CA
    1 day ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our accelerators...  ...that sit at the heart of our on‑vehicle ML inference for ADAS and autonomous...  ...workloads.Build and improve tooling and infrastructure that make it easier to profile, debug, and... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    1 day ago
  • A decentralized AI platform company in the United States is seeking an experienced ML Training Platform Engineer to design and build robust infrastructure for ML training. The ideal candidate has over 5 years in infrastructure and platform engineering, with expertise in... 

    Pluralis Research

    San Francisco, CA
    4 days ago
  •  ...moonshot AI lab focused on mechanistic interpretability, new architectures, and pretraining science. As an ML Engineer, you will build and operate the infrastructure enabling cutting-edge research in training and evaluating large models. You will optimize inference and... 

    Tilde Research

    San Francisco, CA
    4 days ago
  •  ...Ensure that ML models can be effectively developed, deployed, managed, and monitored in Production environments. Productionize ML...  ...training, validation, and deployment utilizing CI/CD practices. Infrastructure management – set up and manage infrastructure for ML workloads... 
    Permanent employment
    Contract work
    Local area

    Robotics Prcocess Automation, LLC

    San Francisco, CA
    2 days ago
  •  ...Job Description We are currently looking for an exceptional engineer to work in the position of Machine Learning Engineer on the Personalization team. The Personalization team at Boomtrain are responsible for designing and building the models and systems that provide... 

    Boomtrain

    San Francisco, CA
    2 days ago
  • $130k - $240k

     ...researchers, designers, growth experts, and engineers rethinking human-computer interaction...  ...well-known angels. About The Role As a ML engineer at Wispr, you’ll play a crucial...  ...features of our voice interface, building infrastructure to handle What are we looking for?... 

    DevExplore

    San Francisco, CA
    3 days ago
  •  ...RaindropRaindrop Raindrop is the monitoring platform for AI agents. Engineering teams at Fortune 100s and the fastest-growing AI companies (...  ...of requests a day + growing. Architect, implement, and scale ML pipelinesQuick iteration without compromising on qualityDeeply... 
    Temporary work

    Uncover

    San Francisco, CA
    2 days ago
  • $150k - $250k

    Garuda Ventures is looking for a Senior Software Engineer (IC) to develop cloud and on-prem systems that will power future factories...  ...software and a strong background in backend systems and infrastructure. The position offers a competitive salary range of $150,000—$... 

    Garuda Ventures

    San Francisco, CA
    2 days ago
  •  ...that every person shapes what gets built. About the Role As a ML engineer at Wispr, you’ll play a crucial role in building the first...  ...and designing new features of our voice interface, building infrastructure to handle What are we looking for? Previous founding or startup... 
    H1b
    Work at office
    Remote work
    Relocation
    Visa sponsorship
    Flexible hours

    Visa Hunt

    San Francisco, CA
    5 days ago
  • $128.7k - $261.3k

     ...approaches to model export, kernel development, and performance engineering so that every cycle on our accelerators translates into better...  ...tooling that makes that path fast, reliable, and effortless for ML engineers across the AV organization to compile their models.... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    3 days ago
  • $170k - $200k

     ...advantage in the robotics revolution. The Instawork Robotics ML Engineer will help build and scale the technology powering physical...  ...with expertise in distributed systems, cloud computing on AWS infrastructure, and scalable data processing for large-scale datasets.An... 
    Hourly pay
    Internship
    Local area
    Shift work

    Instawork

    San Francisco, CA
    2 days ago
  • $212k - $318k

     ...creator economy and are looking for a Senior Machine Learning Engineer, Infrastructure to support our mission.This role is based in San Francisco...  ...TeamYou'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    2 days per week
    3 days per week

    Patreon

    San Francisco, CA
    7 hours ago
  •  ...creating a genuinely open, collaborative path to frontier‑scale AI. We’re looking for an ML Training Platform Engineer to architect, build, and scale the foundational infrastructure powering our decentralised ML training platform. You will own core systems spanning... 
    Work experience placement

    Pluralis Research

    San Francisco, CA
    4 days ago
  • $197.3k - $225.1k

     ...Lead AI/ML Engineer (Platform, kubeflow) Overview At Capital One, we are creating responsible and reliable AI systems, changing...  ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    3 days ago
  • Faire is seeking a Staff Machine Learning Platform Engineer to design, improve, and operate a scalable ML platform that accelerates model training,...  ...Spark, Delta Lake, MLflow, Python and SQL, cloud/infrastructure-as-code, and strong MLOps practices. #J-18808-Ljbffr... 
    Remote job
    Local area

    Faire Inc

    San Francisco, CA
    4 days ago
  • Google Cloud is seeking a Customer Engineer in San Francisco to partner with technical sales as an AI/ML subject matter expert, helping customers and partners understand Google Cloud and develop creative cloud solutions. You will engage in proofs of concepts, present to... 

    Google Inc.

    San Francisco, CA
    3 days ago
  • $166k - $244k

    A leading technology company is seeking a Research Software Engineer to develop next-generation technologies that enhance communication through advanced software and hardware. This role involves collaborating with researchers, optimizing machine learning performance, and... 

    Google

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to ML Infrastructure Engineer. Be the first to apply!