Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

$350k

Thinking Machines Lab

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

You'll work side by side with research teams during active model runs - debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.

What You'll Do
  • Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completion
  • Partner directly with research teams during active model runs, embedding with them to unblock training and speed up iteration
  • Debug failures across the full stack - accelerators, networking, storage, schedulers, and training frameworks - and drive issues to root cause
  • Build monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stalling
  • Improve checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of compute
  • Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads
  • Participate in an on-call rotation supporting production model runs
  • Write postmortems and turn recurring failure patterns into permanent infrastructure fixes
Skills & Qualifications
Minimum Qualifications
  • 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production
  • Track record debugging complex failures across distributed systems - networking, hardware, kernel, or scheduler issues
  • Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system
  • Solid grounding in Linux systems internals and networking fundamentals
  • Comfortable owning production systems, including participating in on-call rotations
Preferred Qualifications
  • Experience operating GPU or TPU training clusters at scale
  • Familiarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloads
  • Experience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)
  • Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performance
  • Experience building observability tooling purpose-built for ML training, not just general infrastructure
  • A track record of thriving in fast-changing, research-driven environments where priorities shift with the science
Logistics
  • Location: This role is based in San Francisco, CA.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in New York, NY vacancy
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered...  ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    New York, NY
    5 days ago
  • $215k - $350k

    We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering... 
    Suggested
    Worldwide
    Home office

    Fortinet

    New York, NY
    1 day ago
  •  ...Palona’s AI agents operate continuously in production, handle real-time guest interactions...  ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product:...  ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth... 
    Suggested
    Temporary work

    Palona AI

    New York, NY
    5 days ago
  •  ...AI Infrastructure SpecialistAs vCluster's AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical...  ...role exists to make that happen.As an AI Infrastructure Engineer, your role will include:Lead Technical Deployments: Drive... 
    Suggested
    Remote work
    Flexible hours

    vCluster

    New York, NY
    4 days ago
  • $200k

     ...AI Infrastructure Engineer Location: New York (4 Days Onsite) Base Salary: $200k + 50% bonus This is a rare opportunity to shape the AI foundations of a complex, global organisation at a pivotal moment in its technology journey. You will play a central role... 
    Suggested
    Work at office

    Harnham

    New York, NY
    4 days ago
  • $112.29k - $161.28k

     ...live and work, in ways they value.In an AI era, we remain ""Guided By Humanity,"" a...  ...OverviewWe are looking for a Senior AI & Cloud Engineer to design, build, and operate...  .... You will bridge the gap between cloud infrastructure, data engineering, and applied AI — developing... 
    Temporary work
    Freelance
    Work at office
    Local area
    Flexible hours

    Publicis Media

    New York, NY
    4 days ago
  • $120k - $240k

     ...About Build AI for the Built World: Build has created the agentic AI stack for institutional...  ...most important built projects - digital infrastructure, energy, industrial - from concept to...  ...create and configure workflows without engineering. The interfaces through which clients... 
    Full time
    Live in

    Build Technologies

    New York, NY
    3 days ago
  • $180k - $225k

    AI Infrastructure Engineer - Agent Sandbox PlatformAs a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both... 
    Full time
    Immediate start
    Remote work

    Scale AI

    New York, NY
    2 days ago
  •  ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and...  ...of deep industry knowledge and applied AI and data engineering. We help the world’s leading Resources and Utilities organizations... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    New York, NY
    3 days ago
  • $175k - $200k

     ...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant...  ...into production, and our external agent infrastructure (the Agency Connect MCP server, which...  ...early partner adoption. Both need a senior engineering owner, and both need to grow into the... 
    Full time
    Internship

    OUTFRONT Media

    New York, NY
    2 days ago
  • $250.8k - $286.2k

    Senior Lead AI Engineer (GenAI Platform Services, Agentic Platform) Overview: At Capital One, we are creating responsible and reliable...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    5 days ago
  • $200k - $230k

     ...and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco...  ...are seeking a Director of AI Platform Engineering to lead the design, development, and...  ...platform engineers, architect critical infrastructure, and drive the strategy for multi-agent... 
    Ongoing contract
    Full time
    Casual work
    Work at office
    Flexible hours

    SS&C Technologies

    New York, NY
    1 day ago
  • $152.29k - $250.2k

    As the Head of AI Platform Engineering - Execution Plane, you will lead the development and implementation of our enterprise platform’s execution...  ...design reviews, architecture standards, testing, CI/CD, infrastructure automation, incident response, SLOs, runbooks, and... 
    Full time
    Local area
    Visa sponsorship
    Work visa
    Flexible hours

    Guardian Life Insurance

    New York, NY
    4 days ago
  •  ...through walls to get things done the right way, we want to build the future of wealth management with you.The RoleAs an AI-Native Data Platform Engineer at Farther, you will design and own the canonical data foundations powering our financial AI systems. This role sits... 

    Farther Advisors

    New York, NY
    4 days ago
  • $230k - $290k

     ...York( Digital Solutions Group ) - Platform Engineering /Full Time /RemoteAHEAD helps large...  ...same enterprises design, build, and run AI agent platforms on top of it. This role...  ...delivery leadership attached. You will write infrastructure code, agent code, and the documents that... 
    Full time
    Work at office
    Shift work

    AHEAD

    New York, NY
    2 days ago
  • The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data... 

    Goldman Sachs

    New York, NY
    4 days ago
  • $269.1k - $307.2k

     ...Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Full time
    Part time
    Work at office
    Remote work

    Capital One

    New York, NY
    3 days ago
  •  ...The Opportunity MassMutual's AI Platform Engineering team is seeking an impact-driven AI Platform Engineer to serve as the technical anchor...  .... The team operates at the intersection of cloud infrastructure, AI/ML systems, and developer experience—delivering foundational... 

    MassMutual

    New York, NY
    2 days ago
  • $205k - $235k

     ...With deep functional and sector expertise, paired with innovative AI-powered technology and an investor mindset, we partner with...  ...Within the EY-Parthenon service line, the EY Growth Platforms AI ML Engineering Director will collaborate with Business Leaders, Data... 
    Full time
    For contractors
    Work experience placement
    Summer holiday
    Flexible hours

    EY (Ernst & Young)

    New York, NY
    3 days ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time,... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  • $250.8k - $286.2k

     ...Senior Lead AI Engineer (GenAI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  • $152k - $241.5k

     ...weight models are foundational to American AI leadership and cybersecurity, and that...  ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling...  ...and maintain the agent harness.Evaluation infrastructure: Build the systems we use to run and... 
    Full time
    Remote work

    Nvidia

    New York, NY
    4 days ago
  • Google Cloud is seeking a Customer Engineer specializing in Cloud AI to accelerate technical wins and adoption of complex workloads. You will partner with technical sales, write code for prototypes, and deliver demos that showcase new AI solutions to customers. You will... 

    Google

    New York, NY
    1 day ago
  • United States Digital Space LLC is seeking an experienced AI/ML architect for a remote engagement supporting a Fortune 100 pharmaceutical...  ...citizenship or Green Card and a Google Cloud Professional Data Engineer certification. You will leverage Gemini Enterprise, Vertex AI... 
    Remote job

    United States Digital Space LLC

    New York, NY
    3 days ago
  • $60 - $70 per hour

    Senior Technical Recruiter - AI Infrastructure and Engineering Remote, US $60 - $70 Job Title Senior Technical Recruiter - AI Infrastructure and Engineering Location Remote, US Salary Range $60 - $70 Job Description About the Mission We support companies building large-... 
    Remote work

    The Leadership Agency Inc.

    New York, NY
    3 days ago
  • Crusoe Cloud is seeking a Solutions Engineer to help enterprise customers deploy AI/ML workloads on Crusoe's GPU infrastructure. Based in New York City, you will run technical discovery, deliver demos and PoCs, and ensure customers land successfully on the platform. This... 
    Full time

    Linuxconfig

    New York, NY
    5 days ago
  • $120k - $150k

     ...Job Description Job Description Lead Cloud AI Engineer (Lead / Senior) Location: Various U.S. Federal Client Sites (Onsite)...  ...provider of enterprise IT modernization, cloud transformation, and infrastructure solutions supporting U.S. Federal Government agencies. As we... 
    Full time
    Temporary work
    Immediate start

    PGTEK

    New York, NY
    4 days ago
  • As an AI Platform Engineer for AI & Emerging Tech, you will drive AI platform enablement across the enterprise. This role sits at the intersection...  ...guardrails (budgets, alerts, quota enforcement) using Infrastructure as Code. Act as a subject... 

    Luxoft

    New York, NY
    a month ago
  • SOCOTEC is seeking a Software Engineer to join our US engineering team and contribute to our AI platform and product portfolio. This full-stack role involves modern...  ...React frontends, Python backends, and data infrastructure. You will ship features end-to-end, own components... 

    SOCOTEC

    New York, NY
    1 day ago
  • Alden Labs in New York City is seeking a founding engineer to help shape the future of healthcare AI. This in-person, full-time role invites you to own core...  ...integrations, intelligent workflows, and voice AI infrastructure. You will move fast, wear many hats across... 
    Full time

    Alden Labs, Inc.

    New York, NY
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!