Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

$350k

Thinking Machines Lab

AI Infrastructure Engineer

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.

You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.

What You'll Do
  • Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completion
  • Partner directly with research teams during active model runs, embedding with them to unblock training and speed up iteration
  • Debug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root cause
  • Build monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stalling
  • Improve checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of compute
  • Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads
  • Participate in an on-call rotation supporting production model runs
  • Write postmortems and turn recurring failure patterns into permanent infrastructure fixes
Minimum Qualifications
  • 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production
  • Track record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issues
  • Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system
  • Solid grounding in Linux systems internals and networking fundamentals
  • Comfortable owning production systems, including participating in on-call rotations
Preferred Qualifications
  • Experience operating GPU or TPU training clusters at scale
  • Familiarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloads
  • Experience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)
  • Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performance
  • Experience building observability tooling purpose-built for ML training, not just general infrastructure
  • A track record of thriving in fast-changing, research-driven environments where priorities shift with the science
Logistics
  • Location: This role is based in San Francisco, CA.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in New York, NY vacancy
  • We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered...  ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    New York, NY
    1 day ago
  • $175k - $275k

     ...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in...  ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Traversal

    New York, NY
    17 hours ago
  • $105k - $115k

     ...Description Cloud AI Engineer (Mid-Level) Location: Various U.S. Federal Client Sites (Onsite) Travel: Travel required...  ...trusted provider of enterprise cloud modernization, AI, and infrastructure solutions supporting U.S. Federal Government agencies. As we... 
    Suggested
    Full time
    Temporary work
    Immediate start

    Pgtek

    New York, NY
    17 hours ago
  • $160k - $220k

     ...of:     The role   We are looking for an experienced AI Engineer to lead the implementation of Azure AI Foundry within an established...  ...our existing data platform, governance model, and analytics infrastructure.   You will work closely with data engineering,... 
    Suggested
    Full time
    Remote work

    Valtech Se

    New York, NY
    17 hours ago
  • $165k - $200k

     ...can be part of the disciplined collaboration and transcendent thinking as our AI Platform Engineer at Capstone Investment Advisors here. Overview: We are seeking an AI Infrastructure Engineer to design, build, and scale the foundational infrastructure that enables... 
    Suggested
    Minimum wage

    Capstone Investment Advisors

    New York, NY
    3 days ago
  • $200k

     ...AI Infrastructure Engineer Location: New York (4 Days Onsite) Base Salary: $200k + 50% bonus This is a rare opportunity to shape the AI foundations of a complex, global organisation at a pivotal moment in its technology journey. You will play a central role... 
    Work at office

    Harnham

    New York, NY
    1 day ago
  •  ...Palona’s AI agents operate continuously in production, handle real-time guest interactions...  ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product:...  ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth... 
    Temporary work

    Palona AI

    New York, NY
    1 day ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer, Gen AI Platform Overview: At Capital One, we are creating responsible and reliable...  ...personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    New York, NY
    17 hours ago
  • $197.3k - $225.1k

    {"description": "Lead AI Engineer ( MLX, Gen AI Platform Services, Agentic AI) Overview At Capital One, we are...  ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    New York, NY
    17 hours ago
  • $180k - $225k

    AI Infrastructure Engineer - Agent Sandbox PlatformAs a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both... 
    Full time
    Immediate start
    Remote work

    Scale AI

    New York, NY
    3 days ago
  • $152.29k - $250.2k

    As the Head of AI Platform Engineering - Execution Plane, you will lead the development and implementation of our enterprise platform’s execution...  ...design reviews, architecture standards, testing, CI/CD, infrastructure automation, incident response, SLOs, runbooks, and... 
    Full time
    Local area
    Visa sponsorship
    Work visa
    Flexible hours

    Guardian Life Insurance

    New York, NY
    17 hours ago
  •  ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and...  ...of deep industry knowledge and applied AI and data engineering. We help the world’s leading Resources and Utilities organizations... 
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    New York, NY
    4 days ago
  • $200k - $230k

     ...and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco...  ...are seeking a Director of AI Platform Engineering to lead the design, development, and...  ...platform engineers, architect critical infrastructure, and drive the strategy for multi-agent... 
    Ongoing contract
    Full time
    Casual work
    Work at office
    Flexible hours

    SS&C Technologies

    New York, NY
    2 days ago
  • $175k - $200k

     ...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant...  ...into production, and our external agent infrastructure (the Agency Connect MCP server, which...  ...early partner adoption. Both need a senior engineering owner, and both need to grow into the... 
    Full time
    Internship

    OUTFRONT Media

    New York, NY
    3 days ago
  •  ...through walls to get things done the right way, we want to build the future of wealth management with you.The RoleAs an AI-Native Data Platform Engineer at Farther, you will design and own the canonical data foundations powering our financial AI systems. This role sits... 

    Farther Advisors

    New York, NY
    17 hours ago
  • The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data... 

    Goldman Sachs

    New York, NY
    17 hours ago
  • $269.1k - $307.2k

     ...Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Full time
    Part time
    Work at office
    Remote work

    Capital One

    New York, NY
    7 hours ago
  •  ...The OpportunityMassMutual’s AI Platform Engineering team is seeking an impact-driven AI Platform Engineer to serve as the technical anchor...  ...initiatives. The team operates at the intersection of cloud infrastructure, AI/ML systems, and developer experience—delivering... 

    MassMutual

    New York, NY
    3 days ago
  •  ...Job Title: AI Platform Engineer Location: NYC, NY (Hybrid Model) - 10003 Zip code Energy & Utility domain with experience in Google...  ...Platform Engineer: Design and implement AI/ML infrastructure using Vertex AI, Kubeflow, TensorFlow Extended (TFX), and... 
    Local area

    Abode Tech Zone

    New York, NY
    3 days ago
  • $205k - $235k

     ...With deep functional and sector expertise, paired with innovative AI-powered technology and an investor mindset, we partner with...  ...Within the EY-Parthenon service line, the EY Growth Platforms AI ML Engineering Director will collaborate with Business Leaders, Data... 
    Full time
    For contractors
    Work experience placement
    Summer holiday
    Flexible hours

    EY (Ernst & Young)

    New York, NY
    4 days ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (GenAI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    2 days ago
  •  ...Head of AI Platform Engineering, Execution Services About the Company Enterprise-focused organization building secure, scalable AI...  ...background in distributed systems, cloud-native platforms, AI/ML infrastructure, and large-scale data ecosystems, as well as experience... 

    Confidential

    New York, NY
    1 day ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time,... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    17 hours ago
  • $60 - $70 per hour

    Senior Technical Recruiter - AI Infrastructure and Engineering Remote, US $60 - $70 Job Title Senior Technical Recruiter - AI Infrastructure and Engineering Location Remote, US Salary Range $60 - $70 Job Description About the Mission We support companies building large-... 
    Remote work

    The Leadership Agency Inc.

    New York, NY
    4 days ago
  • $152k - $241.5k

     ...weight models are foundational to American AI leadership and cybersecurity, and that...  ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling...  ...and maintain the agent harness.Evaluation infrastructure: Build the systems we use to run and... 
    Full time
    Remote work

    Nvidia

    New York, NY
    17 hours ago
  • Google Cloud is seeking a Customer Engineer specializing in Cloud AI to accelerate technical wins and adoption of complex workloads. You will partner with technical sales, write code for prototypes, and deliver demos that showcase new AI solutions to customers. You will... 

    Google

    New York, NY
    2 days ago
  • United States Digital Space LLC is seeking an experienced AI/ML architect for a remote engagement supporting a Fortune 100 pharmaceutical...  ...citizenship or Green Card and a Google Cloud Professional Data Engineer certification. You will leverage Gemini Enterprise, Vertex AI... 
    Remote job

    United States Digital Space LLC

    New York, NY
    4 days ago
  • Crusoe Cloud is seeking a Solutions Engineer to help enterprise customers deploy AI/ML workloads on Crusoe's GPU infrastructure. Based in New York City, you will run technical discovery, deliver demos and PoCs, and ensure customers land successfully on the platform. This... 
    Full time

    Linuxconfig

    New York, NY
    1 day ago
  • $229.9k - $286.2k

     ...Requirements: ~ Bachelors degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields along with a minimum...  ...as an industry leader. With our advanced technology infrastructure and a talented team, we strive to continue enhancing our... 
    Full time

    Capital One

    New York, NY
    18 days ago
  • As an AI Platform Engineer for AI & Emerging Tech, you will drive AI platform enablement across the enterprise. This role sits at the intersection...  ...guardrails (budgets, alerts, quota enforcement) using Infrastructure as Code. Act as a subject... 

    Luxoft

    New York, NY
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!