AI Infrastructure Engineer
$350kThinking Machines Lab
AI Infrastructure Engineer
The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.
We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.
You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.
What You'll Do
- Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completion
- Partner directly with research teams during active model runs, embedding with them to unblock training and speed up iteration
- Debug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root cause
- Build monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stalling
- Improve checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of compute
- Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads
- Participate in an on-call rotation supporting production model runs
- Write postmortems and turn recurring failure patterns into permanent infrastructure fixes
Minimum Qualifications
- 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production
- Track record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issues
- Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system
- Solid grounding in Linux systems internals and networking fundamentals
- Comfortable owning production systems, including participating in on-call rotations
Preferred Qualifications
- Experience operating GPU or TPU training clusters at scale
- Familiarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloads
- Experience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)
- Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performance
- Experience building observability tooling purpose-built for ML training, not just general infrastructure
- A track record of thriving in fast-changing, research-driven environments where priorities shift with the science
Logistics
- Location: This role is based in San Francisco, CA.
- Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
- Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered... ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...SuggestedFull timeWork experience placementLive inWork at officeLocal area
$175k - $275k
...About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise—already trusted by some of the largest companies in... ...would be possible. The Role As an AI Engineer - Cloud Infrastructure on Traversal’s Infrastructure team, you’ll design,...SuggestedFull timeWork at officeFlexible hours$105k - $115k
...Description Cloud AI Engineer (Mid-Level) Location: Various U.S. Federal Client Sites (Onsite) Travel: Travel required... ...trusted provider of enterprise cloud modernization, AI, and infrastructure solutions supporting U.S. Federal Government agencies. As we...SuggestedFull timeTemporary workImmediate start$160k - $220k
...of: The role We are looking for an experienced AI Engineer to lead the implementation of Azure AI Foundry within an established... ...our existing data platform, governance model, and analytics infrastructure. You will work closely with data engineering,...SuggestedFull timeRemote work$165k - $200k
...can be part of the disciplined collaboration and transcendent thinking as our AI Platform Engineer at Capstone Investment Advisors here. Overview: We are seeking an AI Infrastructure Engineer to design, build, and scale the foundational infrastructure that enables...SuggestedMinimum wage$200k
...AI Infrastructure Engineer Location: New York (4 Days Onsite) Base Salary: $200k + 50% bonus This is a rare opportunity to shape the AI foundations of a complex, global organisation at a pivotal moment in its technology journey. You will play a central role...Work at office- ...Palona’s AI agents operate continuously in production, handle real-time guest interactions... ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product:... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...Temporary work
$229.9k - $262.4k
Senior Lead AI Engineer, Gen AI Platform Overview: At Capital One, we are creating responsible and reliable... ...personalized customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in...Full timePart timeLocal area$197.3k - $225.1k
{"description": "Lead AI Engineer ( MLX, Gen AI Platform Services, Agentic AI) Overview At Capital One, we are... ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine...Full timePart timeLocal area$180k - $225k
AI Infrastructure Engineer - Agent Sandbox PlatformAs a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both...Full timeImmediate startRemote work$152.29k - $250.2k
As the Head of AI Platform Engineering - Execution Plane, you will lead the development and implementation of our enterprise platform’s execution... ...design reviews, architecture standards, testing, CI/CD, infrastructure automation, incident response, SLOs, runbooks, and...Full timeLocal areaVisa sponsorshipWork visaFlexible hours- ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and... ...of deep industry knowledge and applied AI and data engineering. We help the world’s leading Resources and Utilities organizations...Full timeWork experience placementLive inWork at officeLocal area
$200k - $230k
...and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco... ...are seeking a Director of AI Platform Engineering to lead the design, development, and... ...platform engineers, architect critical infrastructure, and drive the strategy for multi-agent...Ongoing contractFull timeCasual workWork at officeFlexible hours$175k - $200k
...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant... ...into production, and our external agent infrastructure (the Agency Connect MCP server, which... ...early partner adoption. Both need a senior engineering owner, and both need to grow into the...Full timeInternship- ...through walls to get things done the right way, we want to build the future of wealth management with you.The RoleAs an AI-Native Data Platform Engineer at Farther, you will design and own the canonical data foundations powering our financial AI systems. This role sits...
- The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data...
$269.1k - $307.2k
...Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeWork at officeRemote work- ...The OpportunityMassMutual’s AI Platform Engineering team is seeking an impact-driven AI Platform Engineer to serve as the technical anchor... ...initiatives. The team operates at the intersection of cloud infrastructure, AI/ML systems, and developer experience—delivering...
- ...Job Title: AI Platform Engineer Location: NYC, NY (Hybrid Model) - 10003 Zip code Energy & Utility domain with experience in Google... ...Platform Engineer: Design and implement AI/ML infrastructure using Vertex AI, Kubeflow, TensorFlow Extended (TFX), and...Local area
$205k - $235k
...With deep functional and sector expertise, paired with innovative AI-powered technology and an investor mindset, we partner with... ...Within the EY-Parthenon service line, the EY Growth Platforms AI ML Engineering Director will collaborate with Business Leaders, Data...Full timeFor contractorsWork experience placementSummer holidayFlexible hours$229.9k - $262.4k
...Sr. Lead AI Engineer (GenAI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area- ...Head of AI Platform Engineering, Execution Services About the Company Enterprise-focused organization building secure, scalable AI... ...background in distributed systems, cloud-native platforms, AI/ML infrastructure, and large-scale data ecosystems, as well as experience...
$229.9k - $262.4k
...Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time,...Full timePart timeLocal area$60 - $70 per hour
Senior Technical Recruiter - AI Infrastructure and Engineering Remote, US $60 - $70 Job Title Senior Technical Recruiter - AI Infrastructure and Engineering Location Remote, US Salary Range $60 - $70 Job Description About the Mission We support companies building large-...Remote work$152k - $241.5k
...weight models are foundational to American AI leadership and cybersecurity, and that... ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling... ...and maintain the agent harness.Evaluation infrastructure: Build the systems we use to run and...Full timeRemote work- Google Cloud is seeking a Customer Engineer specializing in Cloud AI to accelerate technical wins and adoption of complex workloads. You will partner with technical sales, write code for prototypes, and deliver demos that showcase new AI solutions to customers. You will...
- United States Digital Space LLC is seeking an experienced AI/ML architect for a remote engagement supporting a Fortune 100 pharmaceutical... ...citizenship or Green Card and a Google Cloud Professional Data Engineer certification. You will leverage Gemini Enterprise, Vertex AI...Remote job
- Crusoe Cloud is seeking a Solutions Engineer to help enterprise customers deploy AI/ML workloads on Crusoe's GPU infrastructure. Based in New York City, you will run technical discovery, deliver demos and PoCs, and ensure customers land successfully on the platform. This...Full time
$229.9k - $286.2k
...Requirements: ~ Bachelors degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields along with a minimum... ...as an industry leader. With our advanced technology infrastructure and a talented team, we strive to continue enhancing our...Full time- As an AI Platform Engineer for AI & Emerging Tech, you will drive AI platform enablement across the enterprise. This role sits at the intersection... ...guardrails (budgets, alerts, quota enforcement) using Infrastructure as Code. Act as a subject...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!
- ai research engineer New York, NY
- machine learning ai engineer New York, NY
- ai developer New York, NY
- senior ai engineer New York, NY
- ai engineer New York, NY
- ai ml engineer New York, NY
- ai prompt engineer New York, NY
- ai engineer remote New York, NY
- data infrastructure engineer New York, NY
- infrastructure engineering manager New York, NY



