AI Infrastructure Engineer
$350kThinking Machines Lab
About Thinking Machines The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it. About the Role We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines. You'll work side by side with research teams during active model runs - debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again. What You'll Do
Minimum Qualifications
- Own the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completion
- Partner directly with research teams during active model runs, embedding with them to unblock training and speed up iteration
- Debug failures across the full stack - accelerators, networking, storage, schedulers, and training frameworks - and drive issues to root cause
- Build monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stalling
- Improve checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of compute
- Build internal tools that reduce toil and improve cluster utilization across post-training and RL workloads
- Participate in an on-call rotation supporting production model runs
- Write postmortems and turn recurring failure patterns into permanent infrastructure fixes
Minimum Qualifications
- 4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in production
- Track record debugging complex failures across distributed systems - networking, hardware, kernel, or scheduler issues
- Strong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a system
- Solid grounding in Linux systems internals and networking fundamentals
- Comfortable owning production systems, including participating in on-call rotations
- Experience operating GPU or TPU training clusters at scale
- Familiarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloads
- Experience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)
- Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performance
- Experience building observability tooling purpose-built for ML training, not just general infrastructure
- A track record of thriving in fast-changing, research-driven environments where priorities shift with the science
- Location: This role is based in San Francisco, CA.
- Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
- Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in New York, NY vacancy
- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention for the next era of digital solutions powered... ...(BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang),...SuggestedFull timeWork experience placementLive inWork at officeLocal area
$215k - $350k
We are seeking an AI Infrastructure Engineer to build, operate, and continuously enhance the Linux and GPU-based infrastructure that powers our AI platforms and performance testing environments. This is a highly hands-on infrastructure, automation, and performance engineering...SuggestedWorldwideHome office- ...Palona’s AI agents operate continuously in production, handle real-time guest interactions... ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product:... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...SuggestedTemporary work
- ...AI Infrastructure SpecialistAs vCluster's AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical... ...role exists to make that happen.As an AI Infrastructure Engineer, your role will include:Lead Technical Deployments: Drive...SuggestedRemote workFlexible hours
$200k
...AI Infrastructure Engineer Location: New York (4 Days Onsite) Base Salary: $200k + 50% bonus This is a rare opportunity to shape the AI foundations of a complex, global organisation at a pivotal moment in its technology journey. You will play a central role...SuggestedWork at office$112.29k - $161.28k
...live and work, in ways they value.In an AI era, we remain ""Guided By Humanity,"" a... ...OverviewWe are looking for a Senior AI & Cloud Engineer to design, build, and operate... .... You will bridge the gap between cloud infrastructure, data engineering, and applied AI — developing...Temporary workFreelanceWork at officeLocal areaFlexible hours$120k - $240k
...About Build AI for the Built World: Build has created the agentic AI stack for institutional... ...most important built projects - digital infrastructure, energy, industrial - from concept to... ...create and configure workflows without engineering. The interfaces through which clients...Full timeLive in$180k - $225k
AI Infrastructure Engineer - Agent Sandbox PlatformAs a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform — the secure, high-performance code execution layer powering our agentic workflows, deployed across both...Full timeImmediate startRemote work- ...combine our strength in technology and leadership in cloud, data and AI with unmatched industry experience, functional expertise and... ...of deep industry knowledge and applied AI and data engineering. We help the world’s leading Resources and Utilities organizations...Full timeWork experience placementLive inWork at officeLocal area
$175k - $200k
...dedicated owner for OUTFRONT's internal AI platform. Today, our internal AI assistant... ...into production, and our external agent infrastructure (the Agency Connect MCP server, which... ...early partner adoption. Both need a senior engineering owner, and both need to grow into the...Full timeInternship$250.8k - $286.2k
Senior Lead AI Engineer (GenAI Platform Services, Agentic Platform) Overview: At Capital One, we are creating responsible and reliable... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area$200k - $230k
...and technology.Job DescriptionDirector, AI Platform EngineeringLocations: San Francisco... ...are seeking a Director of AI Platform Engineering to lead the design, development, and... ...platform engineers, architect critical infrastructure, and drive the strategy for multi-agent...Ongoing contractFull timeCasual workWork at officeFlexible hours$152.29k - $250.2k
As the Head of AI Platform Engineering - Execution Plane, you will lead the development and implementation of our enterprise platform’s execution... ...design reviews, architecture standards, testing, CI/CD, infrastructure automation, incident response, SLOs, runbooks, and...Full timeLocal areaVisa sponsorshipWork visaFlexible hours- ...through walls to get things done the right way, we want to build the future of wealth management with you.The RoleAs an AI-Native Data Platform Engineer at Farther, you will design and own the canonical data foundations powering our financial AI systems. This role sits...
$230k - $290k
...York( Digital Solutions Group ) - Platform Engineering /Full Time /RemoteAHEAD helps large... ...same enterprises design, build, and run AI agent platforms on top of it. This role... ...delivery leadership attached. You will write infrastructure code, agent code, and the documents that...Full timeWork at officeShift work- The OpportunityJoin a team building the data foundations that support the firm’s AI and analytics capabilities. This role sits within the engineering effort to develop a modern Lakehouse and AI data platform that enables reliable, well-governed and high-performing data...
$269.1k - $307.2k
...Distinguished AI Engineer - Agentic AI Platform (Remote Eligible) At Capital One, we are creating responsible and reliable AI systems... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeWork at officeRemote work- ...The Opportunity MassMutual's AI Platform Engineering team is seeking an impact-driven AI Platform Engineer to serve as the technical anchor... .... The team operates at the intersection of cloud infrastructure, AI/ML systems, and developer experience—delivering foundational...
$205k - $235k
...With deep functional and sector expertise, paired with innovative AI-powered technology and an investor mindset, we partner with... ...Within the EY-Parthenon service line, the EY Growth Platforms AI ML Engineering Director will collaborate with Business Leaders, Data...Full timeFor contractorsWork experience placementSummer holidayFlexible hours$229.9k - $262.4k
...Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time,...Full timePart timeLocal area$250.8k - $286.2k
...Senior Lead AI Engineer (GenAI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine...Full timePart timeLocal area$152k - $241.5k
...weight models are foundational to American AI leadership and cybersecurity, and that... ...scrutiny. Our AI Safety & Security Engineering team builds and evaluates AI-powered tooling... ...and maintain the agent harness.Evaluation infrastructure: Build the systems we use to run and...Full timeRemote work- Google Cloud is seeking a Customer Engineer specializing in Cloud AI to accelerate technical wins and adoption of complex workloads. You will partner with technical sales, write code for prototypes, and deliver demos that showcase new AI solutions to customers. You will...
- United States Digital Space LLC is seeking an experienced AI/ML architect for a remote engagement supporting a Fortune 100 pharmaceutical... ...citizenship or Green Card and a Google Cloud Professional Data Engineer certification. You will leverage Gemini Enterprise, Vertex AI...Remote job
$60 - $70 per hour
Senior Technical Recruiter - AI Infrastructure and Engineering Remote, US $60 - $70 Job Title Senior Technical Recruiter - AI Infrastructure and Engineering Location Remote, US Salary Range $60 - $70 Job Description About the Mission We support companies building large-...Remote work- Crusoe Cloud is seeking a Solutions Engineer to help enterprise customers deploy AI/ML workloads on Crusoe's GPU infrastructure. Based in New York City, you will run technical discovery, deliver demos and PoCs, and ensure customers land successfully on the platform. This...Full time
$120k - $150k
...Job Description Job Description Lead Cloud AI Engineer (Lead / Senior) Location: Various U.S. Federal Client Sites (Onsite)... ...provider of enterprise IT modernization, cloud transformation, and infrastructure solutions supporting U.S. Federal Government agencies. As we...Full timeTemporary workImmediate start- As an AI Platform Engineer for AI & Emerging Tech, you will drive AI platform enablement across the enterprise. This role sits at the intersection... ...guardrails (budgets, alerts, quota enforcement) using Infrastructure as Code. Act as a subject...
- SOCOTEC is seeking a Software Engineer to join our US engineering team and contribute to our AI platform and product portfolio. This full-stack role involves modern... ...React frontends, Python backends, and data infrastructure. You will ship features end-to-end, own components...
- Alden Labs in New York City is seeking a founding engineer to help shape the future of healthcare AI. This in-person, full-time role invites you to own core... ...integrations, intelligent workflows, and voice AI infrastructure. You will move fast, wear many hats across...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!
Related searches
- ai research engineer New York, NY
- machine learning ai engineer New York, NY
- ai developer New York, NY
- senior ai engineer New York, NY
- ai engineer New York, NY
- ai ml engineer New York, NY
- ai prompt engineer New York, NY
- ai engineer remote New York, NY
- data infrastructure engineer New York, NY
- infrastructure engineering manager New York, NY


