ML Infrastructure Engineer
$350kThinking Machines Lab
AI Infrastructure EngineerThe mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.We're hiring an AI Infrastructure Engineer to keep our post-training and reinforcement learning (RL) systems fast, reliable, and easy for researchers to iterate on. Think of this as a production engineering or site reliability role built around model training: you'll own the health of the training runs, clusters, and pipelines that power post-training and RL at Thinking Machines.You'll work side by side with research teams during active model runs — debugging failures in real time, hardening infrastructure against the next class of problem, and building the tooling and automation that let researchers spend their time on the science instead of babysitting jobs. This role has real ownership: you'll be the person a research team calls when a run stalls at 2am, and the person who makes sure it doesn't happen again.What You'll DoOwn the reliability, performance, and uptime of large-scale post-training and RL training jobs, from launch through completionPartner directly with research teams during active model runs, embedding with them to unblock training and speed up iterationDebug failures across the full stack — accelerators, networking, storage, schedulers, and training frameworks — and drive issues to root causeBuild monitoring, alerting, and automated recovery so runs self-heal or fail fast instead of silently stallingImprove checkpointing, fault tolerance, and job scheduling so hardware failures cost minutes, not days of computeBuild internal tools that reduce toil and improve cluster utilization across post-training and RL workloadsParticipate in an on-call rotation supporting production model runsWrite postmortems and turn recurring failure patterns into permanent infrastructure fixesMinimum Qualifications4+ years of experience as a production engineer, site reliability engineer, or infrastructure engineer operating large-scale distributed systems in productionTrack record debugging complex failures across distributed systems — networking, hardware, kernel, or scheduler issuesStrong software engineering skills in Python and/or Go/C++, with the judgment to know when to script a fix versus build a systemSolid grounding in Linux systems internals and networking fundamentalsComfortable owning production systems, including participating in on-call rotationsPreferred QualificationsExperience operating GPU or TPU training clusters at scaleFamiliarity with post-training and RL techniques (e.g., RLHF, PPO, DPO) and the infrastructure challenges specific to them, such as reward model serving, rollout generation, and mixed training/inference workloadsExperience with distributed training frameworks (e.g., PyTorch, Ray) and job schedulers (e.g., Slurm, Kubernetes)Experience with high-performance networking (e.g., InfiniBand, RDMA, NCCL) and its role in distributed training performanceExperience building observability tooling purpose-built for ML training, not just general infrastructureA track record of thriving in fast-changing, research-driven environments where priorities shift with the scienceLogisticsLocation: This role is based in San Francisco, CA.Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
$184.9k - $250.2k
We are seeking a Machine Learning Engineer to work directly alongside our research scientists... ...in the real world. This is a hands-on ML role: you will train policies, debug... ...need to think seriously about training infrastructure — managing GPU clusters, optimizing distributed...SuggestedInternshipFlexible hours$160k - $240k
Senior ML Platform Engineer - Artificial Intelligence Location New York Business Area Engineering and CTO Ref # 10050078 Description & Requirements Bloomberg’s Engineering AI department has 400+ AI practitioners building highly sought after products...SuggestedTemporary workFor contractorsWork experience placement$180k - $300k
..., our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly... ...pipelines to deliver reliable end-to-end machine learning (ML) workflowsCollaborate with ML researchers and engineers to...SuggestedWork experience placement$205k - $235k
...of the way - enabling you to shape your future with confidence.Within the EY-Parthenon service line, the EY Growth Platforms AI ML Engineering Director will collaborate with Business Leaders, Data Scientists, Project Managers, and other team members to design, build,...SuggestedFull timeFor contractorsWork experience placementSummer holidayFlexible hours- Triangle Analytics seeks an ML systems engineer to build and scale machine learning components powering our core platform. You will turn... ...reliable predictions at scale, with ownership over modeling infrastructure. You will work across applied ML research and production...Suggested
$250k - $300k
A cloud financial management company is seeking a Mid-Senior Software Engineer to build scalable systems for cloud expense management. Responsibilities include developing backend services, addressing complex data problems, and collaborating with a team to enhance the platform...- Jobtailor is seeking an experienced platform/DevOps engineer to design and maintain infrastructure for ML training, deployment, and inference. You will own CI/CD, IaC, observability, and operational tooling to support reliable, scalable ML workloads. Responsibilities include...
- Career Techniques in New York is seeking an experienced ML researcher to expand our distributed ML stack, benchmark workloads, and stress-test both software and hardware across the platform. You will build high-level abstractions, integrate tooling with data frameworks...
- Goldman Sachs Bank AG in New York is seeking a Vice President in Software Engineering. The successful candidate will drive machine learning projects and develop scalable models. With over 10 years of experience, a degree in Computer Science, and expertise in Python and...
- Jane Street is seeking an engineer with strong machine learning foundations to advance our ML platform and research workflows. You will contribute to API design... ...and experience building training and inference infrastructure, with familiarity across PyTorch, Jax, or...
$197.3k - $225.1k
...Lead AI/ML Engineer (Platform, kubeflow) Overview At Capital One, we are creating responsible and reliable AI systems, changing... ...customer experiences. Our investments in technology infrastructure and world-class talent - along with our deep experience in machine...Full timePart timeLocal area- Snap Inc. is seeking a Machine Learning Engineer to join our Generative ML team in New York. You will develop innovative ML technology, work on cutting-edge generative pipelines for image, video, language or audio generation, and deliver on-device experiences. You will...
- The Goldman Sachs Group is seeking an experienced Machine Learning Engineer in New York to develop and deploy advanced ML models in a dynamic environment. Candidates should have over 10 years of experience in scalable ML systems and strong coding skills. This role involves...
$292k - $417.2k
...ads optimization that shape the future of streaming.We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design. In this role, you will own the strategic direction...Full timeTemporary workLocal areaFlexible hours$180k - $230k
...Arlo's underwriting is the core of the business, and it runs on machine learning at serious scale. We're hiring a Senior ML Infrastructure Engineer to build and own the infrastructure that powers it — from training models on tens of millions of patients and hundreds of...$130k - $147k
...United States. Job Description We are seeking a Senior ML/AI Platform Engineer to help build and operate Curinos' Databricks-native AI... ...will be responsible for designing and implementing the infrastructure, tooling, automation, observability, and governance capabilities...Part timeWork at officeRemote workWork from homeFlexible hours- Brain Co. seeks a Machine Learning Engineer for its Platform team to build core ML capabilities that power multiple products across regulated institutions. You will focus on turning pod needs into reusable platform features, and ensure production-grade quality across deployments...
- ...in the enterprise through purpose-built ML systems that learn from Rippling's proprietary... ...at Rippling.As a Staff Machine Learning Engineer, you will own the end-to-end ML... ...Build robust evaluation and experimentation infrastructure: offline benchmarks, A/B testing, and...Work at office3 days per week
$212k - $318k
...creator economy and are looking for a Senior Machine Learning Engineer, Infrastructure to support our mission.This role is based in San Francisco... ...TeamYou'll join the Relevance team, whose mission is to build the ML systems that power how fans discover creators and how...Work at officeLocal areaRemote workWorldwideFlexible hours2 days per week3 days per week$170k - $220k
...more than 100,000 clients nationwide. Our ML and AI capabilities are expanding rapidly—... ...AI products, and developer tooling—and the infrastructure underneath needs to scale with them. As our first dedicated ML Platform Engineer, you'll define the technical direction and...Full timeWork at officeLocal area$10k
...help define and contribute to a growing ML platform with the goal of accelerating the... ...like data science and backend engineering. You will meet our internal customers where... ...GenAI Platform and underlying operational infrastructure to enable Product impact. Develop, maintain...Work experience placement$229.9k - $286.2k
...Masters or doctoral degree in computer science, electrical engineering, mathematics, or a related field (preferred) ~ Experience with... ...teams Make informed decisions regarding machine learning infrastructure based on understanding of modeling techniques, including...Full timeInternship$77k - $202k
...of impactful solutions in a consulting settingWhat You Must Have- Bachelor's Degree- At least 4 years of experience in software engineering or data engineeringWhat Sets You Apart- Master's Degree in Computer Science, Data Engineering, Software Engineering preferred- Experience...Full timeH1b- ...A leading AI solutions provider in New York is seeking a talented ML Engineer to design, build, and deploy innovative ML systems. The role involves collaborating with cross-functional teams and engaging directly with customers to deliver impactful AI solutions. Candidates...
- ...Design, build, and deploy production‑grade ML systems with end‑to‑end ownership of the... ...in continuous improvement of the ML infrastructure and processes for scalability and performance... ...years of professional experience in ML engineering. Strong programming skills in Python (...Full time
$140k - $155k
...intelligence, and software-driven Field Engineering that drives real transformation on-site... ...across customer sites Work closely with ML engineers to accelerate the feedback loop between deployed models and training infrastructure Contribute to custom Visia implementations...Work experience placementFlexible hours- ...Insights 2023) focused on responsible development of large language models (LLMs). (Learn more atdynamo.ai) Our team of Ph.D.s and engineers pushes the boundaries of AI research, with real-world impact. We offer a fast-paced, collaborative environment free from...
- ...disability, or status as a protected veteran.**Job Description:****ML Engineer****3M Health Care is now Solventum****At Solventum, we... ...by implementing automated workflows, managing cloud infrastructure, and ensuring our AI services are secure and scalable.**Key...H1bRemote work
$160k - $240k
A leading financial technology company in New York is seeking a Senior MLOps Engineer to help enhance AI systems and processes. Candidates should have at least 4 years of experience in programming, knowledge of machine learning frameworks, and experience with cloud technologies...- ...Consumer Electronics in 2019, 2022, and 2023. We are looking to bring on an experienced Data Engineer to join Eight Sleep’s machine learning team. This role involves monitoring production ML systems, building tools and pipelines for data analysis, dataset curation, training...Full timeSleeping nightsFlexible hoursNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to ML Infrastructure Engineer. Be the first to apply!
- data scientist machine learning engineer New York, NY
- machine learning ai engineer New York, NY
- computer vision machine learning engineer New York, NY
- machine learning engineer New York, NY
- ai ml engineer New York, NY
- machine learning software engineer New York, NY
- entry level machine learning engineer New York, NY
- junior machine learning research engineer New York, NY
- senior ml engineer New York, NY
- data infrastructure engineer New York, NY


