Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.
What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing! ) postmortem action items into process and tooling changes.
- ...and (3) building the research and infra for horizontal integrations, such... ...frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they... ...with experience in some layer of ML infrastructure. - Have worked...SuggestedFull time
$182.52k - $297k
...Plaid Machine Learning Infrastructure (ML Infra) team is responsible for creating and maintaining... ...that enable efficient, scalable, reliable, responsible and secure machine learning... ...Plaid use cases to streamline feature engineering for both batch and real-time streaming features...SuggestedFull timeWork experience placement- ...threats while improving the safety and reliability of frontier models in security-... ...The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced... ...environments. Have experience in ML systems, security products, cyber...SuggestedFull time
- ...researchers to use and maximally utilized Optimizing and improving ML data loading transport and storage in highly distributed fully... ...scaled ChatGPT and GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built next-generation...SuggestedFull time
- ...building AI agents that can reliably do everyday digital... ...of the AI technical staff to join the founding team... ...Responsibilities: Scale infra for post-training of... ...Work closely with product engineers to translate cutting‑... ...for: Experience with ML infrastructure (GPU clusters...SuggestedWork at officeRelocationVisa sponsorship
- ...threats while improving the safety and reliability of frontier models in security-... ...The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced... ...environments. Have experience in ML systems, security products, cyber...Full time
- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in... ...power providers, leveraging cutting-edge Machine Learning (ML) and Artificial Intelligence (AI) technologies. Our mission...Work at officeWeekend work
- ...treatment. What We Look for in a Great Engineer You have the intensity and technical... ...both the TypeScript and Python/ML deployment pipelines to support high-velocity... ...feature release while maintaining the highest reliability. DevX Support: Support Developer Experience...Work at office
- ...among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best... ...you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP...Local areaRemote workVisa sponsorshipWork visaFlexible hours
$300k
...practical experience in computer science, engineering, machine learning, or a related field.... ...of post-graduate software engineering or ML engineering experience, excluding internships... ...fundamentals and experience building reliable, maintainable systems. Proficiency in...Full timeInternshipVisa sponsorshipRelocation package$250k
...the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud... ...workloads. The role involves working closely with platform, ML, and infrastructure teams to improve reliability, automation...Full timeRemote work$210k - $240k
Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay... ...that powers our core platform—including data pipelines, ML workloads, and real-time analytics systems. This is a hands‑...Full time$175k - $250k
...: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote... ...unavailable. Modality: On-Site only. Must live within... ...scalability, performance, and reliability across environments. What You... ...working with AI infrastructure, ML pipelines, or GPU orchestration...Full timeRemote workRelocationRelocation package$227.2k - $324.5k
About the Role:As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams... ...identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if...Full timeTemporary workLocal areaFlexible hours- ...analytical database for multimodal, multi-rate data streams, on top of the open source Vortex file format. Our users are AI/ML researchers and AI infra engineers developing models in complex domains, such as weather & climate, financial, time-series, genomics, point-clouds,...Full timeWork at office
$293.6k - $335.1k
...creating responsible and reliable AI systems, changing... ...applications of AI & ML are bringing humanity... ...class applied science and engineering teams to deliver our... ...through this site. Capital One Financial... ...measures is crucial to safeguarding your information from...Full timePart timeLocal area- ...or data scientist can scale an ML application from their laptop... ...Anyscale is looking for a Software Engineer to join the Infrastructure... ...your laptop. As part of the Infra team, we build the scalable, secure... ...features to enhance the reliability, performance, scalability, and...Full time
- ...three exceptional Founding Software Engineers to help us scale the computational biology... ...biology tooling, ensuring reliability, scalability, and performance as we grow... ...ship quickly Build/operate services (infra + ML + web) and iterate with customer feedback...Full timeRelocation
$220k - $247.5k
...playing games. We are looking for a Senior Machine Learning Engineer to join our Revenue ML team at Discord. This role sits at the intersection of... .... Partner closely with Shop, Game Commerce, Revenue Infra, ML Infra and Data Engineering teams to define ML requirements...Full timeSeasonal work$150k - $200k
...their life’s work. About the Role: As a Software Engineer on Collections Infra, you’ll help scale the infrastructure behind Notion’s database... ...-native work, those systems need to become faster, more reliable, and ready for much higher concurrency. This team sits...Full timeLocal area- ...a small, fast-moving team of engineers focused on delivering a world-... ...team responsible for building reliable, high-performance infrastructure... ...closely with researchers, infra teams, and product engineers to... ...Have worked with GPU-based ML workloads and understand the performance...Full time
- ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE... ...become incidents Partner with SRE and Infra teams to ensure Capacity reflects the... ...and you follow through ~ Interest in AI/ML infrastructure; familiarity with GPU infrastructure...Full timeFlexible hours
- ...We are looking for Software Engineers who are "Product Architects"... ...our cloud or their own private infra. The Agentic Engine: We are... ...interactions between users and ML models. You’ll be responsible... ...interactions feel snappy and reliable. The "Design Partner" Sprint...Full time
$220k - $260k
...this role As a Senior Backend/Infra Engineer at Amigo, you'll build the... ...conversations a month, so concurrency, reliability, and clean design are the job... ...high bar You can work on site in New York City or San... ...with production AI or ML systems Benefits (available...Full timeFlexible hours$269.1k - $307.2k
Distinguished AI Engineer (Agentic AI Platform)... ...creating responsible and reliable AI systems,... ...applications of AI & ML are bringing humanity... ...model minutiae or infra plumbing. You... ...office hours, mentoring Staff, Principal and... ...available through this site. Capital...Full timePart timeWork at officeLocal area- ...What you’ll do As a Software Engineer, Infrastructure at Sierra, you... ...’s infrastructure secure, reliable, and scalable, enabling product... ...databases, retrieval systems, and ML models. Develop and... ...startup environment or platform/infra-focused team. Our values...Full timeFlexible hours
- ...developer or data scientist can scale an ML application from their laptop to the... ...with high levels of performance and reliability. We're looking for engineers with systems software experience... .../ distributed libraries, test infra improvements, debugging, and longer-...Full timeWork experience placement
$2,000 per month
...Our vision is to build a world where AI/ML and analytics are powered by decentralized... ...Role As a Founding Principal Software Engineer , you will help build out the next generation... .../ analytics OSS projects or internal data infra products Experience working in digital-...Full time- ...Staff Software Engineer San Francisco, CA Hybrid About us... ...founders). You own reliability, security, and... ...designed. Use AI/ML to turn freeform doctors... ...Identifying the biggest infra cost opportunities and... ...Quarterly company off-sites with the team ⛷️ ~...Full timeWork at officeRemote workWork from homeRelocationFlexible hours
- ...us and help build the platform engineers turn to to ship AI products.... ...running on our platform are fast, reliable, and cost‑efficient. As part... ..., model performance, and infra, helping to define how developers... ...engineering fundamentals and curiosity. ML experience is a plus, but not...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!
- engineering aide San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA
- assistant engineer San Francisco, CA
- software engineer staff San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- staff data engineer San Francisco, CA
- technology administrator San Francisco, CA



