Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff+ Site Reliability Engineer, Safeguards ML Infra

Full-time

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.

What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing! ) postmortem action items into process and tooling changes.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Staff+ Site Reliability Engineer, Safeguards ML Infra in San Francisco, CA vacancy
  •  ...and (3) building the research and infra for horizontal integrations, such...  ...frontier RL training runs fast, reliable, and unblocked. You will work across engineering and infrastructure problems as they...  ...with experience in some layer of ML infrastructure. - Have worked... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  • $182.52k - $297k

     ...Plaid Machine Learning Infrastructure (ML Infra) team is responsible for creating and maintaining...  ...that enable efficient, scalable, reliable, responsible and secure machine learning...  ...Plaid use cases to streamline feature engineering for both batch and real-time streaming features... 
    Suggested
    Full time
    Work experience placement

    Plaid

    San Francisco, CA
    14 hours ago
  •  ...threats while improving the safety and reliability of frontier models in security-...  ...The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced...  ...environments. Have experience in ML systems, security products, cyber... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...researchers to use and maximally utilized Optimizing and improving ML data loading transport and storage in highly distributed fully...  ...scaled ChatGPT and GPT-4 to hundreds of millions of users, engineered the foundations of autonomous driving, built next-generation... 
    Suggested
    Full time

    The Generalist

    San Francisco, CA
    14 hours ago
  •  ...building AI agents that can reliably do everyday digital...  ...of the AI technical staff to join the founding team...  ...Responsibilities: Scale infra for post-training of...  ...Work closely with product engineers to translate cutting‑...  ...for: Experience with ML infrastructure (GPU clusters... 
    Suggested
    Work at office
    Relocation
    Visa sponsorship

    Yutori

    San Francisco, CA
    4 days ago
  •  ...threats while improving the safety and reliability of frontier models in security-...  ...The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced...  ...environments. Have experience in ML systems, security products, cyber... 
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in...  ...power providers, leveraging cutting-edge Machine Learning (ML) and Artificial Intelligence (AI) technologies. Our mission... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    4 days ago
  •  ...treatment. What We Look for in a Great Engineer You have the intensity and technical...  ...both the TypeScript and Python/ML deployment pipelines to support high-velocity...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer Experience... 
    Work at office

    Latent

    San Francisco, CA
    1 day ago
  •  ...among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best...  ...you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP... 
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    AgileEngine

    San Francisco, CA
    5 days ago
  • $300k

     ...practical experience in computer science, engineering, machine learning, or a related field....  ...of post-graduate software engineering or ML engineering experience, excluding internships...  ...fundamentals and experience building reliable, maintainable systems. Proficiency in... 
    Full time
    Internship
    Visa sponsorship
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    17 days ago
  • $250k

     ...the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud...  ...workloads. The role involves working closely with platform, ML, and infrastructure teams to improve reliability, automation... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $210k - $240k

    Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay...  ...that powers our core platform—including data pipelines, ML workloads, and real-time analytics systems. This is a hands‑... 
    Full time

    Alembic Technologies

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote...  ...unavailable. Modality: On-Site only. Must live within...  ...scalability, performance, and reliability across environments. What You...  ...working with AI infrastructure, ML pipelines, or GPU orchestration... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    1 day ago
  • $227.2k - $324.5k

    About the Role:As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams...  ...identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    5 days ago
  •  ...analytical database for multimodal, multi-rate data streams, on top of the open source Vortex file format. Our users are AI/ML researchers and AI infra engineers developing models in complex domains, such as weather & climate, financial, time-series, genomics, point-clouds,... 
    Full time
    Work at office

    Spiral

    San Francisco, CA
    14 hours ago
  • $293.6k - $335.1k

     ...creating responsible and reliable AI systems, changing...  ...applications of AI & ML are bringing humanity...  ...class applied science and engineering teams to deliver our...  ...through this site. Capital One Financial...  ...measures is crucial to safeguarding your information from... 
    Full time
    Part time
    Local area

    SupportFinity

    San Francisco, CA
    2 days ago
  •  ...or data scientist can scale an ML application from their laptop...  ...Anyscale is looking for a Software Engineer to join the Infrastructure...  ...your laptop. As part of the Infra team, we build the scalable, secure...  ...features to enhance the reliability, performance, scalability, and... 
    Full time

    Anyscale

    San Francisco, CA
    14 hours ago
  •  ...three exceptional Founding Software Engineers to help us scale the computational biology...  ...biology tooling, ensuring reliability, scalability, and performance as we grow...  ...ship quickly Build/operate services (infra + ML + web) and iterate with customer feedback... 
    Full time
    Relocation

    Tamarind Bio

    San Francisco, CA
    14 hours ago
  • $220k - $247.5k

     ...playing games. We are looking for a Senior Machine Learning Engineer to join our Revenue ML team at Discord. This role sits at the intersection of...  .... Partner closely with Shop, Game Commerce, Revenue Infra, ML Infra and Data Engineering teams to define ML requirements... 
    Full time
    Seasonal work

    Discord

    San Francisco, CA
    14 hours ago
  • $150k - $200k

     ...their life’s work. About the Role: As a Software Engineer on Collections Infra, you’ll help scale the infrastructure behind Notion’s database...  ...-native work, those systems need to become faster, more reliable, and ready for much higher concurrency. This team sits... 
    Full time
    Local area

    Notion

    San Francisco, CA
    14 hours ago
  •  ...a small, fast-moving team of engineers focused on delivering a world-...  ...team responsible for building reliable, high-performance infrastructure...  ...closely with researchers, infra teams, and product engineers to...  ...Have worked with GPU-based ML workloads and understand the performance... 
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE...  ...become incidents Partner with SRE and Infra teams to ensure Capacity reflects the...  ...and you follow through ~ Interest in AI/ML infrastructure; familiarity with GPU infrastructure... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  •  ...We are looking for Software Engineers who are "Product Architects"...  ...our cloud or their own private infra. The Agentic Engine: We are...  ...interactions between users and ML models. You’ll be responsible...  ...interactions feel snappy and reliable. The "Design Partner" Sprint... 
    Full time

    Hilbert's Ai

    San Francisco, CA
    14 hours ago
  • $220k - $260k

     ...this role As a Senior Backend/Infra Engineer at Amigo, you'll build the...  ...conversations a month, so concurrency, reliability, and clean design are the job...  ...high bar You can work on site in New York City or San...  ...with production AI or ML systems   Benefits (available... 
    Full time
    Flexible hours

    Amigo

    San Francisco, CA
    15 days ago
  • $269.1k - $307.2k

    Distinguished AI Engineer (Agentic AI Platform)...  ...creating responsible and reliable AI systems,...  ...applications of AI & ML are bringing humanity...  ...model minutiae or infra plumbing. You...  ...office hours, mentoring Staff, Principal and...  ...available through this site. Capital... 
    Full time
    Part time
    Work at office
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    14 hours ago
  •  ...What you’ll do As a Software Engineer, Infrastructure at Sierra, you...  ...’s infrastructure secure, reliable, and scalable, enabling product...  ...databases, retrieval systems, and ML models. Develop and...  ...startup environment or platform/infra-focused team. Our values... 
    Full time
    Flexible hours

    Sierra Limited

    San Francisco, CA
    14 hours ago
  •  ...developer or data scientist can scale an ML application from their laptop to the...  ...with high levels of performance and reliability. We're looking for engineers with systems software experience...  .../ distributed libraries, test infra improvements, debugging, and longer-... 
    Full time
    Work experience placement

    Anyscale

    San Francisco, CA
    14 hours ago
  • $2,000 per month

     ...Our vision is to build a world where AI/ML and analytics are powered by decentralized...  ...Role As a Founding Principal Software Engineer , you will help build out the next generation...  .../ analytics OSS projects or internal data infra products Experience working in digital-... 
    Full time

    Nextdata Technologies Inc

    San Francisco, CA
    14 hours ago
  •  ...Staff Software Engineer San Francisco, CA Hybrid About us...  ...founders). You own reliability, security, and...  ...designed. Use AI/ML to turn freeform doctors...  ...Identifying the biggest infra cost opportunities and...  ...Quarterly company off-sites with the team ⛷️ ~... 
    Full time
    Work at office
    Remote work
    Work from home
    Relocation
    Flexible hours

    Metriport

    San Francisco, CA
    14 hours ago
  •  ...us and help build the platform engineers turn to to ship AI products....  ...running on our platform are fast, reliable, and cost‑efficient. As part...  ..., model performance, and infra, helping to define how developers...  ...engineering fundamentals and curiosity. ML experience is a plus, but not... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!