Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff+ Site Reliability Engineer, Safeguards ML Infra

Full-time

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job.

What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter. What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing!

Vacancy posted 25 days ago
Similar jobs that could be interesting for youBased on the Staff+ Site Reliability Engineer, Safeguards ML Infra in San Francisco, CA vacancy
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Suggested
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $148.5k - $223.9k

     ...Salesforce. Job Details Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco. Working closely with...  ...performance and reliability. Understanding of AI/ML concepts applied to operations (e.g., anomaly... 
    Suggested
    Worldwide
    Weekend work

    100 Salesforce, Inc.

    San Francisco, CA
    11 hours ago
  •  ...treatment. What We Look for in a Great Engineer Tool Proficiency: You are highly...  ...streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer... 
    Suggested
    Work at office

    Latent

    San Francisco, CA
    2 days ago
  • $250k

     ...the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud...  ...workloads. The role involves working closely with platform, ML, and infrastructure teams to improve reliability, automation... 
    Suggested
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and...  ...logs, traces—so issues surface before users notice. Automate infra provisioning and config with Terraform, Helm and Kubernetes Operators... 
    Suggested

    Tala Health

    San Francisco, CA
    4 days ago
  •  ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth...  ...mechanisms that react to our workload. Deploy ML systems across the company. Qualifications Has... 
    Worldwide
    Home office
    Flexible hours

    Coda

    San Francisco, CA
    2 days ago
  • $300 per month

     ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you will define and codify...  ...the gold standards of day‑2 operations for our ML infrastructure platform. You will envision and... 
    Flexible hours

    Baseten

    San Francisco, CA
    2 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay...  ...that powers our core platform—including data pipelines, ML workloads, and real-time analytics systems. This is a hands‑... 
    Full time

    Alembic Technologies

    San Francisco, CA
    3 days ago
  •  ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and...  ...infrastructure that powers our core platform—including data pipelines, ML workloads, and real-time analytics systems.This is a hands-... 

    Alembic Limited

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote...  ...unavailable. Modality: On-Site only. Must live within...  ...scalability, performance, and reliability across environments. What You’...  ...working with AI infrastructure, ML pipelines, or GPU orchestration... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    3 days ago
  •  ...lab in San Francisco, seeks a Senior Platform Engineer to evolve infrastructure, tooling, and shared platform capabilities for reliable, secure services across cloud and edge...  ...collaborate with application, services, AI/ML, and security teams to raise developer velocity... 

    Jobleads-US

    San Francisco, CA
    2 days ago
  • $252k - $315k

     ...Our Generative AI Data Engine powers the world’s...  ...horizontal, high-impact L6 Staff Fullstack Engineer &...  ...and incentives, to safeguarding data integrity through...  ...at the intersection of ML, operations, and analytics...  ...for scalability, reliability, and performance... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $250k

     ...without traditional infrastructure limitations. As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale...  ...scheduling, inference optimisation, and distributed systems reliability, working alongside highly technical teams at the... 
    Full time
    San Francisco, CA
    more than 2 months ago
  • $204k - $216k

     ...reasoning behind MINERVA and COGENT into fast, reliable, affordable answers for members. You...  ...to applied AI, research, and platform engineering, and you make the difference between a...  ...infrastructure engineering to get right. The AI/ML Infrastructure and Systems Engineer owns... 

    Sapience AI Corporation

    San Francisco, CA
    3 days ago
  • $130.6k - $192k

     ...state-of-the-art platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-technical audiences to come up with...  ...experienced veterans of backend, web, statistical and data infra engineers and work closely with the data science community... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    4 hours ago
  •  ...California. The Role: As a Platform Engineer , you’ll be responsible for...  ...work will be essential to ensuring the reliability and reproducibility of ML workloads, the safety and control of...  ...closely with ML engineers, DevOps, and infra teams to improve system reliability... 
    Work at office
    Relocation package

    Zyphra

    San Francisco, CA
    a month ago
  • $260k - $300k

     .... About this role As a Staff Backend/Infra Engineer at Concurrence, you'll build...  ...a month, so concurrency, reliability, and clean design are the job...  ...bar You can work on site in New York City or San Francisco...  ...with production AI or ML systems Benefits (available... 
    Full time
    Flexible hours

    Concurrence

    San Francisco, CA
    20 days ago
  • $190.8k - $267.1k

     ...and the UserYou will collaborate closely with client media infra, backend, Feeds, and ML ranking teams. Your role is to translate powerful...  ...love.ð Lead Cross-Functional ExecutionAct as a strategic engineering partner to Product, Design, and Data Science. You will help... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    1 day ago
  •  ...in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we...  ...Familiarity with deployment pipelines, CI/CD, or infra-as-code Experience improving system observability (e.g.... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...high-volume data replication simple, reliable, and scalable for engineering teams. Our platform powers mission...  ...visibility, customer-facing analytics, and AI/ML workloads. We’re trusted by teams...  ..., Redis, Kafka and Elasticsearch Infra: Terraform, Kubernetes, and Helm on... 
    Visa sponsorship

    Artie's

    San Francisco, CA
    5 days ago
  • $7.3 per hour

     ...The role We’re looking for a world‑class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...just maintain existing ones. You thrive in a zero‑to‑one infra environment. High‑velocity execution: you have a strong... 

    Blaxel (YC X25)

    San Francisco, CA
    2 days ago
  •  ...for the future. As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research...  ...provide the best experience to our technical staff. You will leverage IaC, Automation, and...  ..., Dask, Spark). Familiarity with ML frameworks (PyTorch/Tensorflow, JAX,... 
    Local area

    The Voleon Group

    Berkeley, CA
    1 day ago
  •  ...on HaluEval, the CTGT Policy Engine (paired with GPT-120B OSS) outperformed...  ...large language models more reliable, controllable, and performant...  ..., and graph databases Infra: Docker, Kubernetes, Terraform...  ...providers and customer VPCs ML: Self hosted models on... 

    CTGT

    San Francisco, CA
    2 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    1 day ago
  • $204k - $306k

     ...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from...  ...week in our San Francisco Office.The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta, Inc.

    San Francisco, CA
    1 day ago
  • Software Engineer, Platform Arena Intelligence is looking...  ...users that scales, is reliable, and makes the...  ...meaningful backend work. Staff: extensive, deep experience...  ...custom). Background in AI/ML infrastructure, model...  ...with the modern AI infra stack (vLLM, LiteLLM, LangChain... 
    Permanent employment
    Remote work
    Flexible hours
    Shift work
    3 days per week

    Arena AI

    San Francisco, CA
    4 days ago
  •  ...Security Engineer Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models...  ...If you want to build massive-scale ML systems that will define the way the... 
    H1b

    Exa Labs

    San Francisco, CA
    2 days ago
  •  ...gap between our GTM team, core engineering team, and the unique, complex...  ...feedback loop to the product and infra teamssurfacing usability gaps,...  ...the latest developments in ML/AI This is an in person role...  ...approaches fail at reliably extracting information in complex... 
    Work at office
    Local area

    Reducto

    San Francisco, CA
    4 days ago
  •  ...hiring three exceptional Founding Software Engineers to help us scale the computational...  ...computational biology tooling, ensuring reliability, scalability, and performance as we grow...  ...ship quickly Build/operate services (infra + ML + web) and iterate with customer feedback... 
    Relocation

    Tamarind Bio, Inc

    San Francisco, CA
    3 days ago
  • $75k - $100k

     ...bringing together breakthroughs in AI, systems engineering, and product design. Our team is made up...  ...with distributed systems, cloud infra, and high-performance services Proficiency...  ...Experience with compilers, developer tools, or ML systems Even if you don’t meet every... 
    Full time
    Worldwide

    Emergent

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!