Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff+ Site Reliability Engineer, Safeguards ML Infra

Full-time

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.

What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing! ) postmortem action items into process and tooling changes.

Vacancy posted 12 days ago
Similar jobs that could be interesting for youBased on the Staff+ Site Reliability Engineer, Safeguards ML Infra in Washington DC vacancy
  • $143.7k - $194.4k

     ...available, scalable and distributed engineering systems for one of the...  ...Management (ADM) team in Ads AI Core Infra owns the central datalake for...  ..., business analysts, ML engineers, research scientists...  ...(design patterns, reliability and scaling) of new and existing... 
    Suggested
    Internship
    Flexible hours

    Amazon

    Washington DC
    a month ago
  • $136.2k - $214.01k

     ...across their people and AI workflows. Our mission is simple: safeguard the digital world and empower people to work securely and...  ...in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the... 
    Suggested
    Full time
    Flexible hours

    Proofpoint

    Laurel, MD
    4 days ago
  •  ...Description Role Overview We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the...  ...(SLOs) and Service Level Indicators (SLIs) for critical AI/ML services. Error Budgeting: Manage error budgets to balance... 
    Suggested
    Local area

    Tiger Analytics Inc.

    Washington DC
    27 days ago
  • $175k - $250k

    Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting...  ..., performance, and reliability across environments. What You...  ...working with AI infrastructure, ML pipelines, or GPU orchestration... 
    Suggested
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    2 days ago
  • $143.7k - $194.4k

     ...this at scale. The Software Development Engineer will design, build, and maintain cloud-based...  ..., and Datacenter Operations to manage AI/ML infrastructure.Key job...  ...design or architecture (design patterns, reliability and scaling) of new and existing systems... 
    Suggested
    Internship
    Flexible hours

    Amazon

    Washington DC
    8 days ago
  • $232k - $319k

     ...scale the service with great people and reliable, cost-effective, and efficient infrastructure...  ....  What you’ll be doing  Lead the Infra platform and shared services org and...  ...Accelerate the velocity of SRE and product engineering by developing robust platforms, powerful... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    14 days ago
  • $174k - $238k

     ...defining work. We're all in on this mission. If you are too, let's talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    2 days ago
  • $182k - $250.8k

     ...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    Washington DC
    12 hours ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    3 days ago
  • $207k - $284.9k

     ...on this mission. If you are too, let's talk.Senior Manager, Site Reliability EngineeringSecure Every Identity, from AI to HumanIdentity is...  ...mission. If you are too, let's talk.The Federal Operations Engineering GroupOkta's Federal Operations team supports government customers... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours
    Day shift

    Okta

    Washington DC
    1 day ago
  • $134.1k - $241.4k

     ...looking for an amazingly talented Release Train Engineer to join our team! In this role you will get to support the Joint Staff and Chief Digital and Artificial Intelligence...  ...Joint Staff, CDAO, or enterprise-level DoD AI/ML and data integration initiatives.Security... 
    Full time
    Work at office
    Flexible hours
    Shift work

    Parsons

    Arlington, VA
    12 hours ago
  • $200k - $287.5k

     ...gets done.Senior Software Engineer — Cortex TrainingThe Snowflake ML Platform team's mission...  ...that ships fast & sweats reliability and the researchers behind...  ...strongest candidates pair deep infra skills with real post-...  ...on the Snowflake Careers Site for salary and benefits... 

    Snowflake

    Washington DC
    2 days ago
  • $194k - $267k

     ...something more than once, automate it” and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    14 days ago
  • $90k - $130k

     ...solutions save agencies thousands of hours, safeguard national security, and strengthen health...  ...Credence has an immediate opening for a Site Reliability SME who has hands-on experience working as a Cloud Operations Engineer with experience in IT operations to join our... 
    Temporary work
    Work experience placement
    Immediate start
    Worldwide

    Credence

    McLean, VA
    6 days ago
  •  ...artificial intelligence (Al), machine learning (ML) and cross-domain transfer systems to...  ...While the work is primarily conducted on-site at our client location in Bethesda, MD, we...  ...Responsibilities:Work with a team of system integration engineers to successfully deploy release candidates... 
    Remote work
    Flexible hours

    Xcelerate Solutions

    Bethesda, MD
    12 hours ago
  • $107.9k - $195.05k

     ...artificial intelligence (Al), machine learning (ML) and cross-domain transfer systems to...  ...While the work is primarily conducted on-site at our client location in Bethesda, MD, we...  ...include:Work with a team of system integration engineers to successfully deploy release candidates... 
    Full time
    Remote work
    Flexible hours

    Leidos

    Bethesda, MD
    12 hours ago
  • $127.1k - $178.72k

     ...contributes to team efforts in engineering, analytics, and technical...  ...engineering, machine learning, or AI/ML applications in aviation or...  ...Benefits page on our Careers site.Compensation at Noblis is determined...  .... For part time or on-call staff, compensation is... 
    Permanent employment
    Full time
    Contract work
    Part time
    Local area
    Remote work

    Noblis

    Washington DC
    4 days ago
  • $174k - $238k

     .... We're all in on this mission. If you are too, let's talk. The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products Group (EPG). Our mission is to build highly reliable, scalable... 
    Local area
    Worldwide
    Flexible hours

    Okta

    Washington DC
    14 days ago
  • $107.9k - $195.05k

     ...is looking for a highly skilled platform engineer with deep expertise in operating systems,...  ...for the mission customers.This is a 100% on-site position. All work must be performed at the...  ...with Kubernetes cluster management and AI/ML workflow orchestration (Argo, Airflow, and... 
    Full time

    Leidos

    Bethesda, MD
    12 hours ago
  • $135k - $200k

     ...children, and more.The RoleWe are seeking a Forward Deployed Software Engineer to join a newly-formed team focused on developing advanced...  ...sensors, drones, robotics, edge AI, or related fields.Knowledge of AI/ML, sensor processing, and edge/on-prem hardware and networking is... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    4 days ago
  • $145k - $200k

     ...The RoleWe are seeking a Senior Software Engineer to join a customer-facing product engineering...  ...autonomous platformsEngineer scalable, reliable, and fail-safe systems capable of...  ...embedded systems, or roboticsKnowledge of AI/ML, sensor processing, or swarm behaviors is... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Work from home
    Relocation package

    Palantir Technologies

    Washington DC
    1 day ago
  • $130k - $175k

     ...make ongoing improvements to accuracy, reliability, and efficiency Treat customer demand signals...  ...of experience building and deploying AI/ML solutions in production...  ...Type: Full-timeCategory: Software Development, Engineering & ApplicationsSalaried: Salaried
    Work experience placement
    Work at office
    Remote work

    ECS Federal

    Arlington, VA
    3 days ago
  • $148.5k - $223.9k

     ...for an experienced and hands-on software engineer to join our team to build and scale the next...  ...AI and Machine Learning: Leverage AI/ML models for intelligent data analysis, anomaly...  ...strengthen Salesforce’s defenses and safeguard our customers’ trust against evolving cyber... 
    Full time

    Salesforce

    Washington DC
    1 day ago
  • We are seeking an experienced Platform Engineer to automate and operate infrastructure, containers...  ..., databases, data pipelines, and AI/ML workloads.Implement monitoring, logging,...  ..., networking, performance, and platform reliability issues.Develop technical documentation covering... 

    ANALYTICA

    Bethesda, MD
    12 hours ago
  •  ...agencies thousands of hours, safeguard national security, and strengthen...  ...for a Senior Applications Engineer who will be primarily...  ...Addressing Systems (DAAS), AI/ML program.The Senior Applications...  ...CASE) tools, while providing reliable cost and schedule estimates to... 
    Immediate start
    Worldwide

    Credence Management Solutions

    Mc Lean, VA
    2 days ago
  •  ...Division is seeking a Lead Systems Engineer with a blend of cloud...  ...time role will be embedded on-site with the sponsor in Bethesda,...  ...as code, CI/CD, DevSecOps, AI/ML, data analytics, and automated...  ...needs.Mentor and guide MITRE staff supporting related work efforts... 
    Full time
    For contractors
    Internship
    Local area

    Mitre

    Bethesda, MD
    2 days ago
  • $69.4k - $158k

     ...programming and familiarity with high reliability computing platforms and...  ...’s degree in CS, EE, Engineering, or Software EngineeringNice...  ...developmentExperience with AI/ML software safety and test and...  ...Resource page on our Careers site and reviewing Our Employee Benefits... 
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Arlington, VA
    2 days ago
  • $83.59k - $150k

     ...advanced intelligence, surveillance, cyber, engineering, and all-domain capabilities that...  ...ships and defense technology solutions that safeguard our seas, sky, land, space and cyber....  ...artificial intelligence, machine learning (AI/ML) experts; engineers; technologists;... 
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide

    HII Mission Technologies Division

    Washington DC
    2 days ago
  • $186.4k - $233k

     ...lifelong, Scale customers. Our Solutions Engineers ensure customers' first experiences with...  ...willing to relocateBackground working in AI/ML, particularly Generative AI and Large...  ...About Us:At Scale, our mission is to develop reliable AI systems for the world's most important... 
    Full time

    Scale AI

    Washington DC
    4 days ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!