Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job.
What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter. What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing!
$152.5k - $205k
...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical...SuggestedFlexible hours$148.5k - $223.9k
...Salesforce. Job Details Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in SanFrancisco. Working closely with... ...performance and reliability. Understanding of AI/ML concepts applied to operations (e.g., anomaly...SuggestedWorldwideWeekend work- ...treatment. What We Look for in a Great Engineer Tool Proficiency: You are highly... ...streamline both the TypeScript and Python/ML deployment pipelines to support high-velocity... ...feature release while maintaining the highest reliability. DevX Support: Support Developer...SuggestedWork at office
$250k
...the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud... ...workloads. The role involves working closely with platform, ML, and infrastructure teams to improve reliability, automation...SuggestedFull timeRemote work- ...arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and... ...logs, traces—so issues surface before users notice. Automate infra provisioning and config with Terraform, Helm and Kubernetes Operators...Suggested
- ...be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth... ...mechanisms that react to our workload. Deploy ML systems across the company. Qualifications Has...WorldwideHome officeFlexible hours
$300 per month
...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you will define and codify... ...the gold standards of day‑2 operations for our ML infrastructure platform. You will envision and...Flexible hours$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay... ...that powers our core platform—including data pipelines, ML workloads, and real-time analytics systems. This is a hands‑...Full time- ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and... ...infrastructure that powers our core platform—including data pipelines, ML workloads, and real-time analytics systems.This is a hands-...
$175k - $250k
...: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote... ...unavailable. Modality: On-Site only. Must live within... ...scalability, performance, and reliability across environments. What You’... ...working with AI infrastructure, ML pipelines, or GPU orchestration...Full timeRemote workRelocationRelocation package- ...lab in San Francisco, seeks a Senior Platform Engineer to evolve infrastructure, tooling, and shared platform capabilities for reliable, secure services across cloud and edge... ...collaborate with application, services, AI/ML, and security teams to raise developer velocity...
$252k - $315k
...Our Generative AI Data Engine powers the world’s... ...horizontal, high-impact L6 Staff Fullstack Engineer &... ...and incentives, to safeguarding data integrity through... ...at the intersection of ML, operations, and analytics... ...for scalability, reliability, and performance...Full time$250k
...without traditional infrastructure limitations. As a Senior ML Infrastructure Engineer, the successful candidate will help build and scale... ...scheduling, inference optimisation, and distributed systems reliability, working alongside highly technical teams at the...Full time$204k - $216k
...reasoning behind MINERVA and COGENT into fast, reliable, affordable answers for members. You... ...to applied AI, research, and platform engineering, and you make the difference between a... ...infrastructure engineering to get right. The AI/ML Infrastructure and Systems Engineer owns...$130.6k - $192k
...state-of-the-art platform in industry that enables Product Engineers, Data Scientists, ML Engineers and non-technical audiences to come up with... ...experienced veterans of backend, web, statistical and data infra engineers and work closely with the data science community...Hourly payWork at officeLocal areaRemote workFlexible hours- ...California. The Role: As a Platform Engineer , you’ll be responsible for... ...work will be essential to ensuring the reliability and reproducibility of ML workloads, the safety and control of... ...closely with ML engineers, DevOps, and infra teams to improve system reliability...Work at officeRelocation package
$260k - $300k
.... About this role As a Staff Backend/Infra Engineer at Concurrence, you'll build... ...a month, so concurrency, reliability, and clean design are the job... ...bar You can work on site in New York City or San Francisco... ...with production AI or ML systems Benefits (available...Full timeFlexible hours$190.8k - $267.1k
...and the UserYou will collaborate closely with client media infra, backend, Feeds, and ML ranking teams. Your role is to translate powerful... ...love.ð Lead Cross-Functional ExecutionAct as a strategic engineering partner to Product, Design, and Data Science. You will help...For contractorsWork experience placement- ...in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we... ...Familiarity with deployment pipelines, CI/CD, or infra-as-code Experience improving system observability (e.g....WorldwideShift work
$150k - $250k
...high-volume data replication simple, reliable, and scalable for engineering teams. Our platform powers mission... ...visibility, customer-facing analytics, and AI/ML workloads. We’re trusted by teams... ..., Redis, Kafka and Elasticsearch Infra: Terraform, Kubernetes, and Helm on...Visa sponsorship$7.3 per hour
...The role We’re looking for a world‑class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ...just maintain existing ones. You thrive in a zero‑to‑one infra environment. High‑velocity execution: you have a strong...- ...for the future. As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research... ...provide the best experience to our technical staff. You will leverage IaC, Automation, and... ..., Dask, Spark). Familiarity with ML frameworks (PyTorch/Tensorflow, JAX,...Local area
- ...on HaluEval, the CTGT Policy Engine (paired with GPT-120B OSS) outperformed... ...large language models more reliable, controllable, and performant... ..., and graph databases Infra: Docker, Kubernetes, Terraform... ...providers and customer VPCs ML: Self hosted models on...
$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$204k - $306k
...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from... ...week in our San Francisco Office.The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week- Software Engineer, Platform Arena Intelligence is looking... ...users that scales, is reliable, and makes the... ...meaningful backend work. Staff: extensive, deep experience... ...custom). Background in AI/ML infrastructure, model... ...with the modern AI infra stack (vLLM, LiteLLM, LangChain...Permanent employmentRemote workFlexible hoursShift work3 days per week
- ...Security Engineer Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models... ...If you want to build massive-scale ML systems that will define the way the...H1b
- ...gap between our GTM team, core engineering team, and the unique, complex... ...feedback loop to the product and infra teamssurfacing usability gaps,... ...the latest developments in ML/AI This is an in person role... ...approaches fail at reliably extracting information in complex...Work at officeLocal area
- ...hiring three exceptional Founding Software Engineers to help us scale the computational... ...computational biology tooling, ensuring reliability, scalability, and performance as we grow... ...ship quickly Build/operate services (infra + ML + web) and iterate with customer feedback...Relocation
$75k - $100k
...bringing together breakthroughs in AI, systems engineering, and product design. Our team is made up... ...with distributed systems, cloud infra, and high-performance services Proficiency... ...Experience with compilers, developer tools, or ML systems Even if you don’t meet every...Full timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!
- software engineer staff San Francisco, CA
- technology administrator San Francisco, CA
- assistant engineer San Francisco, CA
- staff data engineer San Francisco, CA
- assistant electrical engineer San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA




