Staff+ Site Reliability Engineer, Safeguards ML Infra
Anthropic
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.
What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing! ) postmortem action items into process and tooling changes.
$227.2k - $324.5k
About the Role:As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams... ...identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if...SuggestedFull timeTemporary workLocal areaFlexible hours$125k - $150k
...minute that determine cost, reliability, and margin, but the... ...Overview As a Senior Software Engineer - Backend & AI Infra focus, you will play a... ...designing or supporting AI/ML systems in production is a... ...travel occasionally to customer sites to understand real-world constraints...SuggestedFull timeWork experience placementLive inWork at officeVisa sponsorshipFlexible hours- ...role As a Senior Backend/Infra Engineer at Amigo, you'll build the core... ...a month, so concurrency, reliability, and clean design are the job... ...high bar You can work on site in New York City or San Francisco... ...Experience with production AI or ML systems Benefits (...SuggestedFull timeFlexible hours
$207k - $300k
Lead a team of Software/Systems Engineers on projects for users and be directly responsible... ...of engineers on large-scale projects.Site Reliability Engineering (SRE) combines software and... ...a Software Engineer chose to join SRE.Infra Bigtable SRE is responsible for running...Suggested$320k
...s mission is to create reliable, interpretable, and steerable... ...committed researchers, engineers, policy experts, and... ...hand in hand. Node Infra owns the full lifecycle... ...) for distributed ML workloads. ~ Demonstrated... ..., we expect all staff to be in one of our offices...SuggestedFull timeWork at officeVisa sponsorshipFlexible hours- ...GPU Systems / AI Infrastructure Engineer (NYC) Location: New York City (Hybrid / On-site preferred) Comp: Competitive +... ...equity (Series A-C / high-growth AI infra) About the Role We’re hiring... ...Collaborate closely with ML researchers and infra engineers to...Full time
$200k - $300k
...Staff+ Software Engineer, Backend and Infra Haize Labs takes AI-based applications from proof-of-concept to production... ...We eliminate risk and improve the reliability of LLM-based applications by... ...Have experience doing applied AI/ML Have strong familiarity with security...Visa sponsorship$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous... ...is a role for engineers who want it all — design work, ML systems, and operational ownership. On-call and incident response...Full timeTemporary workLocal areaRemote workWorldwide$252k - $315k
...Our Generative AI Data Engine powers the world’s... ...horizontal, high-impact L6 Staff Fullstack Engineer &... ...and incentives, to safeguarding data integrity through... ...at the intersection of ML, operations, and analytics... ...for scalability, reliability, and performance Mentor...Full time- ...business continuity. Appgate safeguards enterprises and government... ...Role We are seeking Staff and Principal Engineers who thrive in ambiguity,... ...product teams to operationalize ML/AI models on top of the... ...to ensure scalability, reliability, and compliance. Lead and...Full timeTemporary workRemote workWork from homeWorldwideHome officeFlexible hours3 days per week
$191k - $226k
...than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads. This role sits on our Platform Engineering team. You...Remote workWork visaFlexible hours$168k - $200k
...For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the... ...data pipelines, ML workflows, and infra components using GitHub Actions... ...ensure the safety of patients and staff, many of our clients require post-...$135k - $200k
...modern security threats.Our mission is to safeguard systems and data by developing innovative... ...: We are seeking a Lead Penetration Test Engineer with extensive experience in penetration... ...across AWS, Azure, or GCP. • Knowledge of AI/ML security and adversarial testing methods,...Second jobLive inWorldwideFlexible hours$200k - $250k
...This innovative field blends AI, engineering, and materials science,... ...everything secure, observable, and reliable in production. We work in a... ...-discipline domain—robotics, ML, and experimental automation—and... .../Grafana, etc.). Hybrid infra (cloud + on-prem), containerization...Full time- ...re looking for incredible Senior Backend Engineers to help us build an ambitious roadmap .... ...across the entire backend stack, from data infra to product endpoints Build extensible... ...improve accuracy, speed, and cost of AI/ML models. Establish best practices for scalability...Full timeWork at office
$1,000 per month
...the Role We're hiring a Senior Software Engineer to join the product engineering team. You'... ...defined spec to start moving. Real AI/ML experience. You've shipped AI features to... ...Typescript, AWS, Render, Snowflake DevOps/infra depth: CI/CD, monitoring, security...Full timeWork at officeRemote workAll shiftsFlexible hours$147.2k - $224.3k
...complexity of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale... ...idempotency, and invoice-accuracy safeguards to ensure revenue data is always... ...~ AI-native company — your infra directly powers AI agents used by leading...Full timeContract workTemporary workWork at officeImmediate startRemote workFlexible hours- ...Description Job Description Staff / Principal Platform Engineer Location: New York... ...continuity. AppGate safeguards Fortune 500 enterprises and... ...experience operationalizing AI/ML systems, and you treat... ...shapes the architecture, reliability and core platform that defines...For contractorsWorldwide
$140k - $160k
...better. We’re hiring an exceptional ML Engineer to join our team (Boston or NYC office).... ..., ensuring scalability, efficiency, and reliability. What you’ll do Design and develop... ...Support our engineering team on ongoing infra improvements such as end to end testing...Full timeWork at office$146.8k - $272.6k
...every day. As a Lead Software Engineer (Staff Engineer), you will own core... ...that turn frontier models into reliable, production‑grade workflows... ...cut across AI, product, and infra (e.g., a new orchestration layer... ...closely with AI/ML engineers, researchers, designers...Full timeWork at officeLocal areaFlexible hours- ..., superior protection and seamless interoperability. AppGate safeguards Fortune 500 enterprises worldwide. Learn more at appgate.com. About the Role We're looking for a AI/ML Engineer (Senior/Staff/Principal) - Threat Detection who will design, build, and operationalize...Full timeFor contractorsWorldwide
- Applied AI Engineer Duration: 6-12 months contract to hire Location: NYC or San Francisco,... ...deploy software and biological AI systems to safeguard humanity. The same AI architectures that... ...from Palantir and applied biological ML engineers from MIT's Broad Institute and...Contract work
$200k - $400k
...power Decagon: networking, data, ML serving, developer platform,... ...five focus areas: Core Infra: The foundational cloud stack... ...infrastructure‑as‑code—to ensure reliability, scale, and cost efficiency.... ...hiring a Senior Infrastructure Engineer to design, build, and operate...Full timeWork at officeLocal area$182k - $250.8k
...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great... ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work$190k - $260k
...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving...Full timeWork experience placementWork at officeLocal areaRemote workHome office$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...gap between our GTM team, core engineering team, and the unique, complex... ...feedback loop to the product and infra teams-surfacing usability gaps... ...the latest developments in ML/AI This is an in person role... ...approaches fail at reliably extracting information in complex...Work at officeLocal area
$175k - $275k
...We're looking for an AI Engineer to help build the agentic platform... ...harness engineering that make it reliable enough for healthcare finance.... ...on model, framework, and infra Turn capabilities into an ecosystem... ...curiosity about it Applied ML or model-evaluation research...Full timeWork at office- ...to improve their business. Founded by engineers — and customer obsessed — we leap at every... ...Model APIs, to state a few. Improve reliability, latency, and efficiency of distributed... ...Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!
- staff data engineer New York, NY
- project engineer assistant project manager New York, NY
- senior staff engineer New York, NY
- senior staff systems engineer New York, NY
- assistant chief engineer New York, NY
- engineering aide New York, NY
- software engineer staff New York, NY
- staff design engineer New York, NY
- assistant engineering manager New York, NY
- assistant engineer New York, NY



