Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff+ Site Reliability Engineer, Safeguards ML Infra

Full-time

Anthropic

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: The Safeguards ML Infra team designs, builds, and operates the production infrastructure that powers Claude's safety systems. We own the critical backend services that ensure safety on the token generation path, and we own the operational work of getting those systems safely into production: standing up safeguards for every new model launch, and deploying new safety classifiers as they ship. Every frontier model release runs through this team – we configure, verify, and roll out safeguards across every platform Claude runs on (1P, AWS Bedrock, GCP Vertex, etc. ), and we lead incident response when issues arise. This role sits at the center of that operational work. You'll ensure safeguards are properly configured and deployed for model launches and own the off-cycle deployment of new safety classifiers — canarying changes, verifying that the right safeguards are provably live on the right models, and holding rollback authority when something looks wrong. Every launch should also shrink the checklist, and the manual verifications should evolve into a system that runs itself. You'll turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. We're looking for engineers with deep experience in production change management at scale — people who have owned deploy pipelines, config management systems, rollout safety, or launch readiness for systems under real production pressure. Familiarity with ML research or transformer architectures is not required — you will learn that on the job. What we prioritize is production judgment: a track record of shipping changes to critical systems safely, and of automating yourself out of the work you did last quarter.

What you'll do: - Launch captain model releases: stand up, configure, and verify safeguards for every new model, and serve as the safeguards point of contact in the launch room during release windows. - Own the off-cycle deployment of new safety classifiers as they ship from research — canarying rollouts, running post-deploy validations, and investigating discrepancies when something looks wrong. - Verify that the right safeguards are provably live on the right models across every deployment platform (1P, AWS Bedrock, GCP Vertex, etc. ), and detect and eliminate configuration drift between them. - Automate yourself out of last quarter's work: turn launch runbooks into tooling, hand-built checks into continuous validation, and one-off deploys into a repeatable pipeline. - Plan to use Claude aggressively to do this! And be a trailblazer that paves the path for safe agentic operations of safety-critical systems. - Build and maintain a safeguards registry with full provenance — what is running in production, on which model, on which platform, and when and by whom it was deployed. - Participate in on-call and operational-duty rotations covering service incidents, model provisioning, and time-sensitive research and safety launches. You may be a good fit if you: - Have owned production change management at scale — deploy pipelines, config management systems, canary analysis — and have strong opinions about what "verified" means. - Have run high-stakes releases: served as a launch captain, incident commander, or release owner for systems where a bad deploy has real consequences, and are energized rather than drained by being in the critical path. - Have meaningful on-call experience for production systems, including incident response and postmortem-driven improvements — and a track record of turning (and fixing! ) postmortem action items into process and tooling changes.

Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Staff+ Site Reliability Engineer, Safeguards ML Infra in New York, NY vacancy
  • $227.2k - $324.5k

    About the Role:As a Staff Software Engineer on the ML Infrastructure team, you will collaborate closely with the Machine Learning and Product teams...  ...identifying & optimizing latency, cost, and efficiency of our infra. Lead large scale cross functional refactorings if... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    New York, NY
    3 days ago
  •  ...role As a Senior Backend/Infra Engineer at Amigo, you'll build the core...  ...a month, so concurrency, reliability, and clean design are the job...  ...high bar You can work on site in New York City or San Francisco...  ...Experience with production AI or ML systems   Benefits (... 
    Suggested
    Full time
    Flexible hours

    Amigo S.r.l.

    New York, NY
    1 day ago
  • $125k - $150k

     ...minute that determine cost, reliability, and margin, but the...  ...Overview As a Senior Software Engineer - Backend & AI Infra focus, you will play a...  ...designing or supporting AI/ML systems in production is a...  ...travel occasionally to customer sites to understand real-world constraints... 
    Suggested
    Full time
    Work experience placement
    Live in
    Work at office
    Visa sponsorship
    Flexible hours

    Cvector Energy

    New York, NY
    1 day ago
  •  ...GPU Systems / AI Infrastructure Engineer (NYC) Location: New York City (Hybrid / On-site preferred) Comp: Competitive +...  ...equity (Series A-C / high-growth AI infra) About the Role We’re hiring...  ...Collaborate closely with ML researchers and infra engineers to... 
    Suggested
    Full time
    New York, NY
    more than 2 months ago
  • $139k - $257.55k

     ...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous...  ...is a role for engineers who want it all — design work, ML systems, and operational ownership. On-call and incident response... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    10 hours ago
  • $252k - $315k

     ...Our Generative AI Data Engine powers the world’s...  ...horizontal, high-impact L6 Staff Fullstack Engineer &...  ...and incentives, to safeguarding data integrity through...  ...at the intersection of ML, operations, and analytics...  ...for scalability, reliability, and performance Mentor... 
    Full time

    Scale Ai, Inc.

    New York, NY
    1 day ago
  • $191k - $226k

     ...healthier, faster. About the role We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads. This role sits on our Platform Engineering team. You... 
    Remote work
    Work visa
    Flexible hours

    Garner Health

    New York, NY
    2 days ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems... 
    Flexible hours

    Baseten

    New York, NY
    2 days ago
  •  ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE)...  ...work closely with backend, data, and ML engineers to keep the platform highly...  ...safety through reliability checks and safeguards. Contribute to CI/CD pipelines (GitHub... 
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    New York, NY
    5 days ago
  •  ...Job Description Staff / Principal DevOps Engineer Location: New York...  ...continuity. AppGate safeguards Fortune 500...  ...Elasticsearch) running reliably in production. Infrastructure...  ...but new to DevOps/infra, run design reviews,...  ...Operationalize AI/ML: Build the MLOps... 
    For contractors
    Worldwide

    AppGate Cybersecurity, Inc.

    New York, NY
    21 days ago
  • $168k - $200k

     ...For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the...  ...data pipelines, ML workflows, and infra components using GitHub Actions...  ...ensure the safety of patients and staff, many of our clients require post-... 

    Datavant

    New York, NY
    27 days ago
  • Basis is seeking a Site Reliability Engineer to ensure the reliability, scalability, and performance of our AI-powered accounting platform. You’ll join a high-leverage infrastructure team at the intersection of product and platform, owning systems that keepBasis fast, secure... 

    getbasis.ai

    New York, NY
    4 days ago
  •  ...helping professionals advance their careers. Role: Site Reliability Engineer - Cloud / DevOps Location: New York City, NY Hybrid...  ...orchestration and cloud-native infrastructure. Experience with AI/ML workload infrastructure is a plus. NMK Global Inc.... 
    Full time
    3 days per week

    NMK GLOBAL, INC

    New York, NY
    3 days ago
  • $135k - $200k

     ...modern security threats.Our mission is to safeguard systems and data by developing innovative...  ...: We are seeking a Lead Penetration Test Engineer with extensive experience in penetration...  ...across AWS, Azure, or GCP. • Knowledge of AI/ML security and adversarial testing methods,... 
    Second job
    Live in
    Worldwide
    Flexible hours

    AppCast

    New York, NY
    1 day ago
  •  ...Description Job Description Staff / Principal Platform Engineer Location: New York...  ...continuity. AppGate safeguards Fortune 500 enterprises and...  ...experience operationalizing AI/ML systems, and you treat...  ...shapes the architecture, reliability and core platform that defines... 
    For contractors
    Worldwide

    AppGate Cybersecurity, Inc.

    New York, NY
    25 days ago
  •  ...We’re looking for a Software Engineer to join the founding team to...  ...get hands-on with advanced AI/ML work. What You’ll Build...  ..., and observability to ensure reliability, performance, and cost control...  ...Background in developer tools, infra, or technical products Experience... 
    Full time

    Root Access

    New York, NY
    1 day ago
  •  ...re looking for incredible Senior Backend Engineers to help us build an ambitious roadmap ....  ...across the entire backend stack, from data infra to product endpoints Build extensible...  ...improve accuracy, speed, and cost of AI/ML models. Establish best practices for scalability... 
    Full time
    Work at office

    Bevel

    New York, NY
    1 day ago
  • $200k - $250k

     ...This innovative field blends AI, engineering, and materials science,...  ...everything secure, observable, and reliable in production.  We work in a...  ...-discipline domain—robotics, ML, and experimental automation—and...  .../Grafana, etc.). Hybrid infra (cloud + on-prem), containerization... 
    Full time

    Radical Ai

    New York, NY
    1 day ago
  • $1,000 per month

     ...the Role We're hiring a Senior Software Engineer to join the product engineering team. You'...  ...defined spec to start moving. Real AI/ML experience. You've shipped AI features to...  ...Typescript, AWS, Render, Snowflake DevOps/infra depth: CI/CD, monitoring, security... 
    Full time
    Work at office
    Remote work
    All shifts
    Flexible hours

    Evvy

    New York, NY
    1 day ago
  • $146.8k - $272.6k

     ...every day. As a Lead Software Engineer (Staff Engineer), you will own core...  ...that turn frontier models into reliable, production‑grade workflows...  ...cut across AI, product, and infra (e.g., a new orchestration layer...  ...closely with AI/ML engineers, researchers, designers... 
    Full time
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    New York, NY
    2 days ago
  •  ..., superior protection and seamless interoperability. AppGate safeguards Fortune 500 enterprises worldwide. Learn more at appgate.com. About the Role We're looking for a AI/ML Engineer (Senior/Staff/Principal) - Threat Detection who will design, build, and operationalize... 
    Full time
    For contractors
    Worldwide

    Appgate Cybersecurity, Inc.

    New York, NY
    1 day ago
  • $182k - $250.8k

     ...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great...  ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with... 
    Permanent employment
    Local area
    Remote work
    Worldwide
    Flexible hours
    Weekend work
    Weekday work

    Okta

    New York, NY
    4 days ago
  • $190k - $260k

     ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    3 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    4 days ago
  • $194k - $267k

     ...career-defining work. We're all in on this mission. If you are too, let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud... 
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    1 day ago
  • $194k - $267k

     ...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    New York, NY
    2 days ago
  • $200k - $400k

     ...power Decagon: networking, data, ML serving, developer platform,...  ...five focus areas: Core Infra: The foundational cloud stack...  ...infrastructure‑as‑code—to ensure reliability, scale, and cost efficiency....  ...hiring a Senior Infrastructure Engineer to design, build, and operate... 
    Full time
    Work at office
    Local area

    Decagon

    New York, NY
    1 day ago
  •  ...is a full-time role as a Senior SRE with a focus on Kubernetes, infra, security & platforms, located onsite in New York, NY.About...  ...our "hands off culture" by having very few meetings and giving engineers and creators a broad scope of responsibility and autonomy. You'... 
    Full time
    Work at office

    Hadrius

    New York, NY
    1 day ago
  • $130k - $250k

    What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people...  ...digital possibility? Begin your journey here.Securities Frontline Site Reliability Engineers (SREs) play a critical role in our fast-paced... 
    Full time
    Temporary work
    Part time
    Immediate start

    Goldman Sachs

    New York, NY
    2 days ago
  •  ...gap between our GTM team, core engineering team, and the unique, complex...  ...feedback loop to the product and infra teams-surfacing usability gaps...  ...the latest developments in ML/AI This is an in person role...  ...approaches fail at reliably extracting information in complex... 
    Work at office
    Local area

    Reducto

    New York, NY
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff+ Site Reliability Engineer, Safeguards ML Infra. Be the first to apply!