Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer - AI Infrastructure

$350k
Full-time

Looking for a role with plenty of growth opportunities?

Join a rapidly growing AI infrastructure provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The organization partners with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads.

This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering deployment, GPU health, distributed training, networking, and incident response. Working as a senior technical leader, the role focuses on improving infrastructure reliability, operational standards, and platform performance across complex AI environments.

Ready to make a move? Get in touch and apply today!

Responsibilities:

  • Lead high-priority incident response across distributed GPU infrastructure environments
  • Diagnose and resolve issues across the stack, including PyTorch, NCCL, CUDA, drivers, networking fabrics, and hardware layers
  • Own day-to-day operational health of large-scale GPU fleets, including lifecycle management, validation, firmware rollouts, upgrades, and repair workflows
  • Build and maintain observability systems, GPU telemetry platforms, automated remediation tooling, and health-check frameworks
  • Define and scale operational practices, including on-call rotations, escalation processes, incident response, and postmortem standards
  • Partner closely with infrastructure, product, and platform engineering teams to improve reliability and scalability
  • Participate in customer-facing technical discussions, incident reviews, architecture workshops, and workload planning sessions
  • Influence physical infrastructure design, including rack layouts, power density, cooling strategies, burn-in procedures, and network topology decisions
  • Mentor engineers across reliability engineering, systems operations, and incident management practices
  • Contribute to a long-term reliability strategy for hyperscale AI infrastructure environments

Skills/Must Have:

  • Multiple years of hands-on experience operating large-scale GPU infrastructure environments
  • Proven Staff-level SRE or infrastructure engineering experience supporting mission-critical production systems
  • Deep expertise with NVIDIA GPU platforms including H100, H200, B200, or GB200 systems
  • Strong understanding of GPU memory hierarchy, ECC behaviour, NVLink, NVSwitch, thermal management, and hardware failure analysis
  • Production experience with InfiniBand, RoCE, and high-performance distributed training fabrics
  • Deep understanding of distributed AI training technologies, including NCCL, CUDA, PyTorch Distributed, FSDP, DeepSpeed, and Megatron
  • Strong software engineering skills in Go, Python, or Rust
  • Experience building production-grade automation, tooling, fleet management systems, or reliability platforms
  • Hands-on experience with Kubernetes GPU environments, Slurm, or HPC schedulers
  • Strong Linux systems expertise, including kernel tuning, CUDA lifecycle management, cgroups/namespaces, BPF/performance analysis, and firmware operations
  • Calm, structured incident response capabilities within high-pressure production environments
  • Ability to communicate effectively with highly technical customers, providers, and executive stakeholders

Desirable Skills:

  • Experience building custom GPU fleet health systems or fabric controllers
  • Expertise with distributed storage systems such as VAST, Weka, Lustre, or GPFS
  • Experience optimizing distributed training efficiency, checkpointing, and multi-thousand-GPU job performance
  • Background supporting enterprise AI infrastructure customers in customer-facing technical roles
  • Open-source contributions within the GPU, Kubernetes, or AI infrastructure ecosystem
  • Public speaking, technical writing, or community leadership within AI infrastructure or HPC domain

Benefits:

  • Huge stock options 
  • Company bonus 
  • Unlimited PTO
  • 401K + 4% match

Salary:

  • $350,000 base salary

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer - AI Infrastructure in San Francisco, CA vacancy
  • $127k - $249k

     ...are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on...  ...have redefined the data platform for the AI era, enabling builders to create, transform... 
    Suggested
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  •  ...UsAlembic is the pioneering Causal AI platform. We help the world's...  ...built on Grace Blackwell infrastructure — one of the fastest private supercomputers...  ...under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the... 
    Suggested

    Alembic

    San Francisco, CA
    1 day ago
  • $157k - $239k

     ...Francisco, CA / Golden, COInfrastructure - Cloud Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud...  ...observation, IoT connectivity, on-orbit AI, national security missions, and more. Leveraging... 
    Suggested
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    2 days ago
  • $215k - $275k

     ...Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale,...  ...date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    1 day ago
  •  ...London and Amsterdam.Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and...  ...for every product team.As a Staff Site Reliability Engineer on Release Engineering...  ...fast and safe even as AI-assisted development increases... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  • $148.5k - $223.9k

     ...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents...  ....Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco....  ...closely with counterparts in the Infrastructure and R&D organizations, this... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • $152.5k - $205k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the world...  ...’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll...  ...infrastructure behind critical digital-assets, AI, and application workloads. You will... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and...  ...shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded... 
    Temporary work

    TextNow

    San Francisco, CA
    2 days ago
  • $167.7k - $245.2k

     ...approximately 2 days per week on-site at Cisco offices in...  ...the future of AI resilience. Our team provides...  ...intended, improving reliability and reducing risks....  ...Site Reliability Engineer (SRE), you will build,...  ...platform and production infrastructure. You will own the operational... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $114.3k - $235.32k

     ...you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful...  ...grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and...  ...observability, and operational maturity of our infrastructure and delivery ecosystem. The ideal... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    22 hours ago
  •  ...combination of proprietary infrastructure and software, we...  ...” end-to-end. You use AI to work smarter and solve...  ....About the teamThe Engineering team at Airwallex is a...  ...together to build scalable, reliable, and secure products...  ...you’ll doAs a Senior Site Reliability Engineer,... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  •  ...is now part of Superhuman, the AI productivity platform on a mission...  ...looking for an SRE to join our infrastructure team. This role will be responsible...  ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible...  ...for a range of critical infrastructure and operational functions...  ...that ensure cluster reliability and security (e.g., CoreDNS,...  ...redefined the data platform for the AI era, enabling builders to... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $165k - $241.4k

     ...don’t own. Powered by AI and an unmatched set of...  ...the Federal region’s infrastructure and operations, such as...  ...looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and...  ...see the Cisco careers site to discover more benefits... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $186.9k - $267.7k

     ...2 days per week on-site at Cisco offices in...  ...defining the future of AI resilience. Together...  ...intended, improving reliability and reducing risks....  ...observability and control.As a Staff Site Reliability Engineer (SRE), you will...  ..., lead major infrastructure initiatives, and drive... 
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    4 days ago
  • $174k - $239k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...the trusted, neutral infrastructure that enables organizations...  ...too, let's talk.The TeamThe Site Reliability team is dedicated to architecting...  ...platform reliability and engineering velocity.The ideal candidate... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $197.3k - $313.7k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New...  ...of Site Reliability Engineering to spearhead the evolution...  ...Platform, Architecture, Security, Infrastructure, and Product, you will ensure... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    3 days ago
  • $220k - $235k

    Ironclad is the leading AI contracting platform that transforms...  ...a strategic, high-output Staff/Senior Staff SRE to define the...  ...cloud platform and champion engineering excellence across Ironclad. In...  ...strategic direction for the Site Reliability Engineering team and our broader... 
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    3 days ago
  • $204k - $306k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...the trusted, neutral infrastructure that enables organizations...  ...are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure...  ...The IDaaS Site Reliability Engineering GroupOkta authenticates,... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    3 days ago
  • $195k - $257.5k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the...  ...What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll...  ...product, and security teams, you’ll build AI-powered tooling and automation that... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  •  ...identity security, delivering an AI-powered platform that...  ...cloud-native systems. As a Staff Platform Engineer, you will play a critical role...  ...role. You will own reliability for major platform domains,...  ...and maintaining the shared infrastructure services and platforms that... 

    Saviynt

    San Francisco, CA
    a month ago
  • $165k - $241.4k

     ...those beyond their ownership. Leveraging AI and an unparalleled set of cloud,...  ...ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with...  ...reliability.Manage a rapidly growing infrastructure capable of handling substantial daily... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    22 hours ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural...  ...and namespace management, raise reliability, and contribute to upstream Kubernetes....  ...The role includes mentoring and building AI-assisted tooling for triage and automation... 

    Socket

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,...  ...unavailable. Modality: On-Site only. Must live within commuting...  ...interact with generative AI. They are the team behind a...  ..., performance, and reliability across environments. What... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    13 hours ago
  • $150k - $200k

     ...Runpod is the AI Developer Cloud. More than one...  ...inflection point for AI infrastructure, and we're building...  ...announcement: . The Reliability team owns the availability...  ...standards across engineering Designing incident response...  ...systems. As a Site Reliability Engineer on... 
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    GrabJobs

    San Francisco, CA
    4 days ago
  •  ...developers, platform engineers, and IT staff to improve system design...  ..., service quality, reliability, security, and...  ...years of experience in Site Reliability Engineering...  ...Engineering, or related infrastructure roles supporting...  ...Experience managing AI tooling on Kubernetes... 
    Work at office
    Remote work

    GrabJobs

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer - AI Infrastructure. Be the first to apply!