Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer - AI Infrastructure

$350k
Full-time

Looking for a role with plenty of growth opportunities?

Join a rapidly growing AI infrastructure provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The organization partners with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads.

This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering deployment, GPU health, distributed training, networking, and incident response. Working as a senior technical leader, the role focuses on improving infrastructure reliability, operational standards, and platform performance across complex AI environments.

Ready to make a move? Get in touch and apply today!

Responsibilities:

  • Lead high-priority incident response across distributed GPU infrastructure environments
  • Diagnose and resolve issues across the stack, including PyTorch, NCCL, CUDA, drivers, networking fabrics, and hardware layers
  • Own day-to-day operational health of large-scale GPU fleets, including lifecycle management, validation, firmware rollouts, upgrades, and repair workflows
  • Build and maintain observability systems, GPU telemetry platforms, automated remediation tooling, and health-check frameworks
  • Define and scale operational practices, including on-call rotations, escalation processes, incident response, and postmortem standards
  • Partner closely with infrastructure, product, and platform engineering teams to improve reliability and scalability
  • Participate in customer-facing technical discussions, incident reviews, architecture workshops, and workload planning sessions
  • Influence physical infrastructure design, including rack layouts, power density, cooling strategies, burn-in procedures, and network topology decisions
  • Mentor engineers across reliability engineering, systems operations, and incident management practices
  • Contribute to a long-term reliability strategy for hyperscale AI infrastructure environments

Skills/Must Have:

  • Multiple years of hands-on experience operating large-scale GPU infrastructure environments
  • Proven Staff-level SRE or infrastructure engineering experience supporting mission-critical production systems
  • Deep expertise with NVIDIA GPU platforms including H100, H200, B200, or GB200 systems
  • Strong understanding of GPU memory hierarchy, ECC behaviour, NVLink, NVSwitch, thermal management, and hardware failure analysis
  • Production experience with InfiniBand, RoCE, and high-performance distributed training fabrics
  • Deep understanding of distributed AI training technologies, including NCCL, CUDA, PyTorch Distributed, FSDP, DeepSpeed, and Megatron
  • Strong software engineering skills in Go, Python, or Rust
  • Experience building production-grade automation, tooling, fleet management systems, or reliability platforms
  • Hands-on experience with Kubernetes GPU environments, Slurm, or HPC schedulers
  • Strong Linux systems expertise, including kernel tuning, CUDA lifecycle management, cgroups/namespaces, BPF/performance analysis, and firmware operations
  • Calm, structured incident response capabilities within high-pressure production environments
  • Ability to communicate effectively with highly technical customers, providers, and executive stakeholders

Desirable Skills:

  • Experience building custom GPU fleet health systems or fabric controllers
  • Expertise with distributed storage systems such as VAST, Weka, Lustre, or GPFS
  • Experience optimizing distributed training efficiency, checkpointing, and multi-thousand-GPU job performance
  • Background supporting enterprise AI infrastructure customers in customer-facing technical roles
  • Open-source contributions within the GPU, Kubernetes, or AI infrastructure ecosystem
  • Public speaking, technical writing, or community leadership within AI infrastructure or HPC domain

Benefits:

  • Huge stock options 
  • Company bonus 
  • Unlimited PTO
  • 401K + 4% match

Salary:

  • $350,000 base salary

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer - AI Infrastructure in San Francisco, CA vacancy
  • $127k - $249k

     ...are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on...  ...have redefined the data platform for the AI era, enabling builders to create, transform... 
    Suggested
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  • $157k - $239k

     ...Francisco, CA / Golden, COInfrastructure - Cloud Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud...  ...observation, IoT connectivity, on-orbit AI, national security missions, and more. Leveraging... 
    Suggested
    Full time
    Temporary work

    Loft Orbital

    San Francisco, CA
    2 days ago
  •  ...UsAlembic is the pioneering Causal AI platform. We help the world's...  ...built on Grace Blackwell infrastructure — one of the fastest private supercomputers...  ...under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the... 
    Suggested

    Alembic

    San Francisco, CA
    1 day ago
  • $167.7k - $245.2k

     ...on our products to run their critical infrastructure of network switches, security...  ...these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a...  ...automation, project management, and AI-assisted engineering tools such as Codex... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $232k - $319k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...service with great people and reliable, cost-effective, and...  ...velocity of SRE and product engineering by developing robust platforms... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  •  ...London and Amsterdam.Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and...  ...for every product team.As a Staff Site Reliability Engineer on Release Engineering...  ...fast and safe even as AI-assisted development increases... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  •  ...is now part of Superhuman, the AI productivity platform on a mission...  ...looking for an SRE to join our infrastructure team. This role will be responsible...  ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    4 days ago
  • $167.7k - $245.2k

     ...don’t own. Powered by AI and an unmatched set of...  ...the Federal region’s infrastructure and operations, such as...  ...looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and...  ...see the Cisco careers site to discover more benefits... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible...  ...for a range of critical infrastructure and operational functions...  ...that ensure cluster reliability and security (e.g., CoreDNS,...  ...redefined the data platform for the AI era, enabling builders to... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $139.76k - $287.75k

     ...Possible.At Pinterest, AI isn't just a feature,...  ...are seeking a Senior Site ReliabilityEngineer to...  ...instrumental in advancing the reliability, scalability,...  ...operational maturity of our infrastructure and delivery ecosystem...  ...is a highly hands-on engineer with strong production... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    3 days ago
  •  ...complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and...  ...database tuningUses enterprise-authorized AI capabilities within the work environment to... 

    JP Morgan Chase

    San Francisco, CA
    15 hours ago
  • $165k - $225.6k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents...  ....Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco....  ...closely with counterparts in the Infrastructure and R&D organizations, this... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • $152.5k - $205k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the world...  ....What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform...  ...help teams safely adopt and operate AI-enabled tooling and workflows by... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the world...  ...’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll...  ...infrastructure behind critical digital-assets, AI, and application workloads. You will... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and...  ...shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded... 
    Temporary work

    TextNow

    San Francisco, CA
    2 days ago
  •  ...combination of proprietary infrastructure and software, we...  ...” end-to-end. You use AI to work smarter and solve...  ....About the teamThe Engineering team at Airwallex is a...  ...together to build scalable, reliable, and secure products...  ...you’ll doAs a Senior Site Reliability Engineer,... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  • $190.8k - $267.1k

     ...grow its business. The reliability of our Ads systems...  ...partners closely with Ads Engineering teams to improve...  ....We're looking for a Staff Site Reliability Engineer...  ...shape the future of Ads infrastructure at Reddit.What you’ll...  ...intelligence (AI). You will have the opportunity... 
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    San Francisco, CA
    3 days ago
  • $165k - $227k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $220k - $235k

    Ironclad is the leading AI contracting platform that transforms...  ...a strategic, high-output Staff/Senior Staff SRE to define the...  ...cloud platform and champion engineering excellence across Ironclad. In...  ...strategic direction for the Site Reliability Engineering team and our broader... 
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    3 days ago
  • $217k - $303.9k

     ...to scale globally, reliability and performance are...  ...critical than ever. The Site Experience SRE team...  ...the intersection of infrastructure, product engineering, and user experience...  ...are looking for a Staff Site Reliability Engineer...  ...intelligence (AI). You will have the... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    4 days ago
  • $174k - $239k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $195k - $257.5k

     ...applications, and programmable blockchain infrastructure. Circle’s platform includes the...  ...What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll...  ...product, and security teams, you’ll build AI-powered tooling and automation that... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $204k - $306k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...the trusted, neutral infrastructure that enables organizations...  ...are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure...  ...The IDaaS Site Reliability Engineering GroupOkta authenticates,... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    4 days ago
  • $207k - $300k

     ...changes that improve reliability and velocity.Practice...  ...in Computer Science or Engineering, or a related field.Experience...  ...in ML model serving infrastructure, real-time low-latency...  ...organizations.Site Reliability Engineering...  ...like the Gemini-powered AI Coach and Fitbit hardware... 

    Google

    San Francisco, CA
    4 days ago
  • $167.7k - $245.2k

     ...those beyond their ownership. Leveraging AI and an unparalleled set of cloud,...  ...ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with...  ...reliability.Manage a rapidly growing infrastructure capable of handling substantial daily... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  •  ...the TeamUDAIP is U.S. Bank next generation cloud-based data and AI platform to enable data success across business units within...  ...high level, UDAIP success includes industry leading hybrid cloud infrastructure, innovative data capabilities with leverage of many open-source... 

    Phenom People

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer - AI Infrastructure. Be the first to apply!