Staff Site Reliability Engineer - AI Infrastructure
$350kLooking for a role with plenty of growth opportunities?
Join a rapidly growing AI infrastructure provider delivering large-scale compute solutions for AI training and inference across global cloud and GPU environments. The organization partners with leading AI companies and infrastructure providers to build reliable, high-performance platforms supporting next-generation AI workloads.
This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale GPU infrastructure, covering deployment, GPU health, distributed training, networking, and incident response. Working as a senior technical leader, the role focuses on improving infrastructure reliability, operational standards, and platform performance across complex AI environments.
Ready to make a move? Get in touch and apply today!
Responsibilities:
- Lead high-priority incident response across distributed GPU infrastructure environments
- Diagnose and resolve issues across the stack, including PyTorch, NCCL, CUDA, drivers, networking fabrics, and hardware layers
- Own day-to-day operational health of large-scale GPU fleets, including lifecycle management, validation, firmware rollouts, upgrades, and repair workflows
- Build and maintain observability systems, GPU telemetry platforms, automated remediation tooling, and health-check frameworks
- Define and scale operational practices, including on-call rotations, escalation processes, incident response, and postmortem standards
- Partner closely with infrastructure, product, and platform engineering teams to improve reliability and scalability
- Participate in customer-facing technical discussions, incident reviews, architecture workshops, and workload planning sessions
- Influence physical infrastructure design, including rack layouts, power density, cooling strategies, burn-in procedures, and network topology decisions
- Mentor engineers across reliability engineering, systems operations, and incident management practices
- Contribute to a long-term reliability strategy for hyperscale AI infrastructure environments
Skills/Must Have:
- Multiple years of hands-on experience operating large-scale GPU infrastructure environments
- Proven Staff-level SRE or infrastructure engineering experience supporting mission-critical production systems
- Deep expertise with NVIDIA GPU platforms including H100, H200, B200, or GB200 systems
- Strong understanding of GPU memory hierarchy, ECC behaviour, NVLink, NVSwitch, thermal management, and hardware failure analysis
- Production experience with InfiniBand, RoCE, and high-performance distributed training fabrics
- Deep understanding of distributed AI training technologies, including NCCL, CUDA, PyTorch Distributed, FSDP, DeepSpeed, and Megatron
- Strong software engineering skills in Go, Python, or Rust
- Experience building production-grade automation, tooling, fleet management systems, or reliability platforms
- Hands-on experience with Kubernetes GPU environments, Slurm, or HPC schedulers
- Strong Linux systems expertise, including kernel tuning, CUDA lifecycle management, cgroups/namespaces, BPF/performance analysis, and firmware operations
- Calm, structured incident response capabilities within high-pressure production environments
- Ability to communicate effectively with highly technical customers, providers, and executive stakeholders
Desirable Skills:
- Experience building custom GPU fleet health systems or fabric controllers
- Expertise with distributed storage systems such as VAST, Weka, Lustre, or GPFS
- Experience optimizing distributed training efficiency, checkpointing, and multi-thousand-GPU job performance
- Background supporting enterprise AI infrastructure customers in customer-facing technical roles
- Open-source contributions within the GPU, Kubernetes, or AI infrastructure ecosystem
- Public speaking, technical writing, or community leadership within AI infrastructure or HPC domain
Benefits:
- Huge stock options
- Company bonus
- Unlimited PTO
- 401K + 4% match
Salary:
- $350,000 base salary
$127k - $249k
...are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on... ...have redefined the data platform for the AI era, enabling builders to create, transform...SuggestedLocal areaRemote workWorldwideFlexible hours- ...UsAlembic is the pioneering Causal AI platform. We help the world's... ...built on Grace Blackwell infrastructure — one of the fastest private supercomputers... ...under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the...Suggested
$157k - $239k
...Francisco, CA / Golden, COInfrastructure - Cloud Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud... ...observation, IoT connectivity, on-orbit AI, national security missions, and more. Leveraging...SuggestedFull timeTemporary work$215k - $275k
...Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale,... ...date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of...SuggestedWork at office- ...London and Amsterdam.Plaid's Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and... ...for every product team.As a Staff Site Reliability Engineer on Release Engineering... ...fast and safe even as AI-assisted development increases...SuggestedPermanent employmentWork experience placementWork at officeLocal area
$148.5k - $223.9k
...DetailsAbout SalesforceSalesforce is the #1 AI CRM, where humans with agents... ....Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco.... ...closely with counterparts in the Infrastructure and R&D organizations, this...Full timeWorldwideWeekend work$152.5k - $205k
...applications, and programmable blockchain infrastructure. Circle’s platform includes the world... ...’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll... ...infrastructure behind critical digital-assets, AI, and application workloads. You will...Flexible hours$113.4k - $162k
...for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and... ...shape how TextNow builds and operates its systems in an AI-first environment where intelligent tooling is embedded...Temporary work$167.7k - $245.2k
...approximately 2 days per week on-site at Cisco offices in... ...the future of AI resilience. Our team provides... ...intended, improving reliability and reducing risks.... ...Site Reliability Engineer (SRE), you will build,... ...platform and production infrastructure. You will own the operational...Full timeTemporary workLocal areaFlexible hours2 days per week$114.3k - $235.32k
...you love? It’s Possible.At Pinterest, AI isn't just a feature, it's a powerful... ...grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and... ...observability, and operational maturity of our infrastructure and delivery ecosystem. The ideal...Work at officeLocal areaRelocationRelocation package- ...combination of proprietary infrastructure and software, we... ...” end-to-end. You use AI to work smarter and solve... ....About the teamThe Engineering team at Airwallex is a... ...together to build scalable, reliable, and secure products... ...you’ll doAs a Senior Site Reliability Engineer,...Temporary workLocal areaWorldwide
$165k - $225.6k
Secure Every Identity, from AI to HumanIdentity is the key... ...building the trusted, neutral infrastructure that enables organizations... ...across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the...Permanent employmentLocal areaWorldwideFlexible hours- ...is now part of Superhuman, the AI productivity platform on a mission... ...looking for an SRE to join our infrastructure team. This role will be responsible... ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning...WorldwideHome officeFlexible hours
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible... ...for a range of critical infrastructure and operational functions... ...that ensure cluster reliability and security (e.g., CoreDNS,... ...redefined the data platform for the AI era, enabling builders to...Work at officeLocal areaRemote workWorldwideFlexible hours$165k - $241.4k
...don’t own. Powered by AI and an unmatched set of... ...the Federal region’s infrastructure and operations, such as... ...looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and... ...see the Cisco careers site to discover more benefits...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$194k - $267k
Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential... ...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$186.9k - $267.7k
...2 days per week on-site at Cisco offices in... ...defining the future of AI resilience. Together... ...intended, improving reliability and reducing risks.... ...observability and control.As a Staff Site Reliability Engineer (SRE), you will... ..., lead major infrastructure initiatives, and drive...Full timeTemporary workLocal areaFlexible hours2 days per week$174k - $239k
Secure Every Identity, from AI to HumanIdentity is the key... ...building the trusted, neutral infrastructure that enables organizations... ...across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc...Work experience placementLocal areaWorldwideFlexible hours$194k - $267k
Secure Every Identity, from AI to HumanIdentity is the key to unlocking the potential... ...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
Secure Every Identity, from AI to HumanIdentity is the key... ...the trusted, neutral infrastructure that enables organizations... ...too, let's talk.The TeamThe Site Reliability team is dedicated to architecting... ...platform reliability and engineering velocity.The ideal candidate...Local areaWorldwideFlexible hours$197.3k - $313.7k
...SalesforceSalesforce is the #1 AI CRM, where humans with... ...Salesforce.Job Title: Director, Site Reliability EngineeringLocation: New... ...of Site Reliability Engineering to spearhead the evolution... ...Platform, Architecture, Security, Infrastructure, and Product, you will ensure...Full timeImmediate start$220k - $235k
Ironclad is the leading AI contracting platform that transforms... ...a strategic, high-output Staff/Senior Staff SRE to define the... ...cloud platform and champion engineering excellence across Ironclad. In... ...strategic direction for the Site Reliability Engineering team and our broader...Full timeContract workWork at office$204k - $306k
Secure Every Identity, from AI to HumanIdentity is the key... ...the trusted, neutral infrastructure that enables organizations... ...are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure... ...The IDaaS Site Reliability Engineering GroupOkta authenticates,...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$195k - $257.5k
...applications, and programmable blockchain infrastructure. Circle’s platform includes the... ...What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll... ...product, and security teams, you’ll build AI-powered tooling and automation that...Flexible hours- ...identity security, delivering an AI-powered platform that... ...cloud-native systems. As a Staff Platform Engineer, you will play a critical role... ...role. You will own reliability for major platform domains,... ...and maintaining the shared infrastructure services and platforms that...
$165k - $241.4k
...those beyond their ownership. Leveraging AI and an unparalleled set of cloud,... ...ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with... ...reliability.Manage a rapidly growing infrastructure capable of handling substantial daily...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural... ...and namespace management, raise reliability, and contribute to upstream Kubernetes.... ...The role includes mentoring and building AI-assisted tooling for triage and automation...
$175k - $250k
...Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,... ...unavailable. Modality: On-Site only. Must live within commuting... ...interact with generative AI. They are the team behind a... ..., performance, and reliability across environments. What...Full timeRemote workRelocationRelocation package$150k - $200k
...Runpod is the AI Developer Cloud. More than one... ...inflection point for AI infrastructure, and we're building... ...announcement: . The Reliability team owns the availability... ...standards across engineering Designing incident response... ...systems. As a Site Reliability Engineer on...Remote workVisa sponsorshipWork visaFlexible hours- ...developers, platform engineers, and IT staff to improve system design... ..., service quality, reliability, security, and... ...years of experience in Site Reliability Engineering... ...Engineering, or related infrastructure roles supporting... ...Experience managing AI tooling on Kubernetes...Work at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer - AI Infrastructure. Be the first to apply!
- software engineer staff San Francisco, CA
- assistant engineer San Francisco, CA
- engineering aide San Francisco, CA
- staff engineer San Francisco, CA
- staff security engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- technology administrator San Francisco, CA
- senior staff engineer San Francisco, CA
- staff data engineer San Francisco, CA
- staff design engineer San Francisco, CA

