Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - AI Infrastructure

$300k
Full-time

Are you looking for an exciting new opportunity? 

Join a stealth-mode hyperscale data center startup building a next-generation AI and cloud platform designed for startups and advanced research, powered by thousands of H100, H200, and B200 GPUs available on demand. Their platform supports everything from rapid experimentation to full-scale model training and inference, with flexible orchestration via Slurm, Kubernetes, or direct SSH access. 

This is a rare opportunity to work at the intersection of hyperscale infrastructure and AI, shaping the operational backbone of one of the largest GPU clusters in private deployment.  If you want to build and operate infrastructure for frontier AI workloads, automate systems at petascale, and be part of a founding engineering team, this is the place to do it.

Responsibilities:

  • Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
  • Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
  • Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilization, and data flow.
  • Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
  • Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.

Skills / Must Have:

  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
  • Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
  • Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
  • Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
  • Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.

Benefits:

  • Equity

Salary:

  • $300,000 gross per year
Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - AI Infrastructure in San Francisco, CA vacancy
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Suggested
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    16 hours ago
  • $232k - $319k

     ...Every Identity, from AI to HumanIdentity is the...  ...the trusted, neutral infrastructure that enables organizations...  ...with great people and reliable, cost-effective, and...  ...initiatives across SRE & Infrastructure organization...  ...of SRE and product engineering by developing robust... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    6 hours ago
  • $180.5k - $236.91k

     ...Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team....  ...technical domains such as DevOps, site reliability, and cloud best practices Lead the...  ...reimbursements. Artificial Intelligence (AI): Our AI Guidelines outline the... 
    Suggested
    Remote job
    Full time
    Work at office

    Oscar Health

    San Francisco, CA
    a month ago
  • $350k

     ...Join a rapidly growing AI infrastructure provider delivering large-...  ...infrastructure providers to build reliable, high-performance...  ...opportunity is for a Staff Site Reliability Engineer to lead the reliability of...  ...environments Proven Staff-level SRE or infrastructure... 
    Suggested
    Full time
    San Francisco, CA
    more than 2 months ago
  • $200k - $260k

     ...Nimble Nimble is an AI robotics company building...  ...team of the world's best engineers and operators. If you...  ...Software Engineer to join our Infrastructure Team , focused on building secure, reliable, and scalable...  ...security engineering, or SRE experience. Ability to... 
    Suggested
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    13 days ago
  • $207k - $300k

     ...changes that improve reliability and velocity....  ...strategy for Home SRE.Minimum qualifications...  ...Science or Engineering, or a related field...  ...ML model serving infrastructure, real-time low-latency...  ...engineering organizations.Site Reliability...  ...generation real-time Voice AI. We focus on high-... 
    Worldwide

    Google

    San Francisco, CA
    3 days ago
  • $80 per hour

     ...Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy...  ...Linux systems administration, infrastructure monitoring, incident response, programming...  ...developing or deploying Agentic AI or autonomous automation tools to streamline... 
    Hourly pay
    Full time
    Work at office
    Local area
    Shift work
    Night shift

    Essnova Solutions

    Berkeley, CA
    1 day ago
  •  ...Description Developer & Infrastructure Expert Role Type:...  ...to evaluate AI-powered workflows across...  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-...  ...for accuracy and reliability. Work with AWS,...  ...Cloud Infrastructure Site Reliability... 
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    a month ago
  •  ...and Amsterdam.Plaid's Infrastructure team builds the...  ...and tooling that help engineering teams develop, deploy...  ...product team.As a Staff Site Reliability Engineer on Release Engineering...  ...and safe even as AI-assisted development...  ...in backend systems, SRE, or platform... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  • $190.8k - $267.1k

     ...its business. The reliability of our Ads systems...  ...closely with Ads Engineering to improve reliability...  ...for a Senior Site Reliability Engineer...  ...build, and maintain infrastructure, tooling, and...  ...Drive adoption of SRE best practices including...  ...intelligence (AI). You will have the... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...SalesforceSalesforce is the #1 AI CRM, where humans with...  ...is seeking a senior engineering candidate to join the Site Reliability organization in San...  ...with counterparts in the Infrastructure and R&D organizations, this...  ...protected. The ExperienceAs an SRE, you will be a technical... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  • $113.4k - $162k

     ...everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...and operates its systems in an AI-first environment where intelligent...  ...the design and implementation of new SRE best practices.You'll be a great fit... 
    Temporary work

    TextNow

    San Francisco, CA
    6 hours ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support...  ...that ensure cluster reliability and security (e.g., CoreDNS...  ...data platform for the AI era, enabling builders to... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  •  ...combination of proprietary infrastructure and software, we...  ...-to-end. You use AI to work smarter...  ...About the teamThe Engineering team at Airwallex...  ...build scalable, reliable, and secure products...  ...borders.Our SRE team is breaking new...  ...’ll doAs a Senior Site Reliability Engineer... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    2 days ago
  • $165k - $227k

     ...Every Identity, from AI to HumanIdentity is the...  ...the trusted, neutral infrastructure that enables...  ...too, let's talk.The Engineering OpportunityWe are looking...  ...an experienced Senior Site Reliability Engineer to join Okta...  ...contributor within the EPG SRE organization,... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  •  ...mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase...  ...the Enterprise Technology, Infrastructure Platforms team, you will solve...  ...tuningUses enterprise-authorized AI capabilities within the...  ...experience and practical SRE/observability fundamentals... 

    JP Morgan Chase

    San Francisco, CA
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB’s broader engineering...  ...in engineering the reliable, globally connected,...  ...a talented Senior Site Reliability Engineer (SRE...  ...data platform for the AI era, enabling builders... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    2 days ago
  • $230k - $390k

     ...customer experiences with AI. We are primarily an in-person...  ...you'll doAs a Software Engineer on our Site Reliability team at Sierra, you will...  ...across Sierra’s AI-driven infrastructure. You’ll partner closely with...  ....Define the foundation of SRE practices at Sierra,... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    2 days ago
  •  ...world's most dynamic AI companies, like Cursor...  ...AI research, flexible infrastructure, and seamless developer...  ...help build the platform engineers turn to to ship AI...  ...high performance and reliability. You’ll own scheduling...  ...Partner closely with SRE and Capacity teams to... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...About Us At Hayden AI, we are on a mission...  ...world challenges. The Infrastructure Engineering team is crucial to...  ...performance, maximum reliability, and cost-efficiency...  ...modeling best practices in site reliability,...  ...in modern DevOps and SRE practices, including... 
    Full time
    Shift work

    Hayden Ai

    San Francisco, CA
    more than 2 months ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key...  ...building the trusted, neutral infrastructure that enables organizations...  ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk...  ...Platform that enables our SRE teams and business partners.... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $217k - $303.9k

     ...continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring...  ...by artificial intelligence (AI). You will have the... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    3 days ago
  • $194k - $267k

    Secure Every Identity, from AI to HumanIdentity is the key to unlocking the...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $204k - $306k

     ...Every Identity, from AI to HumanIdentity is the...  ...the trusted, neutral infrastructure that enables...  ...let's talk.Manager, Site Reliability EngineeringSan Francisco...  ...IDaaS Site Reliability Engineering GroupOkta authenticates...  ...doing Managing a team of SRE’s supporting various... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    2 days ago
  • $195k - $257.5k

     ...and programmable blockchain infrastructure. Circle’s platform includes...  ...responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team,...  ...Infrastructure as Code, and modern SRE practices to deliver...  ...teams, you’ll build AI-powered tooling and automation... 
    Flexible hours

    Circle

    San Francisco, CA
    6 hours ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA...  ...Formation Bio is a tech and AI driven pharma company differentiated...  ...will build and operate the infrastructure, delivery systems, and...  ...on infrastructure and SRE fundamentals. About You... 
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    San Francisco, CA
    17 hours ago
  •  ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated access to business data across...  ...in infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    2 days ago
  •  ...now part of Superhuman, the AI productivity platform on a mission...  ...goals, we’re looking for an SRE to join our infrastructure team. This role will be...  ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    6 hours ago
  • $181k - $225k

     ...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming...  ...management industry by building an AI platform for wealth...  ...to hire a high performing SRE to join our growing Platform...  ...contribute to solutions with shared infrastructure that benefit all of... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    4 days ago
  • $130k - $200k

     ...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through...  ...The Role This is a career-level SRE role for someone who wants to own systems... 
    Shift work

    Nscale

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Infrastructure. Be the first to apply!