Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - AI Infrastructure

$300k
Full-time

Are you looking for an exciting new opportunity? 

Join a stealth-mode hyperscale data center startup building a next-generation AI and cloud platform designed for startups and advanced research, powered by thousands of H100, H200, and B200 GPUs available on demand. Their platform supports everything from rapid experimentation to full-scale model training and inference, with flexible orchestration via Slurm, Kubernetes, or direct SSH access. 

This is a rare opportunity to work at the intersection of hyperscale infrastructure and AI, shaping the operational backbone of one of the largest GPU clusters in private deployment.  If you want to build and operate infrastructure for frontier AI workloads, automate systems at petascale, and be part of a founding engineering team, this is the place to do it.

Responsibilities:

  • Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
  • Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
  • Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
  • Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilization, and data flow.
  • Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
  • Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.

Skills / Must Have:

  • 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
  • Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
  • Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
  • Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
  • Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
  • Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
  • Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.

Benefits:

  • Equity

Salary:

  • $300,000 gross per year
Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - AI Infrastructure in San Francisco, CA vacancy
  •  ...mission to create the world's first AI-powered Personal & Entrepreneurial Resource...  ...the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be...  ...Grafana, Datadog, ELK). Automate infrastructure provisioning, deployment, and incident... 
    Suggested
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    2 days ago
  • $350k

     ...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity...  ...to the knowledge and tools to make AI work for their unique needs and...  ...in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility... 
    Suggested
    Local area
    Visa sponsorship
    Work visa
    Relocation package

    Thinking Machines Lab

    San Francisco, CA
    16 hours ago
  • $232k - $319k

     ...Every Identity, from AI to Human Identity is...  ...the trusted, neutral infrastructure that enables organizations...  ...with great people and reliable, cost-effective, and...  ...various initiatives across SRE & Infrastructure...  ...velocity of SRE and product engineering by developing robust... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    1 day ago
  • $350k

     ...Join a rapidly growing AI infrastructure provider delivering large-...  ...infrastructure providers to build reliable, high-performance...  ...opportunity is for a Staff Site Reliability Engineer to lead the reliability of...  ...environments Proven Staff-level SRE or infrastructure... 
    Suggested
    Permanent employment
    San Francisco, CA
    25 days ago
  •  ...the intersection of labor markets and AI research. We partner with leading AI...  ...headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability...  ...systems, partnering directly with infrastructure leadership. You’ll play a... 
    Suggested

    Mercor Inc

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...Category Software Engineering Job Details About...  ...Salesforce is the #1 AI CRM, where humans...  ...candidate to join the Site Reliability organization in San Francisco...  ...counterparts in the Infrastructure and R&D organizations...  ...As an SRE, you will be a technical... 
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    5 days ago
  • $260k - $300k

     ...We are an applied AI lab building end-to-end software...  ...the first AI software engineer. Our team is...  ...both the production reliability of our user-facing...  ...pipelines, deployment infrastructure, and developer tooling...  ...engineering fundamentals; SRE at Cognition means... 

    Cognition Corp

    San Francisco, CA
    5 days ago
  •  ...Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates...  ...software engineering and applies them to infrastructure and operations problems. The main...  ...in Python supported by Gen AI tooling to accelerate development of... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    2 days ago
  •  ...now part of Superhuman, the AI productivity platform on a mission...  ...goals, we're looking for an SRE to join our infrastructure team. This role will be...  ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    4 days ago
  • $155k - $222.6k

     ...Meet the Team The SRE Fleet team is...  ...efficiency of the infrastructure that powers our global...  ...As a team of six engineers distributed across...  ...on automation, reliability, and operational excellence...  ...of experience in Site Reliability...  ...Experience leveraging AI-assisted... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    San Francisco, CA
    4 days ago
  •  ...world's most dynamic AI companies, like Cursor...  ...AI research, flexible infrastructure, and seamless developer...  ...help build the platform engineers turn to to ship AI...  ...high performance and reliability. You’ll own scheduling...  ...Partner closely with SRE and Capacity teams to... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge...  ...AI research, flexible infrastructure, and seamless developer...  ...these as part of the SRE team: Improve Baseten... 
    Flexible hours

    Baseten

    San Francisco, CA
    5 days ago
  •  ...We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating...  ...Required skills ~3+ years in SRE, DevOps, or infrastructure engineering... 

    Blaxel, Inc

    San Francisco, CA
    3 days ago
  • $120k - $168.49k

     ...Site Reliability Engineer, Cloud Infrastructure About Quizlet At Quizlet, our mission is to help every learner achieve...  ...looking for a Site Reliability Engineer (SRE) to join our infrastructure team...  .... Cutting‑edge tech: Generative AI, adaptive learning, cognitive science... 
    Internship
    Work at office
    3 days per week

    Quizlet

    San Francisco, CA
    1 day ago
  • $166.9k - $225.9k

     ...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be...  ...reliability. Our infrastructure runs on AWS across...  ...years of experience in Site Reliability...  ...with AIOps - using AI/ML-based tooling for... 
    Work at office
    Immediate start
    Worldwide
    Monday to Friday
    Flexible hours

    Drata Inc

    San Francisco, CA
    2 days ago
  •  ...customer experiences with AI. We are primarily an in-person...  ...you'll do As a Software Engineer on our Site Reliability team at Sierra, you will...  ...across Sierra’s AI-driven infrastructure. You’ll partner closely with...  ...Define the foundation of SRE practices at Sierra,... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    3 days ago
  •  ...customer experiences with AI. We are primarily an in-person...  ...'ll do As a Software Engineer on our Site Reliability team at Sierra, you will...  ...across Sierra's AI-driven infrastructure. You'll partner closely...  ...Define the foundation of SRE practices at Sierra, influencing... 
    Full time
    Flexible hours

    Sierra

    San Francisco, CA
    2 days ago
  • $200k - $260k

     ...robust and highly scalable Infrastructure as Code model. This...  ...leader driving reliability, automation, and scalability...  ...of classic SRE discipline which include...  ...demands of running agentic AI systems in production:...  ...teams, mentor senior engineers, and be a primary escalation... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    4 days ago
  • $194k - $267k

     ...Secure Every Identity, from AI to Human Identity is the key to unlocking...  ...AI by building the trusted, neutral infrastructure that enables organizations to safely...  ...tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    2 days ago
  • $181k - $263k

     ...requirements. The Global SRE team is responsible for...  ...looking for a Senior Staff Site Reliability Engineer who will set the technical...  ...engineering across LiveRamp's global infrastructure. This is a senior...  ...Familiarity with LLMs and AI-assisted development workflows... 
    Work from home
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    2 days ago
  • $220k - $235k

     ...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements...  ...high-output Staff/Senior Staff SRE to define the future of our...  ...~ Ability to build resilient infrastructure ~ Modern GitOps - Experience with... 
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    2 days ago
  • $230k - $390k

     ...customer experiences with AI. We are primarily an in-person...  ...you'll doAs a Software Engineer on our Site Reliability team at Sierra, you will...  ...across Sierra’s AI-driven infrastructure. You’ll partner closely with...  ....Define the foundation of SRE practices at Sierra,... 
    Full time
    Part time
    Flexible hours

    Sierra

    San Francisco, CA
    3 days ago
  •  ...civilization runs on the same infrastructure: agreements between people...  .... We're building the AI that finally changes that....  ...over the last 12 months. Engineering at Ivo Engineers at Ivo are...  ...We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure... 
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Ivo Inc.

    San Francisco, CA
    4 days ago
  • $180k - $250k

     ...Join to apply for the Cloud Infrastructure Engineer role at Braintrust ....  ...Company Braintrust is the AI observability platform. By...  ...Infrastructure Engineer to help us build reliable, scalable infrastructure and...  ...of experience in DevOps, SRE, or Infrastructure... 
    Flexible hours

    Brain Trust Inc

    San Francisco, CA
    4 days ago
  • $180k - $240k

     ...Senior Cloud Infrastructure Engineer Loft Orbital is revolutionizing...  ...to space by building reliable, shareable satellites...  ...is not your typical SRE role, we apply DevOps...  ...financial support ~ Off-sites and many social events...  ..., on-orbit AI, national security missions... 
    Temporary work
    Work at office
    Relocation package
    Flexible hours

    Loft Orbital

    San Francisco, CA
    2 days ago
  • $117.2k - $176.7k

     ...Category Software Engineering Job Details About...  ...Salesforce is the #1 AI CRM, where humans with...  ...team within the Cloud Infrastructure organization. Platform...  ...team a fast, secure, and reliable path to production....  ...engineers. Partner with SRE, security, and product... 

    Salesforce.Com Inc

    San Francisco, CA
    18 hours ago
  • $180k - $250k

     ...company building the next hyperscaler for AI agents, delivering a cloud platform...  ...opportunity to help define the foundational infrastructure shaping the future of AI development....  ...Collaborate closely with Distributed Systems Engineers and work directly with users. Skills/... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  • $160k - $225k

     ...About Fable Security AI-driven threats and human error are...  ...scale the foundational data infrastructure powering a category-...  ...product Work closely with engineering, data science, and product teams...  ...data pipeline. You’ll ensure reliability, scalability, and... 
    Full time
    Work experience placement
    Relocation package
    Flexible hours

    Fable Security

    San Francisco, CA
    1 day ago
  •  ...builds, and operates critical infrastructure that enables research at...  ...our workloads, while remaining reliable and easy to use. About the...  ...looking for a staff-level software engineer to own production-critical...  ...About OpenAI OpenAI is an AI research and deployment... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...stacks to accelerate the progress of AI applications out into the real world...  ...Anyscale is looking for a Software Engineer to join the Infrastructure team. Anyscale aims to provide the next...  ...Develop features to enhance the reliability, performance, scalability, and observability... 
    Full time

    Anyscale

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Infrastructure. Be the first to apply!