Site Reliability Engineer (SRE) - AI Infrastructure
$300kAre you looking for an exciting new opportunity?
Join a stealth-mode hyperscale data center startup building a next-generation AI and cloud platform designed for startups and advanced research, powered by thousands of H100, H200, and B200 GPUs available on demand. Their platform supports everything from rapid experimentation to full-scale model training and inference, with flexible orchestration via Slurm, Kubernetes, or direct SSH access.
This is a rare opportunity to work at the intersection of hyperscale infrastructure and AI, shaping the operational backbone of one of the largest GPU clusters in private deployment. If you want to build and operate infrastructure for frontier AI workloads, automate systems at petascale, and be part of a founding engineering team, this is the place to do it.
Responsibilities:
- Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
- Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
- Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
- Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilization, and data flow.
- Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
- Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.
Skills / Must Have:
- 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
- Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
- Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
- Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
- Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
- Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
- Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.
Benefits:
- Equity
Salary:
- $300,000 gross per year
- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...SuggestedWork at officeWeekend work
$232k - $319k
...Every Identity, from AI to HumanIdentity is the... ...the trusted, neutral infrastructure that enables organizations... ...with great people and reliable, cost-effective, and... ...initiatives across SRE & Infrastructure organization... ...of SRE and product engineering by developing robust...SuggestedPermanent employmentLocal areaWorldwideFlexible hours$180.5k - $236.91k
...Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team.... ...technical domains such as DevOps, site reliability, and cloud best practices Lead the... ...reimbursements. Artificial Intelligence (AI): Our AI Guidelines outline the...SuggestedRemote jobFull timeWork at office$350k
...Join a rapidly growing AI infrastructure provider delivering large-... ...infrastructure providers to build reliable, high-performance... ...opportunity is for a Staff Site Reliability Engineer to lead the reliability of... ...environments Proven Staff-level SRE or infrastructure...SuggestedFull time$200k - $260k
...Nimble Nimble is an AI robotics company building... ...team of the world's best engineers and operators. If you... ...Software Engineer to join our Infrastructure Team , focused on building secure, reliable, and scalable... ...security engineering, or SRE experience. Ability to...SuggestedLocal areaFlexible hours$207k - $300k
...changes that improve reliability and velocity.... ...strategy for Home SRE.Minimum qualifications... ...Science or Engineering, or a related field... ...ML model serving infrastructure, real-time low-latency... ...engineering organizations.Site Reliability... ...generation real-time Voice AI. We focus on high-...Worldwide$80 per hour
...Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy... ...Linux systems administration, infrastructure monitoring, incident response, programming... ...developing or deploying Agentic AI or autonomous automation tools to streamline...Hourly payFull timeWork at officeLocal areaShift workNight shift- ...Description Developer & Infrastructure Expert Role Type:... ...to evaluate AI-powered workflows across... ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-... ...for accuracy and reliability. Work with AWS,... ...Cloud Infrastructure Site Reliability...Remote jobFor contractors
- ...and Amsterdam.Plaid's Infrastructure team builds the... ...and tooling that help engineering teams develop, deploy... ...product team.As a Staff Site Reliability Engineer on Release Engineering... ...and safe even as AI-assisted development... ...in backend systems, SRE, or platform...Permanent employmentWork experience placementWork at officeLocal area
$190.8k - $267.1k
...its business. The reliability of our Ads systems... ...closely with Ads Engineering to improve reliability... ...for a Senior Site Reliability Engineer... ...build, and maintain infrastructure, tooling, and... ...Drive adoption of SRE best practices including... ...intelligence (AI). You will have the...For contractorsWork experience placement$148.5k - $223.9k
...SalesforceSalesforce is the #1 AI CRM, where humans with... ...is seeking a senior engineering candidate to join the Site Reliability organization in San... ...with counterparts in the Infrastructure and R&D organizations, this... ...protected. The ExperienceAs an SRE, you will be a technical...Full timeWorldwideWeekend work$113.4k - $162k
...everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,... ...and operates its systems in an AI-first environment where intelligent... ...the design and implementation of new SRE best practices.You'll be a great fit...Temporary work$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support... ...that ensure cluster reliability and security (e.g., CoreDNS... ...data platform for the AI era, enabling builders to...Work at officeLocal areaRemote workWorldwideFlexible hours- ...combination of proprietary infrastructure and software, we... ...-to-end. You use AI to work smarter... ...About the teamThe Engineering team at Airwallex... ...build scalable, reliable, and secure products... ...borders.Our SRE team is breaking new... ...’ll doAs a Senior Site Reliability Engineer...Temporary workLocal areaWorldwide
$165k - $227k
...Every Identity, from AI to HumanIdentity is the... ...the trusted, neutral infrastructure that enables... ...too, let's talk.The Engineering OpportunityWe are looking... ...an experienced Senior Site Reliability Engineer to join Okta... ...contributor within the EPG SRE organization,...Local areaWorldwideFlexible hours- ...mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase... ...the Enterprise Technology, Infrastructure Platforms team, you will solve... ...tuningUses enterprise-authorized AI capabilities within the... ...experience and practical SRE/observability fundamentals...
$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB’s broader engineering... ...in engineering the reliable, globally connected,... ...a talented Senior Site Reliability Engineer (SRE... ...data platform for the AI era, enabling builders...Local areaRemote workWorldwideFlexible hours$230k - $390k
...customer experiences with AI. We are primarily an in-person... ...you'll doAs a Software Engineer on our Site Reliability team at Sierra, you will... ...across Sierra’s AI-driven infrastructure. You’ll partner closely with... ....Define the foundation of SRE practices at Sierra,...Full timeFlexible hours- ...world's most dynamic AI companies, like Cursor... ...AI research, flexible infrastructure, and seamless developer... ...help build the platform engineers turn to to ship AI... ...high performance and reliability. You’ll own scheduling... ...Partner closely with SRE and Capacity teams to...Full timeFlexible hours
- ...About Us At Hayden AI, we are on a mission... ...world challenges. The Infrastructure Engineering team is crucial to... ...performance, maximum reliability, and cost-efficiency... ...modeling best practices in site reliability,... ...in modern DevOps and SRE practices, including...Full timeShift work
$194k - $267k
Secure Every Identity, from AI to HumanIdentity is the key... ...building the trusted, neutral infrastructure that enables organizations... ...technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk... ...Platform that enables our SRE teams and business partners....Permanent employmentWork at officeLocal areaWorldwideFlexible hours$217k - $303.9k
...continues to scale globally, reliability and performance are more critical than ever. The Site Experience SRE team sits at the intersection of infrastructure, product engineering, and user experience - ensuring... ...by artificial intelligence (AI). You will have the...For contractorsWork experience placement$194k - $267k
Secure Every Identity, from AI to HumanIdentity is the key to unlocking the... ...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$204k - $306k
...Every Identity, from AI to HumanIdentity is the... ...the trusted, neutral infrastructure that enables... ...let's talk.Manager, Site Reliability EngineeringSan Francisco... ...IDaaS Site Reliability Engineering GroupOkta authenticates... ...doing Managing a team of SRE’s supporting various...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$195k - $257.5k
...and programmable blockchain infrastructure. Circle’s platform includes... ...responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team,... ...Infrastructure as Code, and modern SRE practices to deliver... ...teams, you’ll build AI-powered tooling and automation...Flexible hours$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA... ...Formation Bio is a tech and AI driven pharma company differentiated... ...will build and operate the infrastructure, delivery systems, and... ...on infrastructure and SRE fundamentals. About You...Work experience placementWork at officeLocal areaRelocation3 days per week- ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated access to business data across... ...in infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes...Work at officeLocal areaFlexible hours
- ...now part of Superhuman, the AI productivity platform on a mission... ...goals, we’re looking for an SRE to join our infrastructure team. This role will be... ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning...WorldwideHome officeFlexible hours
$181k - $225k
...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming... ...management industry by building an AI platform for wealth... ...to hire a high performing SRE to join our growing Platform... ...contribute to solutions with shared infrastructure that benefit all of...Work at officeImmediate start3 days per week$130k - $200k
...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through... ...The Role This is a career-level SRE role for someone who wants to own systems...Shift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Infrastructure. Be the first to apply!
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- infrastructure engineer San Francisco, CA
- infrastructure engineering manager San Francisco, CA
- infrastructure developer San Francisco, CA
- remote infrastructure engineer San Francisco, CA
- senior infrastructure engineer San Francisco, CA




