Site Reliability Engineer (SRE) - AI Infrastructure
$300kAre you looking for an exciting new opportunity?
Join a stealth-mode hyperscale data center startup building a next-generation AI and cloud platform designed for startups and advanced research, powered by thousands of H100, H200, and B200 GPUs available on demand. Their platform supports everything from rapid experimentation to full-scale model training and inference, with flexible orchestration via Slurm, Kubernetes, or direct SSH access.
This is a rare opportunity to work at the intersection of hyperscale infrastructure and AI, shaping the operational backbone of one of the largest GPU clusters in private deployment. If you want to build and operate infrastructure for frontier AI workloads, automate systems at petascale, and be part of a founding engineering team, this is the place to do it.
Responsibilities:
- Design, deploy, and maintain large-scale GPU clusters (H100/H200/B200) for training and inference workloads.
- Build automation pipelines for provisioning, scaling, and monitoring compute resources across Slurm and Kubernetes environments.
- Develop observability, alerting, and auto-healing systems for high-availability GPU workloads.
- Collaborate with ML, networking, and platform teams to optimise resource scheduling, GPU utilization, and data flow.
- Implement infrastructure-as-code, CI/CD pipelines, and reliability standards across thousands of nodes.
- Diagnose performance bottlenecks and drive continuous improvements in reliability, latency, and throughput.
Skills / Must Have:
- 7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting large-scale compute environments.
- Strong hands-on experience with Kubernetes and Slurm for cluster orchestration and workload management.
- Deep knowledge of Linux systems, networking, and GPU infrastructure (NVIDIA H100/H200/B200 preferred).
- Proficiency in Python, Go, or Bash for automation, tooling, and performance tuning.
- Experience with observability stacks (Prometheus, Grafana, Loki) and incident response frameworks.
- Familiarity with high-performance computing (HPC) or AI/ML training infrastructure at scale.
- Background in reliability engineering, distributed systems, or hardware acceleration environments is a strong plus.
Benefits:
- Equity
Salary:
- $300,000 gross per year
- ...mission to create the world's first AI-powered Personal & Entrepreneurial Resource... ...the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be... ...Grafana, Datadog, ELK). Automate infrastructure provisioning, deployment, and incident...SuggestedTemporary workWorldwide
$350k
...Site Reliability Engineer (SRE) San Francisco Thinking Machines Lab's mission is to empower humanity... ...to the knowledge and tools to make AI work for their unique needs and... ...in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility...SuggestedLocal areaVisa sponsorshipWork visaRelocation package$232k - $319k
...Every Identity, from AI to Human Identity is... ...the trusted, neutral infrastructure that enables organizations... ...with great people and reliable, cost-effective, and... ...various initiatives across SRE & Infrastructure... ...velocity of SRE and product engineering by developing robust...SuggestedPermanent employmentLocal areaWorldwideFlexible hours$350k
...Join a rapidly growing AI infrastructure provider delivering large-... ...infrastructure providers to build reliable, high-performance... ...opportunity is for a Staff Site Reliability Engineer to lead the reliability of... ...environments Proven Staff-level SRE or infrastructure...SuggestedPermanent employment- ...the intersection of labor markets and AI research. We partner with leading AI... ...headquarters. About the Role As a Site Reliability Engineer (SRE) at Mercor, you’ll own production reliability... ...systems, partnering directly with infrastructure leadership. You’ll play a...Suggested
$148.5k - $223.9k
...Category Software Engineering Job Details About... ...Salesforce is the #1 AI CRM, where humans... ...candidate to join the Site Reliability organization in San Francisco... ...counterparts in the Infrastructure and R&D organizations... ...As an SRE, you will be a technical...WorldwideWeekend work$260k - $300k
...We are an applied AI lab building end-to-end software... ...the first AI software engineer. Our team is... ...both the production reliability of our user-facing... ...pipelines, deployment infrastructure, and developer tooling... ...engineering fundamentals; SRE at Cognition means...- ...Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates... ...software engineering and applies them to infrastructure and operations problems. The main... ...in Python supported by Gen AI tooling to accelerate development of...Immediate startRemote workWorldwide
- ...now part of Superhuman, the AI productivity platform on a mission... ...goals, we're looking for an SRE to join our infrastructure team. This role will be... ...building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning...WorldwideHome officeFlexible hours
$155k - $222.6k
...Meet the Team The SRE Fleet team is... ...efficiency of the infrastructure that powers our global... ...As a team of six engineers distributed across... ...on automation, reliability, and operational excellence... ...of experience in Site Reliability... ...Experience leveraging AI-assisted...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- ...world's most dynamic AI companies, like Cursor... ...AI research, flexible infrastructure, and seamless developer... ...help build the platform engineers turn to to ship AI... ...high performance and reliability. You’ll own scheduling... ...Partner closely with SRE and Capacity teams to...Full timeFlexible hours
- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge... ...AI research, flexible infrastructure, and seamless developer... ...these as part of the SRE team: Improve Baseten...Flexible hours
- ...We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating... ...Required skills ~3+ years in SRE, DevOps, or infrastructure engineering...
$120k - $168.49k
...Site Reliability Engineer, Cloud Infrastructure About Quizlet At Quizlet, our mission is to help every learner achieve... ...looking for a Site Reliability Engineer (SRE) to join our infrastructure team... .... Cutting‑edge tech: Generative AI, adaptive learning, cognitive science...InternshipWork at office3 days per week$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be... ...reliability. Our infrastructure runs on AWS across... ...years of experience in Site Reliability... ...with AIOps - using AI/ML-based tooling for...Work at officeImmediate startWorldwideMonday to FridayFlexible hours- ...customer experiences with AI. We are primarily an in-person... ...you'll do As a Software Engineer on our Site Reliability team at Sierra, you will... ...across Sierra’s AI-driven infrastructure. You’ll partner closely with... ...Define the foundation of SRE practices at Sierra,...Full timeFlexible hours
- ...customer experiences with AI. We are primarily an in-person... ...'ll do As a Software Engineer on our Site Reliability team at Sierra, you will... ...across Sierra's AI-driven infrastructure. You'll partner closely... ...Define the foundation of SRE practices at Sierra, influencing...Full timeFlexible hours
$200k - $260k
...robust and highly scalable Infrastructure as Code model. This... ...leader driving reliability, automation, and scalability... ...of classic SRE discipline which include... ...demands of running agentic AI systems in production:... ...teams, mentor senior engineers, and be a primary escalation...Casual workWork at officeRemote workFlexible hours$194k - $267k
...Secure Every Identity, from AI to Human Identity is the key to unlocking... ...AI by building the trusted, neutral infrastructure that enables organizations to safely... ...tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$181k - $263k
...requirements. The Global SRE team is responsible for... ...looking for a Senior Staff Site Reliability Engineer who will set the technical... ...engineering across LiveRamp's global infrastructure. This is a senior... ...Familiarity with LLMs and AI-assisted development workflows...Work from homeFlexible hoursNight shift$220k - $235k
...Staff/Senior Staff Site Reliability Engineer Ironclad is the leading AI contracting platform that transforms agreements... ...high-output Staff/Senior Staff SRE to define the future of our... ...~ Ability to build resilient infrastructure ~ Modern GitOps - Experience with...Full timeContract workWork at office$230k - $390k
...customer experiences with AI. We are primarily an in-person... ...you'll doAs a Software Engineer on our Site Reliability team at Sierra, you will... ...across Sierra’s AI-driven infrastructure. You’ll partner closely with... ....Define the foundation of SRE practices at Sierra,...Full timePart timeFlexible hours- ...civilization runs on the same infrastructure: agreements between people... .... We're building the AI that finally changes that.... ...over the last 12 months. Engineering at Ivo Engineers at Ivo are... ...We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure...Contract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$180k - $250k
...Join to apply for the Cloud Infrastructure Engineer role at Braintrust .... ...Company Braintrust is the AI observability platform. By... ...Infrastructure Engineer to help us build reliable, scalable infrastructure and... ...of experience in DevOps, SRE, or Infrastructure...Flexible hours$180k - $240k
...Senior Cloud Infrastructure Engineer Loft Orbital is revolutionizing... ...to space by building reliable, shareable satellites... ...is not your typical SRE role, we apply DevOps... ...financial support ~ Off-sites and many social events... ..., on-orbit AI, national security missions...Temporary workWork at officeRelocation packageFlexible hours$117.2k - $176.7k
...Category Software Engineering Job Details About... ...Salesforce is the #1 AI CRM, where humans with... ...team within the Cloud Infrastructure organization. Platform... ...team a fast, secure, and reliable path to production.... ...engineers. Partner with SRE, security, and product...$180k - $250k
...company building the next hyperscaler for AI agents, delivering a cloud platform... ...opportunity to help define the foundational infrastructure shaping the future of AI development.... ...Collaborate closely with Distributed Systems Engineers and work directly with users. Skills/...$160k - $225k
...About Fable Security AI-driven threats and human error are... ...scale the foundational data infrastructure powering a category-... ...product Work closely with engineering, data science, and product teams... ...data pipeline. You’ll ensure reliability, scalability, and...Full timeWork experience placementRelocation packageFlexible hours- ...builds, and operates critical infrastructure that enables research at... ...our workloads, while remaining reliable and easy to use. About the... ...looking for a staff-level software engineer to own production-critical... ...About OpenAI OpenAI is an AI research and deployment...Full timeWork at officeRelocation package
- ...stacks to accelerate the progress of AI applications out into the real world... ...Anyscale is looking for a Software Engineer to join the Infrastructure team. Anyscale aims to provide the next... ...Develop features to enhance the reliability, performance, scalability, and observability...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - AI Infrastructure. Be the first to apply!
- site reliability engineer San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site reliability engineer remote San Francisco, CA
- senior infrastructure engineer San Francisco, CA
- infrastructure engineering manager San Francisco, CA
- infrastructure engineer San Francisco, CA
- security infrastructure engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- infrastructure developer San Francisco, CA
- principal infrastructure engineer San Francisco, CA




