Director, Site Reliability Engineering -- AI Accelerator Infrastructure
$195k - $285kPhizenix
Job Description
Job Description
The Role
You will build and lead Our Client's Site Reliability Engineering function from the ground up — owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate with our Client on hardware and software deployments.
You are a hands-on engineering leader. You will establish SRE as a discipline at our client, hire and grow the team, set the technical direction, own SLOs for critical systems, and be the senior escalation point when things go wrong — all in parallel. You will partner closely with the Director of DevOps Engineering, whose pipelines and automation run on the infrastructure you own, and work directly with hardware and software development teams to ensure HPC infrastructure meets their workload requirements.
What You Will DoLeadership & Organizational Build-Out
- Own the SRE function end-to-end: define the team's charter, establish SRE as a discipline within our client's engineering culture, and drive buy-in across hardware, software, and executive stakeholders who have operated without a dedicated SRE team.
- Hire, develop, and retain a team of 3–5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement from day one.
- Define the SRE technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model — and execute against it with your hands on the keyboard where needed.
- Serve as the senior technical escalation for critical incidents — guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches.
- Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.
- Partner with the Director of DevOps Engineering to align infrastructure reliability with pipeline and automation delivery; the two functions must operate as a unified platform.
- Direct a dedicated Data Center & Lab Technician team — your hands and feet across on-premises and colocation facilities; set their work priorities, establish operational standards, and ensure physical infrastructure execution aligns with the SRE technical roadmap.
- Establish SRE process from a zero baseline: define SLIs and SLOs, build error budgets, design on-call rotations, and create the incident management framework our client currently lacks.
- Own 24×7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services — designing for failure domains, progressive delivery, and strict change control at every tier.
- Own the full observability stack (metrics, traces, logs) and instrument it from the ground up — Prometheus, Grafana, and/or Datadog — with SLO visibility, alert design, and E2E signal quality.
- Evolve incident and problem management into a data-driven discipline: automated triage workflows, pattern detection across recurring failures, and every P0/P1 producing a written RCA with tracked systemic fixes.
- Own FinOps and capacity planning as a unified discipline across all three infrastructure tiers — cloud (AWS, Azure, GCP), colocation, and on-premises: establish spend visibility and attribution across every tier, model TCO comparatively, drive workload placement decisions based on cost and performance, and anticipate infrastructure needs for new silicon programs and customer deployments.
- Own the migration from ad-hoc JBOD-based storage and point-in-time snapshots to an enterprise-grade shared storage platform spanning on-premises, colocation, and cloud tiers — covering architecture, vendor selection, data protection design (snapshots, replication, DR), and integration with HPC workloads and development environments
There is no automation baseline today. You will build it.
- Drive IaC-first discipline across the team — Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management; this capability is currently absent and you will establish it.
- Build self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle.
- Instrument the team's own development practices — runbooks, change governance, deployment pipelines for infrastructure code — establishing standards that scale as the team grows.
- Build a documentation culture from scratch: runbooks, architecture diagrams, and operational playbooks maintained as living artifacts — not a one-time project.
- Design and scale a follow-the-sun on-call model as our client expands globally; the framework you build now will be the foundation the team inherits.
- Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our client's accelerator roadmap.
- Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 15+ years in SRE, infrastructure engineering, or production engineering.
- 5+ years leading SRE or infrastructure engineering teams — including experience building or significantly rebuilding a function, not just managing a steady-state team.
- Demonstrated track record of establishing SRE as a discipline in an organization that lacked it: defining SLOs, creating on-call frameworks, standing up observability, and driving cultural change with engineering teams that came from a reactive ops background.
- Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), kernel tuning, and bare-metal operations; hands-on experience with enterprise shared storage platforms (NAS/SAN, NFS/SMB at scale, snapshot and replication architectures) and hybrid-cloud storage integration across on-prem and cloud tiers.
- Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.
- IaC fluency: Terraform and Ansible at production scale — module design, remote state, environment isolation, and change governance.
- Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.
- Full observability stack ownership: Prometheus, Grafana, and/or Datadog — SLO definition, alert design, and E2E signal quality.
- Strong Python and/or Go — production services, not just scripts; automation that touches real infrastructure safely.
- Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership, including stakeholders with no infrastructure background.
- Ability to operate in a high-ambiguity, low-process environment — you build the structure, you don't inherit it.
Strongly Preferred
- Experience operating customer-facing infrastructure or platform services — reliability expectations beyond internal tooling.
- Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink — setup, troubleshooting, and performance tuning.
- HPC job scheduler experience: Slurm, LSF, or equivalent — setup, tuning, and integration with infrastructure automation.
- Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo — unified observability and IaC across all tiers.
- FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations.
- ITIL knowledge or equivalent structured incident/problem/change management framework.
- Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.
California Pay Range
$195,000—$285,000 USD
$155k - $235k
Site Reliability Engineering — AI Accelerator Infrastructure we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible....Suggested- ...SRE team, responsible for the reliability, automation, and observability of the infrastructure that the company runs on. You will... ..., and self-service tooling for engineering teams. Develop networking... ...application layer using structured, AI-assisted workflows —...Suggested
$184k - $287.5k
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale... .... Apply automation and Generative AI/Agentic solutions to minimize manual... ...performant, and supportable. Background with infrastructure automation. Experience running...Suggested$236k - $330k
...and processing of engineering hardware must be performed on site. Minimum qualifications... .... Experience accelerating robot learning... ...latency tele-operation infrastructure, haptic feedback... .... The AI and Infrastructure... ...scale, efficiency, reliability and velocity. Our...SuggestedContract workRemote workWorldwideFlexible hours$168k - $258.75k
...the next generation of AI-powered simulation tools to accelerate hardware and silicon development... .... AI is reshaping how engineering teams operate, and we... ...in AI, silicon, or infrastructure 10+ years of technical... ...projects spanning multiple sites Your base salary will be...Suggested$184k - $287.5k
...graphics, PC gaming, and accelerated computing for more... ...unlimited potential of AI to define the next era... ...group of forward‑thinking engineers tackling some of the... ...and help shape how AI infrastructure runs in production. In... ...projects to enable reliable operation at hyperscale...Remote work$300 per month
Crusoe's mission is to accelerate the abundance of... .... We’re crafting the engine that powers a world where... ...ambitiously with AI — without sacrificing... ...transformative cloud infrastructure. About This Role At... ...our Compute‑focused Site Reliability Engineers are the backbone...Temporary work$152k - $241.5k
...graphics, PC gaming, and accelerated computing for more... ...harness the power of AI to deliver... ...network fabrics. Use IaC(Infrastructure‑as‑Code) and configuration... ...lifecycle management, fleet reliability/auto‑healing, E2E... .... Mentored other engineers and influenced technical...$176k - $333.5k
We are seeking a Senior Infrastructure System Software Engineer with profound expertise in High-Performance Computing (HPC) and AI workload management, as well as Kubernetes-based infrastructure... ...-running system service solutions to accelerate the training of extensive AI models....$124k - $195.5k
...computer graphics, PC gaming, and accelerated computing for more than 25 years.... ...tapping into the unlimited potential of AI to define the next era of... ...dedicated and motivated System Software Engineer who is passionate about AI Infrastructure. You will collaborate with...$207k - $301k
Google is hiring a Staff Software Engineer in Sunnyvale, California, to enhance its AI and Infrastructure capabilities. This role involves collaborating with teams to integrate co-accelerators, developing system software, and addressing complex challenges in kernel and...$262k - $365k
...Master’s degree or PhD in Engineering, Computer Science, or... ...that power Google's AI and High‑Performance Computing (HPC) infrastructure. Your focus will be on... ...‑scale deployment of Accelerators (e.g., GPUs, TPUs, etc... ...unparalleled scale, efficiency, reliability and velocity. Our...Worldwide$220k - $320k
...Description Job Description About the role Own the infrastructure that engineering depends on — Kubernetes clusters, CI/CD pipelines, on-... ...with chip-design and software teams driving DensityAI's AI accelerator program from first silicon through scale-out. What...H1bVisa sponsorshipWork visa$188k - $275k
...frontier model is as much an infrastructure problem as it is a... .... Reporting to the Director of Product Management,... ...challenges that directly accelerate deep learning... ...with the world’s best AI teams, and delivering... ...and CoreWeave platform engineering. This role sits at the...Temporary workCasual workWork at officeRemote workFlexible hours$235.03k - $352.29k
...driver, combining cutting‑edge AI with automotive‑grade... ...deep expertise in large‑scale infrastructure, workload orchestration, as... ...ensuring our researchers and engineers have seamless access to the... ...infrastructure bottlenecks and accelerate the Nuro Driver™ development...$165k - $242k
...Essential Cloud for AI™. Built for pioneers... ...CoreWeave combines superior infrastructure performance with deep... ...expertise to accelerate breakthroughs and turn... ...About The Role Senior engineers are area owners who lead... ..., throughput, and reliability across multiple services...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$216k - $333.5k
Are you an experienced Software Infrastructure Engineer looking for an opportunity in Autonomous Vehicles Software? Come and join our international... ...you. We are looking for great people like you to help us accelerate the next wave of artificial intelligence. If you’re...$207k - $275k
...The Essential Cloud for AI™. Built for pioneers by pioneers... ...combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute... ...platform that lets our engineers ship software quickly, reliably, and safely. We own the...Permanent employmentTemporary workCasual workWork at officeFlexible hours- Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distributed AI-RAN Environments) We challenge conventional limits... ...orchestration at scale. Hands‑on work with GPU‑accelerated systems and high‑performance computing (HPC) environments...
$204k - $343k
...the future of physical AI. Founded in 2017 and... ...company creates the digital infrastructure to bring intelligence... ...About The Role As an Engineering Manager on the ML... ...remove bottlenecks and accelerate the path from experimentation... ...systems that run reliably at massive scale Nice...Full timeFor contractorsFor subcontractor- ...world's largest AI chip, 56 times larger... .... TPM role owns site and data center... ...closely with Hardware Engineering, Inference... ...Cerebras systems are reliably deployed,... ...Engineering AI Cloud Infrastructure & Operations Network... ...AI/ML, HPC, or accelerator-based...
$116k - $189.75k
...board new applications, AI/ML services, and model endpoints on AWS Infrastructure. Make meaningful... ...marketing campaigns and site migrations. Set up Akamai... ...in Computer Science/Engineering or a related field, or... ..., or managing GPU-accelerated infrastructure. Strong...- ...performing team that delivers infrastructure and performance... ...a Lead Infrastructure Engineer at JPMorganChase... ...enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis... ...health care coverage, on-site health and wellness centers...Permanent employment
$248k - $391k
...graphics, PC gaming, and accelerated computing for 30... ...unlimited potential of AI to define the next... ...skilled Principal Software Engineer to join our dynamic... ...performance of our infrastructure both on‑prem and in the... ...engineering, site reliability, or systems architecture...$188k - $275k
CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ..., CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute... ...environments, enjoy tackling challenging engineering problems, and are excited by...Permanent employmentTemporary workCasual workWork at officeWorldwideFlexible hours$200k - $322k
...systems telemetry, and cloud infrastructure operations. You will play a... ...will work closely with our Engineering, Infrastructure, and Software... ...troubleshoot, debug, and manage AI infrastructure effectively.... ..., and trust across the accelerated computing ecosystem. Providing...Worldwide$124k - $195.5k
...computer graphics, PC gaming, and accelerated computing for more than 25... ...the unlimited potential of AI to define the next era of... ...Marketing Manager - Data Center Infrastructure Specialist to join our... ...Marketing, Computer Science, Engineering, or a related field (or equivalent...- ...DESCRIPTION Elevate your engineering prowess to... ...among the top echelon in site reliability. As a Senior Lead... ...Chase within the Infrastructure Platforms and Foundational... ...-authorized AI capabilities within the... ...work environment to accelerate reliability design and...
$156k - $229k
Senior DFT Engineer, Test Infrastructure, Google Cloud Sunnyvale, CA, USA Job Level: Mid Minimum... ...’ll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to... ...unparalleled scale, efficiency, reliability and velocity. Our customers include...Full timeWorldwide- ...forefront of revolutionizing AI computing by reengineering infrastructure at the system level. Our... ...in efficient, more reliable computing at a fraction... ...seeking a skilled Presales Engineer to be the technical... ...hands-on experience in GPU acceleration, Kubernetes, Terraform,...Work at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director, Site Reliability Engineering -- AI Accelerator Infrastructure. Be the first to apply!
- principal cloud engineer Santa Clara, CA
- senior principal engineer Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- general engineer Santa Clara, CA
- principal application developer Santa Clara, CA
- principal engineer Santa Clara, CA
- director of product engineering Santa Clara, CA
- director data engineering Santa Clara, CA
- data center chief engineer Santa Clara, CA
- chief engineer Santa Clara, CA

