Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Site Reliability Engineering -- AI Accelerator Infrastructure

$195k - $285k

Phizenix

Job Description

Job Description

The Role

You will build and lead Our Client's Site Reliability Engineering function from the ground up — owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate with our Client on hardware and software deployments.

You are a hands-on engineering leader. You will establish SRE as a discipline at our client, hire and grow the team, set the technical direction, own SLOs for critical systems, and be the senior escalation point when things go wrong — all in parallel. You will partner closely with the Director of DevOps Engineering, whose pipelines and automation run on the infrastructure you own, and work directly with hardware and software development teams to ensure HPC infrastructure meets their workload requirements.

What You Will Do

Leadership & Organizational Build-Out

  • Own the SRE function end-to-end: define the team's charter, establish SRE as a discipline within our client's engineering culture, and drive buy-in across hardware, software, and executive stakeholders who have operated without a dedicated SRE team.
  • Hire, develop, and retain a team of 3–5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement from day one.
  • Define the SRE technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model — and execute against it with your hands on the keyboard where needed.
  • Serve as the senior technical escalation for critical incidents — guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches.
  • Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.
  • Partner with the Director of DevOps Engineering to align infrastructure reliability with pipeline and automation delivery; the two functions must operate as a unified platform.
  • Direct a dedicated Data Center & Lab Technician team — your hands and feet across on-premises and colocation facilities; set their work priorities, establish operational standards, and ensure physical infrastructure execution aligns with the SRE technical roadmap.
Reliability & Observability — Building From Scratch
  • Establish SRE process from a zero baseline: define SLIs and SLOs, build error budgets, design on-call rotations, and create the incident management framework our client currently lacks.
  • Own 24×7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services — designing for failure domains, progressive delivery, and strict change control at every tier.
  • Own the full observability stack (metrics, traces, logs) and instrument it from the ground up — Prometheus, Grafana, and/or Datadog — with SLO visibility, alert design, and E2E signal quality.
  • Evolve incident and problem management into a data-driven discipline: automated triage workflows, pattern detection across recurring failures, and every P0/P1 producing a written RCA with tracked systemic fixes.
  • Own FinOps and capacity planning as a unified discipline across all three infrastructure tiers — cloud (AWS, Azure, GCP), colocation, and on-premises: establish spend visibility and attribution across every tier, model TCO comparatively, drive workload placement decisions based on cost and performance, and anticipate infrastructure needs for new silicon programs and customer deployments.
  • Own the migration from ad-hoc JBOD-based storage and point-in-time snapshots to an enterprise-grade shared storage platform spanning on-premises, colocation, and cloud tiers — covering architecture, vendor selection, data protection design (snapshots, replication, DR), and integration with HPC workloads and development environments
Automation & Infrastructure as Code — Establishing the Baseline

There is no automation baseline today. You will build it.

  • Drive IaC-first discipline across the team — Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management; this capability is currently absent and you will establish it.
  • Build self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle.
  • Instrument the team's own development practices — runbooks, change governance, deployment pipelines for infrastructure code — establishing standards that scale as the team grows.
Documentation & Global Collaboration
  • Build a documentation culture from scratch: runbooks, architecture diagrams, and operational playbooks maintained as living artifacts — not a one-time project.
  • Design and scale a follow-the-sun on-call model as our client expands globally; the framework you build now will be the foundation the team inherits.
  • Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our client's accelerator roadmap.
What You Will Bring Required
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 15+ years in SRE, infrastructure engineering, or production engineering.
  • 5+ years leading SRE or infrastructure engineering teams — including experience building or significantly rebuilding a function, not just managing a steady-state team.
  • Demonstrated track record of establishing SRE as a discipline in an organization that lacked it: defining SLOs, creating on-call frameworks, standing up observability, and driving cultural change with engineering teams that came from a reactive ops background.
  • Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), kernel tuning, and bare-metal operations; hands-on experience with enterprise shared storage platforms (NAS/SAN, NFS/SMB at scale, snapshot and replication architectures) and hybrid-cloud storage integration across on-prem and cloud tiers.
  • Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.
  • IaC fluency: Terraform and Ansible at production scale — module design, remote state, environment isolation, and change governance.
  • Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.
  • Full observability stack ownership: Prometheus, Grafana, and/or Datadog — SLO definition, alert design, and E2E signal quality.
  • Strong Python and/or Go — production services, not just scripts; automation that touches real infrastructure safely.
  • Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership, including stakeholders with no infrastructure background.
  • Ability to operate in a high-ambiguity, low-process environment — you build the structure, you don't inherit it.

Strongly Preferred

  • Experience operating customer-facing infrastructure or platform services — reliability expectations beyond internal tooling.
  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink — setup, troubleshooting, and performance tuning.
  • HPC job scheduler experience: Slurm, LSF, or equivalent — setup, tuning, and integration with infrastructure automation.
  • Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo — unified observability and IaC across all tiers.
  • FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations.
  • ITIL knowledge or equivalent structured incident/problem/change management framework.
  • Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.

California Pay Range

$195,000—$285,000 USD

Vacancy posted 11 days ago
Similar jobs that could be interesting for youBased on the Director, Site Reliability Engineering -- AI Accelerator Infrastructure in Santa Clara, CA vacancy
  • $155k - $235k

    Site Reliability Engineering — AI Accelerator Infrastructure we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible.... 
    Suggested

    Phizenix

    Santa Clara, CA
    4 days ago
  •  ...SRE team, responsible for the reliability, automation, and observability of the infrastructure that the company runs on. You will...  ..., and self-service tooling for engineering teams. Develop networking...  ...application layer using structured, AI-assisted workflows —... 
    Suggested

    Calance

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale...  .... Apply automation and Generative AI/Agentic solutions to minimize manual...  ...performant, and supportable. Background with infrastructure automation. Experience running... 
    Suggested

    NVIDIA AI

    Santa Clara, CA
    4 days ago
  • $236k - $330k

     ...and processing of engineering hardware must be performed on site. Minimum qualifications...  .... Experience accelerating robot learning...  ...latency tele-operation infrastructure, haptic feedback...  .... The AI and Infrastructure...  ...scale, efficiency, reliability and velocity. Our... 
    Suggested
    Contract work
    Remote work
    Worldwide
    Flexible hours

    Google

    Sunnyvale, CA
    4 days ago
  • $168k - $258.75k

     ...the next generation of AI-powered simulation tools to accelerate hardware and silicon development...  .... AI is reshaping how engineering teams operate, and we...  ...in AI, silicon, or infrastructure 10+ years of technical...  ...projects spanning multiple sites Your base salary will be... 
    Suggested

    NVIDIA AI

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...graphics, PC gaming, and accelerated computing for more...  ...unlimited potential of AI to define the next era...  ...group of forward‑thinking engineers tackling some of the...  ...and help shape how AI infrastructure runs in production. In...  ...projects to enable reliable operation at hyperscale... 
    Remote work

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $300 per month

    Crusoe's mission is to accelerate the abundance of...  .... We’re crafting the engine that powers a world where...  ...ambitiously with AI — without sacrificing...  ...transformative cloud infrastructure. About This Role At...  ...our Compute‑focused Site Reliability Engineers are the backbone... 
    Temporary work

    Crusoe Energy Systems

    Sunnyvale, CA
    4 days ago
  • $152k - $241.5k

     ...graphics, PC gaming, and accelerated computing for more...  ...harness the power of AI to deliver...  ...network fabrics. Use IaC(Infrastructure‑as‑Code) and configuration...  ...lifecycle management, fleet reliability/auto‑healing, E2E...  .... Mentored other engineers and influenced technical... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $176k - $333.5k

    We are seeking a Senior Infrastructure System Software Engineer with profound expertise in High-Performance Computing (HPC) and AI workload management, as well as Kubernetes-based infrastructure...  ...-running system service solutions to accelerate the training of extensive AI models.... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $124k - $195.5k

     ...computer graphics, PC gaming, and accelerated computing for more than 25 years....  ...tapping into the unlimited potential of AI to define the next era of...  ...dedicated and motivated System Software Engineer who is passionate about AI Infrastructure. You will collaborate with... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $207k - $301k

    Google is hiring a Staff Software Engineer in Sunnyvale, California, to enhance its AI and Infrastructure capabilities. This role involves collaborating with teams to integrate co-accelerators, developing system software, and addressing complex challenges in kernel and... 

    Google

    Sunnyvale, CA
    4 days ago
  • $262k - $365k

     ...Master’s degree or PhD in Engineering, Computer Science, or...  ...that power Google's AI and High‑Performance Computing (HPC) infrastructure. Your focus will be on...  ...‑scale deployment of Accelerators (e.g., GPUs, TPUs, etc...  ...unparalleled scale, efficiency, reliability and velocity. Our... 
    Worldwide

    Google

    Sunnyvale, CA
    4 days ago
  • $220k - $320k

     ...Description Job Description About the role Own the infrastructure that engineering depends on — Kubernetes clusters, CI/CD pipelines, on-...  ...with chip-design and software teams driving DensityAI's AI accelerator program from first silicon through scale-out. What... 
    H1b
    Visa sponsorship
    Work visa

    DensityAI

    Mountain View, CA
    17 days ago
  • $188k - $275k

     ...frontier model is as much an infrastructure problem as it is a...  .... Reporting to the Director of Product Management,...  ...challenges that directly accelerate deep learning...  ...with the world’s best AI teams, and delivering...  ...and CoreWeave platform engineering. This role sits at the... 
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Weights & Biases

    Sunnyvale, CA
    1 day ago
  • $235.03k - $352.29k

     ...driver, combining cutting‑edge AI with automotive‑grade...  ...deep expertise in large‑scale infrastructure, workload orchestration, as...  ...ensuring our researchers and engineers have seamless access to the...  ...infrastructure bottlenecks and accelerate the Nuro Driver™ development... 

    Nuro

    Mountain View, CA
    4 days ago
  • $165k - $242k

     ...Essential Cloud for AI™. Built for pioneers...  ...CoreWeave combines superior infrastructure performance with deep...  ...expertise to accelerate breakthroughs and turn...  ...About The Role Senior engineers are area owners who lead...  ..., throughput, and reliability across multiple services... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • $216k - $333.5k

    Are you an experienced Software Infrastructure Engineer looking for an opportunity in Autonomous Vehicles Software? Come and join our international...  ...you. We are looking for great people like you to help us accelerate the next wave of artificial intelligence. If you’re... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $207k - $275k

     ...The Essential Cloud for AI™. Built for pioneers by pioneers...  ...combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute...  ...platform that lets our engineers ship software quickly, reliably, and safely. We own the... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distributed AI-RAN Environments) We challenge conventional limits...  ...orchestration at scale. Hands‑on work with GPU‑accelerated systems and high‑performance computing (HPC) environments... 

    Intelliswift - An LTTS Company

    Sunnyvale, CA
    4 days ago
  • $204k - $343k

     ...the future of physical AI. Founded in 2017 and...  ...company creates the digital infrastructure to bring intelligence...  ...About The Role As an Engineering Manager on the ML...  ...remove bottlenecks and accelerate the path from experimentation...  ...systems that run reliably at massive scale Nice... 
    Full time
    For contractors
    For subcontractor

    Applied Intuition

    Sunnyvale, CA
    4 days ago
  •  ...world's largest AI chip, 56 times larger...  .... TPM role owns site and data center...  ...closely with Hardware Engineering, Inference...  ...Cerebras systems are reliably deployed,...  ...Engineering AI Cloud Infrastructure & Operations Network...  ...AI/ML, HPC, or accelerator-based... 

    Cerebras

    Sunnyvale, CA
    4 days ago
  • $116k - $189.75k

     ...board new applications, AI/ML services, and model endpoints on AWS Infrastructure. Make meaningful...  ...marketing campaigns and site migrations. Set up Akamai...  ...in Computer Science/Engineering or a related field, or...  ..., or managing GPU-accelerated infrastructure. Strong... 

    NVIDIA

    Santa Clara, CA
    1 day ago
  •  ...performing team that delivers infrastructure and performance...  ...a Lead Infrastructure Engineer at JPMorganChase...  ...enterprise-authorized AI capabilities within the work environment to accelerate infrastructure analysis...  ...health care coverage, on-site health and wellness centers... 
    Permanent employment

    J.P. Morgan

    Palo Alto, CA
    14 days ago
  • $248k - $391k

     ...graphics, PC gaming, and accelerated computing for 30...  ...unlimited potential of AI to define the next...  ...skilled Principal Software Engineer to join our dynamic...  ...performance of our infrastructure both on‑prem and in the...  ...engineering, site reliability, or systems architecture... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  • $188k - $275k

    CoreWeave is The Essential Cloud for AI™. Built for pioneers by...  ..., CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute...  ...environments, enjoy tackling challenging engineering problems, and are excited by... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Worldwide
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • $200k - $322k

     ...systems telemetry, and cloud infrastructure operations. You will play a...  ...will work closely with our Engineering, Infrastructure, and Software...  ...troubleshoot, debug, and manage AI infrastructure effectively....  ..., and trust across the accelerated computing ecosystem. Providing... 
    Worldwide

    2100 NVIDIA USA

    Santa Clara, CA
    2 days ago
  • $124k - $195.5k

     ...computer graphics, PC gaming, and accelerated computing for more than 25...  ...the unlimited potential of AI to define the next era of...  ...Marketing Manager - Data Center Infrastructure Specialist to join our...  ...Marketing, Computer Science, Engineering, or a related field (or equivalent... 

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  •  ...DESCRIPTION Elevate your engineering prowess to...  ...among the top echelon in site reliability.  As a Senior Lead...  ...Chase within the Infrastructure Platforms and Foundational...  ...-authorized AI capabilities within the...  ...work environment to accelerate reliability design and... 

    J.P. Morgan

    Palo Alto, CA
    3 days ago
  • $156k - $229k

    Senior DFT Engineer, Test Infrastructure, Google Cloud Sunnyvale, CA, USA Job Level: Mid Minimum...  ...’ll work to shape the future of AI/ML hardware acceleration. You will have an opportunity to...  ...unparalleled scale, efficiency, reliability and velocity. Our customers include... 
    Full time
    Worldwide

    Google

    Sunnyvale, CA
    1 day ago
  •  ...forefront of revolutionizing AI computing by reengineering infrastructure at the system level. Our...  ...in efficient, more reliable computing at a fraction...  ...seeking a skilled Presales Engineer to be the technical...  ...hands-on experience in GPU acceleration, Kubernetes, Terraform,... 
    Work at office
    Remote work

    FlexAI

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Site Reliability Engineering -- AI Accelerator Infrastructure. Be the first to apply!