Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr Staff Site Reliability Engineer, AI Infrastructure

Jobleads-US

At d-Matrix , we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.

Role Overview

d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. You will own systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments. This is hands-on, high-ownership work: you'll build and operate real infrastructure, not manage tickets.

What You Will Do

  • Own reliability and availability across colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.

  • Perform hands-on infrastructure work — server provisioning, OS configuration, networking, storage, and hardware troubleshooting — from bare metal through auto-scaling Kubernetes environments.

  • Lead capacity planning and hardware lifecycle management for your domains, and track cloud spend to support FinOps and workload placement decisions.

  • Drive all provisioning, deployment, and operational changes through Terraform and/or Ansible rather than manual steps, and contribute to shared IaC modules used across the global SRE and data center services teams.

  • Build and document automation that eliminates toil — host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.

  • Design and maintain monitoring, alerting, and SLIs (Prometheus/Grafana, DataDog, Splunk, or equivalent), contributing to AIOps-driven detection workflows.

  • Participate in on-call rotation, triaging and resolving incidents from bare metal to application layer, and produce high-quality RCAs for P0/P1 incidents.

  • Support platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks.

What You Will Bring

  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 7+ years in SRE, infrastructure engineering, or systems administration.

  • Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure — networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.

  • Production IaC experience with Terraform and/or Ansible — writing and maintaining configurations, not just running existing playbooks.

  • Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking.

  • Experience with observability tooling — Prometheus/Grafana, DataDog, Splunk, or equivalent — including building dashboards and writing alert rules.

  • Production-quality Python and/or Bash scripting, paired with incident response experience: structured triage, RCA production, and follow-through on action items.

Preferred Qualifications

  • Experience operating customer-facing infrastructure or platform services with external reliability expectations.

  • Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem.

  • Experience deploying and operating AI-driven infrastructure tools — AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics — in production.

  • HPC job scheduler experience: Slurm, LSF, or equivalent.

  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink.

  • Experience with large-scale infrastructure automation — host lifecycle management, fleet auto-healing, or AIOps-driven operations — building tooling that reduces manual intervention, not just running it.

Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We're committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Sr Staff Site Reliability Engineer, AI Infrastructure in Santa Clara, CA vacancy
  • $175k - $265k

     ...unleashing the potential of generative AI to power the transformation of...  ...Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on —...  ...member of that team, responsible for reliability, automation, and observability across... 
    Senior

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $183k - $247.6k

     ...accelerated servers powering AI/ML training and...  ...Cloud Hardware Development Engineer to define server...  ...Manufacturing Partner sites.The Ideal CandidateYou...  ...next-generation AI/ML infrastructure deployed in datacenters...  ...employees, supervisors, and staff; adhere to standards of... 
    Senior
    Local area
    Worldwide
    Flexible hours
    Day shift

    Amazon

    Cupertino, CA
    2 days ago
  • $165k - $265k

     ...goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging...  ...on Starshield's software and GPU infrastructure, you will design, operate and scale the...  ...validate, and productize solutions for AI clusters (100k+ GPU scale)Develop... 
    Senior
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  •  ...Powered by the Illumio AI Security Graph, our...  ...cyber resilience for the infrastructure, systems, and...  ...running. Location: 5 on-site days a week in Sunnyvale...  ...Team's Vision: Our Engineering team is shaping the future...  ...Senior Site Reliability Engineer (SRE) with a... 
    Senior
    Work experience placement

    Illumio

    Sunnyvale, CA
    5 days ago
  • $207.4k - $259.2k

     ...physical artificial intelligence (“AI”) solutions, and other technologies...  ...a highly experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In this...  ..., implement, and maintain robust infrastructure and automation solutions.ResponsibilitiesImplement... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Visa sponsorship

    Archer Aviation

    San Jose, CA
    2 days ago
  •  ...Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive...  ...collection-monitor. Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series... 
    Senior
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    9 days ago
  • $55 - $60 per hour

     ...Overview: Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the...  ...infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms... 
    Senior
    Temporary work
    Local area

    CYNET SYSTEMS

    Santa Clara, CA
    5 days ago
  •  ...Cloud, is a leader in AI cloud infrastructure serving tens of...  ...is currently Tuesday.Engineering at Lambda is responsible...  ...teams to improve service reliability and deployment...  ...years of experience in Site Reliability Engineering...  ...virtualization technologies, SR-IOV, and... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of...  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building...  ...Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago
  • $168k - $270.25k

     ....Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial...  ...with engineering teams to align infrastructure with their evolving needs, document best...  ...for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $165.5k - $289.6k

     ...It all started when engineer Fred Luddy wrote code that automated...  ...work. Today, ServiceNow is the AI control tower for business reinvention...  ...a highly experienced Senior Staff Cloud FinOps Analyst to lead...  ...Cloud Platform, private infrastructure, and emerging AI services.... 
    Senior
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    20 days ago
  • $148k - $235.75k

     ...into the unlimited potential of AI to define the next era of...  ....Join our team of innovative engineers who are building an AI Data Center...  ..., high-volume telemetry into reliable, job-centric insights and...  ...automation.Manage deployment infrastructure and packaging (Helm + Terraform... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

    NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global...  ...bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.Leading software... 
    Full time
    Work experience placement
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195k - $285k

     ...unleashing the potential of generative AI to power the transformation of...  ...AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we...  ...builds and leads d-Matrix's Site Reliability Engineering... 
    Remote work

    d-Matrix

    Santa Clara, CA
    3 days ago
  • $267k - $356k

     ...Cloud, is a leader in AI cloud infrastructure serving tens of...  ...Tuesday.Lambda's Storage Engineering team is the backbone...  ...industry, which means reliability and performance aren'...  ...across new and existing sites using tools such as...  ...drivers.Experience with SR-IOV and... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

     ...You’ll harness the power of AI to deliver groundbreaking solutions...  ...and network fabrics.Use IaC(Infrastructure‑as‑Code) and config...  ...lifecycle management, fleet reliability/auto-healing, E2E observability...  ..., or Ruby.Mentored other engineers and influenced technical direction... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $165.5k - $289.6k

    It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker...  ...meaningful work. Today, ServiceNow is the AI control tower for business reinvention....  ...within this posting, you are leaving this site and going to a third-party website where... 
    Senior
    Work at office
    Immediate start
    Remote work
    Flexible hours

    SmartRecruiters, Inc.

    Santa Clara, CA
    3 days ago
  • $262k - $364k

     ...architecture and design of the inference and training AI infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while...  ...:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and... 
    Senior

    Google

    Sunnyvale, CA
    1 day ago
  • $192.4k - $275.8k

     ...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...the team's automation direction and infrastructure architecture decisions with regional...  ...connect and protect organizations in the AI era - and beyond. We’ve been... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    5 days ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain...  ...that power NVIDIA’s next-generation AI-driven enterprise products and services...  ...Go, with a focus on automation and infrastructure-as-code.Experience with infrastructure... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $132.6k - $214.5k

     ...and Inclusion. We weave AI into the fabric of...  ...collaborate closely with our engineering teams to develop...  ...health. As a Senior Staff SRE with the Cortex Observability...  ...GCP, to optimize our infrastructure, leveraging cloud-...  ...and ensure the reliability and availability of our... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience...  ...through repair.Experience managing production reliability through on-call duties, incident response, observability... 
    Senior
    Permanent employment
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $100k

    Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations...  ...contributors of all seniorities.We are looking for a talented engineer to build and run the performance infrastructure shared by our RISC-V Software and RISC-V Performance... 
    Senior
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    1 day ago
  • $183.6k - $297k

     ..., Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact...  ...drives great outcomes.Job SummaryThe TeamEngineering - Our engineering team is at the core of our products and connected directly to... 
    Senior
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $137k - $156k

     ...seek talented, passionate, and committed engineers, technologists, and business leaders to...  ...focuses on multi-GPU server systems used for AI, HPC, enterprise computing, and...  ...validation, technical enablement, HPC, AI infrastructure, or a related field.Strong knowledge of... 
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    4 days ago
  • $124.92k - $171.77k

     ...applications, including high-growth ones in AI datacenters, automated driving,...  ...performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped...  ...are seeking a hands-on Senior PCB Layout Engineer to own the physical implementation of... 
    Senior
    Flexible hours

    SiTime

    Santa Clara, CA
    5 days ago
  •  ...builds the world's largest AI chip, 56 times larger than...  ...powered by the Wafer-Scale Engine (WSE). This team will help deliver world-class, ultra-reliable inference infrastructure for leading model builders...  ...and other frontier labs.As a Staff SRE, you will lead the engineering... 
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $183.6k - $297k

     ...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and...  ....Job SummaryThe TeamEngineering - Our engineering team is at the core of our products and...  ...security services and networking infrastructure.Technical Leadership: Lead complex engineering... 
    Senior
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  •  ...builds the world's largest AI chip, 56 times larger than GPUs...  ...the RoleWe're hiring a Staff Engineer to help lead, drive, and contribute...  ...for major technical areas.Reliability & Performance. Architect...  .... Partner with ML, Product, Infrastructure, and Cloud teams to translate... 
    Senior

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $193.3k - $261.5k

     ...the right moment. As AI models outgrow any single...  ....We're looking for an engineer to work at the...  ...leading edge of AI/ML infrastructure.About the team: You'd...  ...architecture (design patterns, reliability and scaling) of new...  ...employees, supervisors, and staff; adhere to standards... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr Staff Site Reliability Engineer, AI Infrastructure. Be the first to apply!