Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr Staff Site Reliability Engineer, AI Infrastructure

$175k - $265k
Full-time

d-Matrix

At d-Matrix , we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive , and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together , we can help shape the endless possibilities of AI.

Role Overview

d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. You will own systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments. This is hands-on, high-ownership work: you'll build and operate real infrastructure, not manage tickets.

What You Will Do

  • Own reliability and availability across colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.

  • Perform hands-on infrastructure work — server provisioning, OS configuration, networking, storage, and hardware troubleshooting — from bare metal through auto-scaling Kubernetes environments.

  • Lead capacity planning and hardware lifecycle management for your domains, and track cloud spend to support FinOps and workload placement decisions.

  • Drive all provisioning, deployment, and operational changes through Terraform and/or Ansible rather than manual steps, and contribute to shared IaC modules used across the global SRE and data center services teams.

  • Build and document automation that eliminates toil — host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.

  • Design and maintain monitoring, alerting, and SLIs (Prometheus/Grafana, DataDog, Splunk, or equivalent), contributing to AIOps-driven detection workflows.

  • Participate in on-call rotation, triaging and resolving incidents from bare metal to application layer, and produce high-quality RCAs for P0/P1 incidents.

  • Support platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks.

What You Will Bring

  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 7+ years in SRE, infrastructure engineering, or systems administration.

  • Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure — networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.

  • Production IaC experience with Terraform and/or Ansible — writing and maintaining configurations, not just running existing playbooks.

  • Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking.

  • Experience with observability tooling — Prometheus/Grafana, DataDog, Splunk, or equivalent — including building dashboards and writing alert rules.

  • Production-quality Python and/or Bash scripting, paired with incident response experience: structured triage, RCA production, and follow-through on action items.

Preferred Qualifications

  • Experience operating customer-facing infrastructure or platform services with external reliability expectations.

  • Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem.

  • Experience deploying and operating AI-driven infrastructure tools — AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics — in production.

  • HPC job scheduler experience: Slurm, LSF, or equivalent.

  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink.

  • Experience with large-scale infrastructure automation — host lifecycle management, fleet auto-healing, or AIOps-driven operations — building tooling that reduces manual intervention, not just running it.

Equal Opportunity Employment Policy

d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.

d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Sr Staff Site Reliability Engineer, AI Infrastructure in Santa Clara, CA vacancy
  •  ...Powered by the Illumio AI Security Graph, our...  ...cyber resilience for the infrastructure, systems, and...  ...running. Location: 5 on-site days a week in Sunnyvale...  ...Team's Vision: Our Engineering team is shaping the...  ...experienced Senior Site Reliability Engineer (SRE) with a... 
    Senior
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  •  ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response... 
    Senior
    Contract work

    VDart

    Santa Clara, CA
    21 hours ago
  • $132.6k - $214.5k

     ...and Inclusion. We weave AI into the fabric of...  ...collaborate closely with our engineering teams to develop...  ...health. As a Senior Staff SRE with the Cortex Observability...  ...GCP, to optimize our infrastructure, leveraging cloud-...  ...and ensure the reliability and availability of our... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $192.4k - $275.8k

     ...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...the team's automation direction and infrastructure architecture decisions with regional...  ...connect and protect organizations in the AI era – and beyond. We’ve been... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    Cisco

    San Jose, CA
    1 day ago
  • $195k - $285k

     ...unleashing the potential of generative AI to power the transformation of...  ...AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we...  ...builds and leads d-Matrix's Site Reliability Engineering... 
    Suggested
    Full time
    Remote work

    d-Matrix

    Santa Clara, CA
    4 days ago
  •  ...Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive...  ...collection-monitor. Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series... 
    Senior
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    2 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building...  ...systems, networking, cloud infrastructure, Kubernetes, databases, capacity management...  ...the technical direction of NVIDIA’s AI Platform Runtime and lead reliability... 
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $152k - $287.5k

     ...DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience...  ...through repair. ~ Experience managing production reliability through on-call duties, incident response,... 
    Senior
    Permanent employment
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $229.9k - $262.4k

     ...Sr. Lead AI Engineer (Gen AI Platform Services) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been...  .... Our investments in technology infrastructure and world‑class talent — along with... 
    Senior
    Local area

    Capital One National Association

    San Jose, CA
    3 days ago
  • $230k - $250k

     ...shaping the future of network reliability, security, and AI‑ready operations. About...  ...building the reliability engineering function at Forward —...  ...closely with engineering, infrastructure, and product to ensure our...  ...6+ years of experience in site reliability engineering, DevOps... 
    Night shift

    Forward

    Santa Clara, CA
    14 hours ago
  • $230k - $250k

     ...foundation for autonomous networking, giving engineers and AI agents the ability to know the...  ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a...  ...will work closely with engineering, infrastructure, and product to ensure our platform... 
    Night shift

    Forward Networks Inc

    Santa Clara, CA
    1 day ago
  • $149.8k - $224.6k

     ...responsibilities of a Technical Support Engineer within a SaaS (Software as a...  ...with a growing focus on Site Reliability Engineering (SRE). The...  ...run, support, and scale an AI Security Public SaaS...  ...building scalable, resilient infrastructure to support AI inference workloads... 
    Local area

    F5

    San Jose, CA
    4 days ago
  • $170k - $200k

     ...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible...  ...maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability,... 
    Full time
    Worldwide

    Virtue AI

    Sunnyvale, CA
    4 days ago
  •  ...responsibilities of a Technical Support Engineer within a SaaS (Software as a...  ...with a growing focus on Site Reliability Engineering (SRE). The ideal...  ...run, support, and scale an AI Security Public SaaS...  ...building scalable, resilient infrastructure to support AI inference workloads... 
    Work at office
    Local area
    Remote work
    Work from home

    F5 Networks

    San Jose, CA
    1 day ago
  •  ...Powered by the Illumio AI Security Graph, our...  ...cyber resilience for the infrastructure, systems, and...  ...running. Location: 5 On-Site Days a Week in Sunnyvale...  ...CA Headquarters Our Engineering team is driven by a culture...  ...on enhancing system reliability and scalability of... 
    Work experience placement
    Immediate start

    Illumio

    Sunnyvale, CA
    3 days ago
  • $110k - $130k

     ...working with the World's leading AI-first Quality Engineering Company? Ready to advance your career...  ...! We are looking for a Site Reliability Engineer to join our growing team...  ...Good understanding of hybrid infrastructure. Expertise with AWS. Expertise... 
    Casual work
    Local area
    Flexible hours

    QualiTest Group

    Santa Clara, CA
    2 days ago
  • $165.5k - $289.6k

     ...It all started when engineer Fred Luddy wrote code that automated...  ...work. Today, ServiceNow is the AI control tower for business reinvention...  ...a highly experienced Senior Staff Cloud FinOps Analyst to lead...  ...Cloud Platform, private infrastructure, and emerging AI services.... 
    Senior
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    13 days ago
  • $122.5k - $175k

     ...believe the future of work is Human + AI and are building an AI-native...  ...at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role...  ...Chief Architect, REIS in the Cloud Infrastructure & Operations department. You are an... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    1 day ago
  •  ...Integrity, and Inclusion. We weave AI into the fabric of everything we do...  ...Lead, mentor, and develop a team of Site Reliability/Production Engineers, providing technical direction,...  ...health of critical Cortex services and infrastructure. Drive improvements in... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    14 hours ago
  • $139k - $242k

     ...is The Essential Cloud for AI™. Built for pioneers by pioneers...  ...CoreWeave combines superior infrastructure performance with deep...  ...services into a cohesive, high-reliability engine of fleet management. This group...  ...problems of scale for multi-site deployment and management of... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    11 days ago
  • $152k - $228k

     ...Role The Platform Engineering team builds, secures...  ...operates scalable infrastructure for cloud-managed SaaS...  ...at customer sites. The Site Reliability Engineering discipline...  ...with DevSecOps. As Staff Site Reliability Engineer...  ..., and creating AI Ops workflows for triage... 
    Permanent employment
    Contract work
    Work at office
    Remote work

    NightDragon Acquisition Corp.

    Santa Clara, CA
    14 hours ago
  • $122.5k - $175k

     ...believe the future of work is Human + AI and are building an AI-native...  ...Zscaler. Role We are looking for a Staff Site Reliability Engineer to join our team. This is a hybrid role...  ...Chief Architect, REIS in the Cloud Infrastructure & Operations department. You are an... 
    Full time
    Work at office
    Local area
    3 days per week

    Zscaler

    San Jose, CA
    3 days ago
  • $140k - $165k

     ...most advanced electronic devices and IT infrastructure, enabling enhanced performance and user...  ...Why Join Us? Build foundational AI infrastructure that powers next-gen enterprise...  ...the Role: We are seeking a hands-on AI Engineer to design, deploy, and maintain on-prem... 
    Senior

    SK hynix memory solutions America Inc.

    San Jose, CA
    8 days ago
  • $140k - $215k

     ...world’s most advanced AI-native platform. We work...  ...would ostensibly be the Sr backend developer who...  ..., and chaos/resilience engineering. Cybersecurity experience...  ...2-3 days per week on-site in one of the posted locations...  ...wide testing tools and infrastructure Define and drive... 
    Senior
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    Sunnyvale, CA
    2 days ago
  • $183k - $247.6k

     ...to shape the future of AI? Join the team...  ...operate next-generation infrastructure that powers breakthrough...  ...hardware, and network engineers, supply chain specialists...  ...Own end-to-end system reliability, proactively identifying...  ..., supervisors, and staff; adhere to standards of... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    3 days ago
  • $175k - $229k

     ...DevOps Engineer Instrumental builds the manufacturing acceleration...  ..., repair cycles—and our AI engines identify insights that...  ...platforms on public cloud infrastructure, AWS preferred. ~ Expert knowledge...  ...ensure ongoing performance, reliability and efficiency. ~ Network/... 
    Senior

    Instrumental Inc

    Palo Alto, CA
    2 days ago
  • $285k - $345k

     ...and curation. We are seeking engineers capable of developing, designing...  ...deploying highly scalable, reliable applications, tools, and...  ...and build large-scale video infrastructure powering Roku's Live and VOD...  ...time off. How will I use AI at Roku? At Roku, we don’... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $160k - $200k

     ...seek talented, passionate, and committed engineers, technologists, and business leaders to...  ...network management software for large-scale AI and data center fabrics. In this role,...  ...hardening. Drive performance, reliability, and debuggability improvements across AI... 
    Senior
    Worldwide
    Shift work

    Supermicro

    San Jose, CA
    1 day ago
  • $160k - $200k

     ...seek talented, passionate, and committed engineers, technologists, and business leaders to join...  ...is seeking a top-notch hands-on Sr. Software Engineer to work on PCIe, SAS/SATA...  ...Responsibilities:Create and maintain several end-to-end AI agent pipelines that handle high-volume... 
    Senior
    Worldwide

    Super Micro Computer

    San Jose, CA
    21 hours ago
  •  ...combining our expertise across connectivity, AI, security and more, we'll map a new way forward...  ...Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that underpins the build pipelines that allow... 
    Full time
    Contract work

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr Staff Site Reliability Engineer, AI Infrastructure. Be the first to apply!