Sr Staff Site Reliability Engineer, AI Infrastructure
$175k - $265kd-Matrix
At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together, we can help shape the endless possibilities of AI. Role Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. You will own systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments. This is hands-on, high-ownership work: you'll build and operate real infrastructure, not manage tickets.What You Will DoOwn reliability and availability across colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.Perform hands-on infrastructure work — server provisioning, OS configuration, networking, storage, and hardware troubleshooting — from bare metal through auto-scaling Kubernetes environments.Lead capacity planning and hardware lifecycle management for your domains, and track cloud spend to support FinOps and workload placement decisions.Drive all provisioning, deployment, and operational changes through Terraform and/or Ansible rather than manual steps, and contribute to shared IaC modules used across the global SRE and data center services teams.Build and document automation that eliminates toil — host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.Design and maintain monitoring, alerting, and SLIs (Prometheus/Grafana, DataDog, Splunk, or equivalent), contributing to AIOps-driven detection workflows.Participate in on-call rotation, triaging and resolving incidents from bare metal to application layer, and produce high-quality RCAs for P0/P1 incidents.Support platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks.What You Will BringBachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 7+ years in SRE, infrastructure engineering, or systems administration.Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure — networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.Production IaC experience with Terraform and/or Ansible — writing and maintaining configurations, not just running existing playbooks.Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking.Experience with observability tooling — Prometheus/Grafana, DataDog, Splunk, or equivalent — including building dashboards and writing alert rules.Production-quality Python and/or Bash scripting, paired with incident response experience: structured triage, RCA production, and follow-through on action items.Preferred QualificationsExperience operating customer-facing infrastructure or platform services with external reliability expectations.Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem.Experience deploying and operating AI-driven infrastructure tools — AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics — in production.HPC job scheduler experience: Slurm, LSF, or equivalent.Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink.Experience with large-scale infrastructure automation — host lifecycle management, fleet auto-healing, or AIOps-driven operations — building tooling that reduces manual intervention, not just running it.Equal Opportunity Employment Policyd-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation. Compensation Range: $175K - $265KLocationSanta ClaraEmployment TypeFull timeLocation TypeHybridDepartmentG&ACompensation$175K – $265K • Offers Equity • Offers BonusThe pay range below is for all roles at this level across all US locations and functions. Individual pay rates depend on a number of factors—including the role’s function and location, as well as the individual’s knowledge, skills, experience, education, and training. We also offer incentive opportunities that reward employees based on individual and company performance. This is in addition to our diverse package of benefits centered around the wellbeing of our employees and their loved ones. In addition to the usual Medical/Dental/Vision/401k, our inclusive rewards plan empowers our people to care for their whole selves. An investment in your future is an investment in ours.
- ...Powered by the Illumio AI Security Graph, our... ...cyber resilience for the infrastructure, systems, and... ...running. Location: 5 on-site days a week in Sunnyvale... ...Team's Vision: Our Engineering team is shaping the... ...experienced Senior Site Reliability Engineer (SRE) with a...SeniorWork experience placementImmediate start
- ...Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive... ...collection-monitor. Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework, default M-series...SeniorFull timeContract workLocal area
$114.4k - $124.8k
...Requirements: At least 5 years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations in large-scale enterprise environments... ...internal microservices, tools, or workflows with AI agent tooling, LLM orchestration, or agentic...SeniorHourly payFull timeTemporary work$165.5k - $289.6k
...It all started when engineer Fred Luddy wrote code that automated... ...work. Today, ServiceNow is the AI control tower for business reinvention... ...a highly experienced Senior Staff Cloud FinOps Analyst to lead... ...Cloud Platform, private infrastructure, and emerging AI services....SeniorFull timeWork at officeImmediate startRemote workFlexible hours$195k - $285k
...unleashing the potential of generative AI to power the transformation of... ...AI inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the chips we... ...builds and leads d-Matrix's Site Reliability Engineering...SuggestedRemote work$165.5k - $289.6k
...Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried... ...they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform...SeniorFull timeWork at officeImmediate startRemote workFlexible hours$262k - $364k
...architecture and design of the inference and training AI infrastructure from SRE side, ensuring it is reliable, scalable, cost effective and performant, while... ...:Master's degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines software and...Senior- ...Job Title : Senior Site Reliability Engineer Location : Santa Clara, CA Contract ENGAGEMENT SUMMARY The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response...SeniorContract work
$132.6k - $214.5k
...and Inclusion. We weave AI into the fabric of... ...collaborate closely with our engineering teams to develop... ...health. As a Senior Staff SRE with the Cortex Observability... ...GCP, to optimize our infrastructure, leveraging cloud-... ...and ensure the reliability and availability of our...SeniorFull timeWork at officeVisa sponsorshipWork visa$152k - $241.5k
...DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience... ...through repair.Experience managing production reliability through on-call duties, incident response, observability...SeniorPermanent employmentFull time$192.4k - $275.8k
...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service... ...the team's automation direction and infrastructure architecture decisions with regional... ...connect and protect organizations in the AI era – and beyond. We’ve been...SeniorFull timeTemporary workLocal areaFlexible hours$198k - $326k
...responsible for scaling LinkedIn's AI model training, feature engineering and serving with hundreds of... ...billions of user queries.Model Training Infrastructure: As an engineer on the AI Training... ...GPU inference at scale.As a Sr. Staff Software Engineer, you will have first...SeniorFor contractorsWork at officeFlexible hours$170k - $277k
...Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the... ...great outcomes. Job Summary We are seeking innovative engineers to design and develop security features for our next‑generation...SeniorFull timeWork at office$183.6k - $297k
...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and... ...drives great outcomes. The Team Engineering - Our engineering team is at the core of... ...security services and networking infrastructure. Technical Leadership: Lead complex engineering...SeniorFull timeWork at office$170k - $277k
...Integrity, and Inclusion. We weave AI into the fabric of everything... ...visionary Senior Principal Engineer/Architect to serve as the... ...a strategic bridge between infrastructure products and operational excellence... ...global systems are not just reliable, but fully self-healing....SeniorFull timeWork at officeVisa sponsorshipWork visaFlexible hours$114.91k - $135.19k
...enterprises. For more information, visit aviatrix.ai.ABOUT THE ROLE: Senior Cloud Network EngineerAs a Sr. Cloud Network Engineer, you will be a part of Aviatrix's Customer... ...BGP, IPsec VPN, virtualization, Linux, and infrastructure software.Experience with Amazon Web...SeniorFull timeTemporary workWork experience placementLocal areaFlexible hoursShift workWeekend workDay shift$91.4k - $187k
The Sr. Network Developer will play a crucial role... ...together the data, infrastructure, applications, and expertise... ...saving care. And with AI embedded across our... ...teams, including engineering, product management, and... ...developed with scalability, reliability, and security in mind....SeniorTemporary workFlexible hours$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building... ...systems, networking, cloud infrastructure, Kubernetes, databases, capacity management... ...the technical direction of NVIDIA’s AI Platform Runtime and lead reliability...Full time$168.2k - $310.1k
...'re looking for a Senior Cloud Security Engineer to help define, build, and scale Adobe's... ...more cloud security domains, such as cloud infrastructure security, cloud logging, identity and... ...businesses to turn ideas into impact, powered by AI and driven by human ingenuity.Our 30,000...SeniorFull timeTemporary workLocal areaImmediate startWorldwide$140k - $215k
...with the world’s most advanced AI-native platform. We work on... ..., you would ostensibly be the Sr backend developer who is also... ...contract, and chaos/resilience engineering. Cybersecurity experience is a... ...requiring 2-3 days per week on-site in one of the posted locations...SeniorFull timeContract workWork experience placementWork at officeLocal area2 days per week3 days per week$151.6k - $245.3k
...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use... ...Career Palo Alto Networks runs a large hybrid infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the...Full timeWork at office$110k - $175k
...most advanced electronic devices and IT infrastructure, enabling enhanced performance and user... ...the Role : As a Firmware Validation Engineer at SK hynix memory solutions, you will... ...JavaScript or python. Good understanding of AI development tools. Fast learner with...Senior$55 - $60 per hour
...Overview: Our client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In this role, the... ...infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms...SeniorTemporary workLocal area$230k - $250k
...shaping the future of network reliability, security, and AI‑ready operations. About... ...building the reliability engineering function at Forward —... ...closely with engineering, infrastructure, and product to ensure our... ...6+ years of experience in site reliability engineering, DevOps...Night shift- ...Technical/Functional Skills: 2+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a related role supporting cloud-based... ...application platforms. Experience leveraging AI-assisted development tools to improve software...Full timeWorldwide
$124.92k - $171.77k
...applications, including high-growth ones in AI datacenters, automated driving,... ...performance, smaller size, lower power, and better reliability. With more than 4 billion devices shipped... ...are seeking a hands-on Senior PCB Layout Engineer to own the physical implementation of...SeniorFlexible hours$230k - $250k
...foundation for autonomous networking, giving engineers and AI agents the ability to know the... ...been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a... ...will work closely with engineering, infrastructure, and product to ensure our platform...Night shift$149.8k - $224.6k
...responsibilities of a Technical Support Engineer within a SaaS (Software as a... ...with a growing focus on Site Reliability Engineering (SRE). The... ...run, support, and scale an AI Security Public SaaS... ...building scalable, resilient infrastructure to support AI inference workloads...Local area- ...Powered by the Illumio AI Security Graph, our... ...cyber resilience for the infrastructure, systems, and... ...running. Location: 5 On-Site Days a Week in Sunnyvale... ...CA Headquarters Our Engineering team is driven by a culture... ...on enhancing system reliability and scalability of...Work experience placementImmediate start
$170k - $200k
...Site Reliability Engineer We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible... ...maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability,...Full timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr Staff Site Reliability Engineer, AI Infrastructure. Be the first to apply!
- software engineer staff Santa Clara, CA
- technology administrator Santa Clara, CA
- assistant engineer Santa Clara, CA
- staff engineer Santa Clara, CA
- senior staff systems engineer Santa Clara, CA
- senior staff engineer Santa Clara, CA
- engineering aide Santa Clara, CA
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA




