Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation
Reqroute
Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation
Location :- Santa Clara, CA (Hybrid)
Work Authorization: USC/GC only
Position Summary & Job Description:-
Client is looking for an experienced Site Reliability Engineer (SRE) to join our Infrastructure Platform Engineering team. In this role, you will help design, scale, and secure our enterprise on-premises and cloud hybrid infrastructure, driving high availability, automation, and operational excellence.
You will work at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms, and automated pipelines. If you are passionate about reducing toil through code, leveraging modern AI agent frameworks (like Model Context Protocol / MCP), and ensuring 24/7 system reliability across massive fleet environments, we want to hear from you.
Key Responsibilities
- Infrastructure Reliability & Server Management: Architect, manage, and scale robust on-premises infrastructure and server fleets, ensuring high availability, performance optimization, and rigorous incident management.
- Configuration Management & Automation: Drive configuration management across our environment using Chef (Cinc) and Infrastructure as Code (IaC) principles to ensure zero-drift and consistent deployments.
- CI/CD & GitOps Pipelines: Design and maintain secure, scalable CI/CD pipelines (GitLab CI/CD) and GitOps workflows for automated system configuration, package rollout, and patch management.
- AI Agent Tooling & Service Building: Build and integrate next-generation internal tools and services utilizing AI agent frameworks and LLM tooling (such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows.
- Observability & Telemetry: Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies.
- Cross-Platform Support: Partner with Windows and Linux engineering teams to maintain secure, compliant server and client environments, enforcing security standards (CIS benchmarks) and automated patching.
Qualifications & Required Skills
- Experience: 5+ years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations within large-scale enterprise environments.
- Configuration Management: Deep expertise in Chef (or Cinc) cookbook development, serverless execution modes, and automated provisioning.
- Infrastructure & Server Operations: Strong mastery of on-premises infrastructure, server management, hardware provisioning, and operating systems architecture.
- CI/CD & Automation: Proven track record of building automated CI/CD pipelines and GitOps workflows using modern version control (Git).
- AI Agent & Tool Building Experience: Hands-on experience building internal microservices, tools, or workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks.
- Observability & Reliability: Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability.
- OS Familiarity: Experience managing and securing Windows infrastructure (alongside Linux) (Note: Windows experience is a key supporting requirement, but the core focus remains on SRE, Chef, and on-prem/hybrid infrastructure).
- Scripting & Development: Proficiency in languages such as Python, Go, PowerShell, or Bash for automation and tooling development.
- ...company for AI and Bitcoin mining infrastructure. Bitdeer is committed to... ...protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other... ...As an entry-level Software Engineer on the SRE / Monitoring Platform...SuggestedFull timeContract workTemporary workInternshipLocal area
$208k - $333.5k
...Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused on designing... ...continuous delivery, and automation. As an Engineering... ...self-service platforms, infrastructure-as-code, standardized delivery... ...with AI agents, agentic workflows, or intelligent...SuggestedFull time$170k - $277k
...Senior Principal Engineer/Architect to serve... ...for our global SRE and Platform Engineering... ...bridge between infrastructure products and... ...requirements into automated, intelligent self... ...are not just reliable, but fully self-healing... ..., LLMs, and agentic workflows into the...SuggestedFull timeWork at officeVisa sponsorshipWork visaFlexible hours$184k - $287.5k
...the effort in crafting agentic AI systems that... ...critically important infrastructure. Be part of a dynamic... ...standards in AI-enhanced automation!What you'll be doing:... ...closely with GPU SW kernel engineers to pinpoint high-... ...that are fast, reliable, and meaningfully better...SuggestedFull time$174k - $252k
...through mechanisms like automation, and evolve... ...changes that improve reliability and velocity.... ...Computer Science, Engineering, a related field,... ...Science or Engineering.Site Reliability Engineering (SRE) is what you get... ...by the Technical Infrastructure team to keep it running...Suggested$104.4k - $171k
...mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of... ...stability engineering through automation and tooling. ~4. Executing production... ...with Linux environments or cloud infrastructure ~ Exceptional system diagnostic...- ...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and... ...change, and build the automation and frameworks that make the... ...automation direction and infrastructure architecture decisions... ...7+ years of experience in SRE, cloud operations, and systems...
$209.7k - $238.25k
...About our Vehicle Software Engineering Teams Our Vehicle... .... Within it, the Tools & Infrastructure team owns the reliability stack for vehicle software... ...operator tooling, and the automation that lets engineering and... ...this team is defining what SRE looks like for vehicle-...Full timeWork at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work3 days per week- ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing... ...more bounded contexts of the NeoCloud SRE platform — the multi-region substrate that... .... Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework...Full timeContract workLocal area
$152k - $241.5k
...and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering... ....What you’ll be doingBuild automation for bare-metal provisioning... ...Experience managing production reliability through on-call duties,...Permanent employmentFull time- ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing... ...that turns novel incidents into new automations. NeoCloud is building an AI-operated... ...-outs, rack and stack, labeling (on-site roles). Execute structured shift handoffs...Full timeLocal areaShift workNight shift
$100k - $200k
...US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be... ...systems. The ideal candidate is passionate about cloud infrastructure, automation, and building reliable, production-grade environments...Full time- ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are looking for an... ...platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill....
$248k - $396.75k
...Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on... ...systems, networking, cloud infrastructure, Kubernetes, databases,... ...and rapidly. Our engineers automate repetitive operational... ...experience building AI agents, agentic workflows, or intelligent...- ...Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote...Full timeLocal areaRemote work
$175k - $265k
...possibilities of AI. Role Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities... ...a core member of that team, responsible for reliability, automation, and observability across colo, on-premises...$195k - $285k
...inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the... ...and leads d-Matrix's Site Reliability... ...DoBuild and lead the SRE function from scratch... ...design.Drive IaC-first automation (Terraform, Ansible)...Remote work- ...a leader in AI cloud infrastructure serving tens of thousands... ...or Staff Software Engineer in Lambda’s Cloud... ...capacity and orchestration, reliability, and developer-facing... ..., including better automation, testing, guardrails,... ..., security, and SRE teams to resolve dependencies...Work at officeLocal areaWork from homeFlexible hours
$160k - $192k
...is the Kubernetes and AI infrastructure platform behind some of the... ...About the Job As an AI engineer, you will be part of internal... ...that help promote agentic workflows and automations that improve employee productivity... ...as a Software Engineer, SRE, ML Engineer, Solutions...Work at officeImmediate startRemote workWork from homeWork visa- ...Description Role 1 - Core Platform Engineer (L1, breadth-first) The... ...defense: incident response, triage, reliability, and automation across the full infrastructure stack. Day-to-day: Runs incident... ...Sourcing note: This is a deep SRE profile, not a pure generalist. The...Night shift
- ...Description Developer & Infrastructure Expert Role Type:... ..., DevOps, SRE, and platform engineering. You will test AI-generated... ...for accuracy and reliability. Work with AWS,... ...Infrastructure Site Reliability... ...workplace connectors, automation, or AI-powered infrastructure...Remote jobFor contractors
$216.78k - $390.2k
...governance framework for Agentic AI across the... ...authority for Agentic AI Design Engineering and will lead the transformation... ...grids, industrial automation, and 5G and cloud infrastructure. With a highly differentiated... ...characterization, test, reliability, applications, and...Local area- ...Cupertino, CA is seeking a Software Engineer Manager for the Agentic Evaluation Platform. You will lead... ...and evaluation harnesses to enable reliable, scalable Siri evaluations across Apple... ...own APIs, and partner with QE and infrastructure teams to deliver high-quality...
$188.5k - $282.7k
...with enterprise AI infrastructure.Nature of the Specialized... ...environments.Engineering secure "Landing Zones... ...) hierarchies and automated policy enforcement.... ...Architecture, or Site Reliability Engineering (SRE).Expert-level proficiency... ...and auditing agentic actions, enforcing...$193.3k - $261.5k
...frontier: a production-grade, agentic AI platform that automates catalog diagnostics —... ...Software Development Engineer who can operate at the intersection... ...close the accuracy and reliability loop for production-... ...of Amazon's technology infrastructure. We maintain...Hourly payInternshipLocal areaFlexible hours$165.2k - $223.6k
...is. We are building Agentic WorkSpaces: a platform... ...how you work, not just infrastructure that delivers... ...already use, delivered reliably from commercial regions... ...Software Development Engineer to design, build, and... ...through instrumentation, automated testing, and pipeline...InternshipWorldwideFlexible hours$70 - $100 per hour
...Job Title: Cloud SRE Engineer - Mandarin Bilingual Position Type... ...Cloud SRE Engineer to own the reliability, stability, and continuous... ...— spanning compute infrastructure (CVM/VMs), networking, and... ...tooling Develop scripts and automation tools to improve operational...Hourly payContract workTemporary workWork experience placement$104.9k - $174.7k
About the Role:The SRE role is... ...for improving the reliability, availability, performance... ...completion.Follow up with engineering, development,... ..., on-premises infrastructure, security, Kubernetes, automation, monitoring, and system... ...of experience in Site Reliability...Full timeLocal area$237.6k - $401.7k
...Software Engineer Manager: Agentic Evaluation Platform Cupertino, California... ...team applies large scale infrastructure, apple platforms and service... ...that let us evaluate Siri reliably and at scale. You'll own... ...to scoring - prioritizing automated triage and durable fixes over...Relocation$207k - $300k
...critical areas within SU SRE, mentoring team... ...members to enhance system reliability and efficiency.... ...of experience as a Site Reliability Engineer.3 years of experience... ...AI Agent, Google Infrastructure.Experience in large-... ...process improvement and automation.Site Reliability...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation. Be the first to apply!
- site reliability engineer Santa Clara, CA
- site reliability engineer sre Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- data infrastructure engineer Santa Clara, CA
- infrastructure engineer Santa Clara, CA
- remote infrastructure engineer Santa Clara, CA
- senior infrastructure engineer Santa Clara, CA
- infrastructure developer Santa Clara, CA
- infrastructure engineering manager Santa Clara, CA
- automation engineer Santa Clara, CA



