Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation

Reqroute

Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation
Location :- Santa Clara, CA (Hybrid)

Work Authorization: USC/GC only

Position Summary & Job Description:-

Client is looking for an experienced Site Reliability Engineer (SRE) to join our Infrastructure Platform Engineering team. In this role, you will help design, scale, and secure our enterprise on-premises and cloud hybrid infrastructure, driving high availability, automation, and operational excellence.

You will work at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms, and automated pipelines. If you are passionate about reducing toil through code, leveraging modern AI agent frameworks (like Model Context Protocol / MCP), and ensuring 24/7 system reliability across massive fleet environments, we want to hear from you.

Key Responsibilities

  • Infrastructure Reliability & Server Management: Architect, manage, and scale robust on-premises infrastructure and server fleets, ensuring high availability, performance optimization, and rigorous incident management.
  • Configuration Management & Automation: Drive configuration management across our environment using Chef (Cinc) and Infrastructure as Code (IaC) principles to ensure zero-drift and consistent deployments.
  • CI/CD & GitOps Pipelines: Design and maintain secure, scalable CI/CD pipelines (GitLab CI/CD) and GitOps workflows for automated system configuration, package rollout, and patch management.
  • AI Agent Tooling & Service Building: Build and integrate next-generation internal tools and services utilizing AI agent frameworks and LLM tooling (such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows.
  • Observability & Telemetry: Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies.
  • Cross-Platform Support: Partner with Windows and Linux engineering teams to maintain secure, compliant server and client environments, enforcing security standards (CIS benchmarks) and automated patching.

Qualifications & Required Skills

  • Experience: 5+ years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations within large-scale enterprise environments.
  • Configuration Management: Deep expertise in Chef (or Cinc) cookbook development, serverless execution modes, and automated provisioning.
  • Infrastructure & Server Operations: Strong mastery of on-premises infrastructure, server management, hardware provisioning, and operating systems architecture.
  • CI/CD & Automation: Proven track record of building automated CI/CD pipelines and GitOps workflows using modern version control (Git).
  • AI Agent & Tool Building Experience: Hands-on experience building internal microservices, tools, or workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks.
  • Observability & Reliability: Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability.
  • OS Familiarity: Experience managing and securing Windows infrastructure (alongside Linux) (Note: Windows experience is a key supporting requirement, but the core focus remains on SRE, Chef, and on-prem/hybrid infrastructure).
  • Scripting & Development: Proficiency in languages such as Python, Go, PowerShell, or Bash for automation and tooling development.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation in Santa Clara, CA vacancy
  • $55 - $60 per hour

     ...client is looking for an experienced Site Reliability Engineer (SRE) to join the Infrastructure Platform Engineering team. In...  ..., driving high availability, automation, and operational excellence. The...  ...infrastructure management and cutting-edge agentic AI tooling, building robust... 
    Suggested
    Temporary work
    Local area

    CYNET SYSTEMS

    Santa Clara, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) Location: Santa Clara Valley (Cupertino), California, Hybrid. Duration: 6+ Months Job...  ...internet services. Build and run systems, infrastructure and applications through automation. Participate in periodic on-call duties.... 
    Suggested

    Zortech Solutions

    Cupertino, CA
    5 days ago
  • $170k - $277k

     ...Senior Principal Engineer/Architect to serve...  ...for our global SRE and Platform Engineering...  ...bridge between infrastructure products and...  ...requirements into automated, intelligent self...  ...are not just reliable, but fully self-healing...  ..., LLMs, and agentic workflows into the... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Palo Alto Networks, Inc.

    Santa Clara, CA
    4 days ago
  • $208k - $333.5k

     ...Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused on designing...  ...continuous delivery, and automation. As an Engineering...  ...self-service platforms, infrastructure-as-code, standardized delivery...  ...with AI agents, agentic workflows, or intelligent... 
    Suggested
    Full time

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...the effort in crafting agentic AI systems that...  ...critically important infrastructure. Be part of a dynamic...  ...standards in AI-enhanced automation!What you'll be doing:...  ...closely with GPU SW kernel engineers to pinpoint high-...  ...that are fast, reliable, and meaningfully better... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $114.4k - $124.8k

     ...least 5 years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations in large-scale...  ...serverless execution modes, and automated provisioning Strong knowledge...  ...tooling, LLM orchestration, or agentic frameworks Expertise in end... 
    Hourly pay
    Full time
    Temporary work

    CYNET SYSTEMS

    Santa Clara, CA
    3 days ago
  • $174k - $252k

     ...through mechanisms like automation, and evolve...  ...changes that improve reliability and velocity....  ...Computer Science, Engineering, a related field,...  ...Science or Engineering.Site Reliability Engineering (SRE) is what you get...  ...by the Technical Infrastructure team to keep it running... 

    Google

    Sunnyvale, CA
    1 day ago
  • $105k - $155k

     ...company for AI and Bitcoin mining infrastructure. Bitdeer is committed to...  ...protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other...  ...As an entry-level Software Engineer on the SRE / Monitoring Platform... 
    Remote job
    Full time
    Contract work
    Temporary work
    Internship
    Local area

    Bitdeer

    San Jose, CA
    a month ago
  •  ...Description Developer & Infrastructure Expert Role Type:...  ..., DevOps, SRE, and platform engineering. You will test AI-generated...  ...for accuracy and reliability. Work with AWS,...  ...Infrastructure Site Reliability...  ...workplace connectors, automation, or AI-powered infrastructure... 
    For contractors
    Remote work

    YO AI Labs

    San Jose, CA
    a month ago
  •  ...Position: Site Reliability Engineering (SRE) Location: Santa Clara, CA (Onsite) Duration: W2 / C2C Contract Experience: 10+ Years...  ..., Jenkins, AWS CodeBuild, AWS CodeDeploy • WS Cloud infrastructure experience including EC2, SSM, vulnerability management... 
    Contract work
    Immediate start

    Syntricate Technologies

    Santa Clara, CA
    3 days ago
  •  ...Site Reliability Engineer (SRE) Share Contractual Sunnyvale, CA PDT - 8450 8-10 Overview: *Must...  ...Reliability Engineering, DevOps or infrastructure focused role • Advanced experience...  ...on experience with CI/CD systems • Automation advocate - you truly believe in... 

    Purple Drive

    Sunnyvale, CA
    3 days ago
  •  ...Title: Site Reliability Engineer (SRE) Location: Location: Sunnyvale, CA (3x/ week onsite...  ...and implement resilient and scalable infrastructure solutions. Operate, monitor,...  ...performance. Develop and implement automation to provision, configure, deploy,... 
    Contract work

    AceStack LLC

    Sunnyvale, CA
    2 days ago
  •  ...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and...  ...change, and build the automation and frameworks that make the...  ...automation direction and infrastructure architecture decisions...  ...7+ years of experience in SRE, cloud operations, and systems... 

    Outshift by Cisco

    San Jose, CA
    8 hours ago
  •  ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing...  ...more bounded contexts of the NeoCloud SRE platform — the multi-region substrate that...  .... Alert, Correlation & SLO: alert-engine-framework, alert-correlation, slo-framework... 
    Full time
    Contract work
    Local area

    Bitdeer

    San Jose, CA
    5 days ago
  • $152k - $241.5k

     ...and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering...  ....What you’ll be doingBuild automation for bare-metal provisioning...  ...Experience managing production reliability through on-call duties,... 
    Permanent employment
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing...  ...that turns novel incidents into new automations. NeoCloud is building an AI-operated...  ...-outs, rack and stack, labeling (on-site roles). Execute structured shift handoffs... 
    Full time
    Local area
    Shift work
    Night shift

    Bitdeer

    San Jose, CA
    5 days ago
  • $100k - $200k

     ...US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be...  ...systems. The ideal candidate is passionate about cloud infrastructure, automation, and building reliable, production-grade environments... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for...  ...on' role — you will build the automation, observability, and tooling that allows... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    1 day ago
  •  ...Role : SRE Engineer Location : San Jose, CA (ONSITE) FULL TIME ONLY...  ...procedures • Writing Python code to automate and develop small functionalities •...  ...time data processing, data Lakehouse infrastructure. • You have experience with Python writing... 
    Full time

    AceStack LLC

    San Jose, CA
    1 day ago
  •  ...Position- SRE Engineer Duration-Contract Location- San Jose, C JD Roles & Responsibilities...  ...-ups, DR planning • Creating and supporting automation scripts (shell/ansible/python) for infrastructure deployments, validations and monitoring to... 
    Contract work
    Immediate start

    Syntricate Technologies

    San Jose, CA
    3 days ago
  • $175k - $229k

     ...DevOps Engineer Instrumental builds the manufacturing...  ...years of DevOps or SRE experience deploying...  ...on public cloud infrastructure, AWS preferred. ~ Expert...  ...ongoing performance, reliability and efficiency. ~...  ...operating systems. Automation, automation,... 

    Instrumental Inc

    Palo Alto, CA
    5 days ago
  •  ...SRE Engineer Location: Sunnyvale CA Rate: DOE Duration: 12+ Months What You Will Do: Identify, develop and execute opportunities...  ...enterprise systems. ~ Proven track record of driving test automation to single click deployments. ~ Java experience with SQL.... 

    Redolent

    Sunnyvale, CA
    3 days ago
  •  ...SRE Engineer St Louis, MO (Onsite from day 1) Client Required Skills: • Bachelor's Degree in Computer Science, Computer Systems...  ...• Experience with web applications and distributed systems infrastructure. • Excellent verbal and written communication to a variety... 

    Omega Solutions

    Santa Clara, CA
    4 days ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on...  ...systems, networking, cloud infrastructure, Kubernetes, databases,...  ...and rapidly. Our engineers automate repetitive operational...  ...experience building AI agents, agentic workflows, or intelligent... 
    Full time

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are looking for an...  ...platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill.... 

    Eitacies Inc

    Santa Clara, CA
    a month ago
  •  ...SRE Devops Engineer Rootshell Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location: Santa... 
    Full time
    Local area
    Remote work

    Rootshell Inc

    Santa Clara, CA
    2 days ago
  • $175k - $265k

     ...possibilities of AI. Role Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on — colocation facilities...  ...a core member of that team, responsible for reliability, automation, and observability across colo, on-premises... 

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $195k - $285k

     ...inference silicon, and the infrastructure underpinning our engineering organization must be as reliable and scalable as the...  ...and leads d-Matrix's Site Reliability...  ...DoBuild and lead the SRE function from scratch...  ...design.Drive IaC-first automation (Terraform, Ansible)... 
    Remote work

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $160k - $192k

     ...is the Kubernetes and AI infrastructure platform behind some of the...  ...About the Job As an AI engineer, you will be part of internal...  ...that help promote agentic workflows and automations that improve employee productivity...  ...as a Software Engineer, SRE, ML Engineer, Solutions... 
    Work at office
    Immediate start
    Remote work
    Work from home
    Work visa

    Spectro Cloud

    San Jose, CA
    3 days ago
  • $216.78k - $390.2k

     ...governance framework for Agentic AI across the...  ...authority for Agentic AI Design Engineering and will lead the transformation...  ...grids, industrial automation, and 5G and cloud infrastructure. With a highly differentiated...  ...characterization, test, reliability, applications, and... 
    Local area

    Onsemi

    San Jose, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation. Be the first to apply!