Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation

ReqRoute,Inc

Role :- Site Reliability Engineer (SRE) Infrastructure & Agentic Automation
Location :- Santa Clara, CA (Hybrid)

Work Authorization: USC/GC only

Position Summary & Job Description:-

Client is looking for an experienced Site Reliability Engineer (SRE) to join our Infrastructure Platform Engineering team. In this role, you will help design, scale, and secure our enterprise on-premises and cloud hybrid infrastructure, driving high availability, automation, and operational excellence.

You will work at the intersection of traditional infrastructure management and cutting-edge agentic AI tooling, building robust services, telemetry platforms, and automated pipelines. If you are passionate about reducing toil through code, leveraging modern AI agent frameworks (like Model Context Protocol / MCP), and ensuring 24/7 system reliability across massive fleet environments, we want to hear from you.

Key Responsibilities

  • Infrastructure Reliability & Server Management: Architect, manage, and scale robust on-premises infrastructure and server fleets, ensuring high availability, performance optimization, and rigorous incident management.
  • Configuration Management & Automation: Drive configuration management across our environment using Chef (Cinc) and Infrastructure as Code (IaC) principles to ensure zero-drift and consistent deployments.
  • CI/CD & GitOps Pipelines: Design and maintain secure, scalable CI/CD pipelines (GitLab CI/CD) and GitOps workflows for automated system configuration, package rollout, and patch management.
  • AI Agent Tooling & Service Building: Build and integrate next-generation internal tools and services utilizing AI agent frameworks and LLM tooling (such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows.
  • Observability & Telemetry: Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies.
  • Cross-Platform Support: Partner with Windows and Linux engineering teams to maintain secure, compliant server and client environments, enforcing security standards (CIS benchmarks) and automated patching.

Qualifications & Required Skills

  • Experience: 5+ years of experience in Site Reliability Engineering, Systems Engineering, or Infrastructure Operations within large-scale enterprise environments.
  • Configuration Management: Deep expertise in Chef (or Cinc) cookbook development, serverless execution modes, and automated provisioning.
  • Infrastructure & Server Operations: Strong mastery of on-premises infrastructure, server management, hardware provisioning, and operating systems architecture.
  • CI/CD & Automation: Proven track record of building automated CI/CD pipelines and GitOps workflows using modern version control (Git).
  • AI Agent & Tool Building Experience: Hands-on experience building internal microservices, tools, or workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks.
  • Observability & Reliability: Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability.
  • OS Familiarity: Experience managing and securing Windows infrastructure (alongside Linux) (Note: Windows experience is a key supporting requirement, but the core focus remains on SRE, Chef, and on-prem/hybrid infrastructure).
  • Scripting & Development: Proficiency in languages such as Python, Go, PowerShell, or Bash for automation and tooling development.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation in Santa Clara, CA vacancy
  •  ...company for AI and Bitcoin mining infrastructure. Bitdeer is committed to...  ...protects, and operates the fleet. The SRE Platform team builds the monitoring and automation substrate that every other...  ...As an entry-level Software Engineer on the SRE / Monitoring Platform... 
    Suggested
    Full time
    Contract work
    Internship
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    15 days ago
  • $169k - $338k

     ...Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering...  ...of next-generation agentic AI systems and intelligent automation solutions that ensure...  ...transformation of traditional SRE practices into AI-...  ...highly automated infrastructure that supports... 
    Suggested
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  • $167.7k - $245.2k

     ...approximately 2 days per week on-site at Cisco offices in...  ...intended, improving reliability and reducing risks....  ...Site Reliability Engineer (SRE), you will build,...  ...platform and production infrastructure. You will own the...  ...deployments, develop automation to improve operational... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Jose, CA
    1 day ago
  • $186.9k - $267.7k

     ...approximately 2 days per week on-site at Cisco offices in...  ...intended, improving reliability and reducing risks....  ...Site Reliability Engineer (SRE), you will provide...  ...strategy, lead major infrastructure initiatives, and drive...  ...across deployment automation, production operations... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...IT's Enterprise AI & Automation team to develop and expand...  ...enterprise-grade agentic AI systems at one of...  ...business results across engineering, IT, supply chain,...  ...candidate must grasp infrastructure aspects from Kubernetes...  ..., latency, cost, reliability, and safety. ~ Comprehensive... 
    Suggested

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  •  ...intelligence via additional agentic computation.Cerebras...  ...a high-performance SRE function to support...  ...by the Wafer-Scale Engine (WSE). This team...  ...world-class, ultra-reliable inference infrastructure for leading model builders...  ...their pain points, automate their toil, and... 
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  •  ...enterprise customers, blending Site Reliability Engineering, Systems Engineering, and...  ...change, and build the automation and frameworks that make the...  ...automation direction and infrastructure architecture decisions...  ...7+ years of experience in SRE, cloud operations, and systems... 

    Outshift by Cisco

    San Jose, CA
    4 days ago
  • $174.72k - $295.68k

     ...You will be a senior engineer on the team building...  ...systems, services, and automation that power how our...  ...operational systems. The agentic layer (NL intake,...  ...distributed systems, building reliable services and...  ...platform services and infrastructure with strong reliability... 
    Full time

    XPENG Motors

    Santa Clara, CA
    2 days ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain...  ...components to eliminate manual work through automation, performance tuning and growing...  ...Go, with a focus on automation and infrastructure-as-code.Experience with... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing...  ...that turns novel incidents into new automations. NeoCloud is building an AI-operated...  ...-outs, rack and stack, labeling (on-site roles). Execute structured shift handoffs... 
    Full time
    Local area
    Shift work
    Night shift

    Bitdeer Technologies Group

    San Jose, CA
    a month ago
  • $178k - $321k

     ...We are safe and reliable, backed by our Proof...  ...(Hive Mind), agentic workflows, data pipelines...  ...data and AI infrastructure everything else...  ...is a two-person engineering team: you deploy,...  ...apps, Drive-synced automation, locally scheduled...  ...external careers site.Notice:All... 

    OKX

    San Jose, CA
    1 day ago
  • $100k - $200k

     ...US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be...  ...systems. The ideal candidate is passionate about cloud infrastructure, automation, and building reliable, production-grade environments... 
    Full time

    OPPO

    Palo Alto, CA
    9 hours ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for...  ...on' role — you will build the automation, observability, and tooling that allows... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    9 hours ago
  • $184k - $287.5k

     ...the effort in crafting agentic AI systems that...  ...critically important infrastructure. Be part of a dynamic...  ...standards in AI-enhanced automation!What you'll be doing:...  ...closely with GPU SW kernel engineers to pinpoint high-...  ...that are fast, reliable, and meaningfully better... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $200k - $322k

     ...Senior Staff Software Engineer to own the engineering...  ...strategic, AI infused automated resolution systems and...  ...:Design and implement agentic AI workflows using LLM...  ...Own the full stack from infrastructure to user facing tools....  ...years experience in SRE, Enterprise Support or... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Job Description Job Description Java SRE Engineer Onsite San Francisco Bay Area Infrastructure Engineer (2 Positions) We are looking for an...  ...platforms. This role is focused on infrastructure, reliability, and automation , with Java exposure as a supporting skill.... 

    Eitacies Inc

    Santa Clara, CA
    a month ago
  •  ...Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE Devops Engineer Fulltime Role for one of our direct client. Role: SRE Devops Engineer Location :- Santa Clara,CA (Remote... 
    Full time
    Local area
    Remote work

    Rootshell Enterprise Technologies

    Santa Clara, CA
    2 days ago
  • $101k - $161k

     ...awards, such as Best Engineering Team, Best Company for...  ...’re looking for Site Reliability Engineers to join our...  ...Service (CVaaS) global SRE team. SREs at Arista...  ...believe in building highly automated and self-sustaining...  ...stack, monitoring infrastructure, and much more. What... 

    Arista Networks

    Santa Clara, CA
    2 days ago
  •  ...Description Developer & Infrastructure Expert Role Type:...  ..., DevOps, SRE, and platform engineering. You will test AI-generated...  ...for accuracy and reliability. Work with AWS,...  ...Infrastructure Site Reliability...  ...workplace connectors, automation, or AI-powered infrastructure... 
    Remote job
    For contractors

    YO AI Labs

    San Jose, CA
    3 days ago
  • $174k - $252k

     ...through mechanisms like automation, and evolve...  ...changes that improve reliability and velocity....  ...Computer Science, Engineering, a related field,...  ...Science or Engineering.Site Reliability Engineering (SRE) is what you get...  ...by the Technical Infrastructure team to keep it running... 

    Google

    Sunnyvale, CA
    4 days ago
  • $262k - $364k

     ...AViD ecosystem have reliability and uptime...  ...performance.Build creative engineering solutions to operations and infrastructure problems, including AI-powered automation and agentic workflows for...  ...to the cross-SRE AI Ops program, driving...  ...a strategic way.Site Reliability... 

    Google

    Mountain View, CA
    4 days ago
  •  ...Job Title: Senior Site Reliability Engineer Kubernetes Platform Location: Remote...  ...~10+ years of experience in SRE, DevOps, or infrastructure engineering ~ Strong experience...  ...Improve system reliability through automation, thoughtful design, and continuous... 
    Full time
    Remote work

    SFE

    San Jose, CA
    1 day ago
  • $70 - $100 per hour

     ...Job Title: Cloud SRE Engineer - Mandarin Bilingual Position Type...  ...Cloud SRE Engineer to own the reliability, stability, and continuous...  ...— spanning compute infrastructure (CVM/VMs), networking, and...  ...tooling Develop scripts and automation tools to improve operational... 
    Hourly pay
    Contract work
    Temporary work
    Work experience placement

    IntelliPro Group Inc.

    Palo Alto, CA
    more than 2 months ago
  • $140k - $165k

     ...advanced electronic devices and IT infrastructure, enabling enhanced...  ...systems. Work on cutting-edge agentic AI — not just chatbots, but...  ...are seeking a hands-on AI Engineer to design, deploy, and maintain...  ...that drive real-world automation. You’ll be responsible for setting... 
    Full time

    SK Hynix Memory Solutions America Inc.

    San Jose, CA
    more than 2 months ago
  • $207k - $300k

     ...Software/Systems Engineers on projects for users...  ...services, build automation to prevent...  ...sustainable multi-site on-call rotations...  ...expertise in Site Reliability Engineering practices...  ...Engineering (SRE) combines software...  ...systems, building infrastructure and eliminating work... 

    Google

    San Jose, CA
    1 day ago
  • $170k - $210k

     ...quality, performance, and reliability matter. As part of...  ...will help define how agentic AI systems are...  .... This is a hands-on engineering role for someone who...  ...copilots that drive automation, accelerate workflows...  ...engineering, product, QA, infrastructure, and cross-functional... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $143k - $191k

     ...building the next generation of Agentic AI products that help...  ...We’re looking for a Software Engineer to design and build highly scalable...  ...intuitive, scalable, and reliable product features. This is a...  ...Webpack, NextJS. Data & Infrastructure: MySQL, Solr, Apache Airflow... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Eightfold

    Santa Clara, CA
    a month ago
  •  ...Job Title: Mid-Senior Site Reliability Engineer Kubernetes Platform Location...  ...8+ years of experience in SRE, DevOps, or platform...  ...fundamentals Experience with Infrastructure as Code (e.g., Terraform)...  ...Build and improve automation, tooling, and CI/CD workflows... 
    Full time

    SFE

    San Jose, CA
    1 day ago
  •  ...scripting) language for automation of build tasks (PowerShell...  ...deploy resources using Infrastructure-as-Code Develop...  ...balancer, DNS, etc. Engineering and coding skills applied...  ...Understand the concepts of Site Reliability Engineering (SRE) to maximize automation,... 
    Local area
    2 days per week

    Insight Global

    San Jose, CA
    4 days ago
  • $144k - $216k

     ...As a Senior Software Engineer on our Core AI team, you will be a key driver of FloQast's...  ...the AI products that power our accounting automation platform and enable our vision of an AI...  ...chatbots, document processing systems, and agentic workflows using Python and modern AI... 
    Full time

    Floqast

    San Jose, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Infrastructure & Agentic Automation. Be the first to apply!