Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Wand AI

Build the Future Workforce

Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it’s already operating at scale inside some of the world’s largest organizations.

Build the Future Workforce

Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it’s already operating at scale inside some of the world’s largest organizations. Wand built the world’s first Agentic Labor Infrastructure enabling governments and global enterprises to create, manage, and scale digital workforces. Our mission is to integrate agent ecosystems into the core of work and business, unlocking a generational leap in the global economy. We’re building the infrastructure that lets humans and AI agents operate together safely, transparently, and at scale.

Join Wand in leading the Agentic Shift

Wand is building a high-performing global team who take full ownership of what they build. We lead by example, move fast, make data-aware decisions, and continuously push for more‑always with a focus on delivering real value to customers. You would be joining a world-class team that combines deep research expertise and real-world product execution, with experience spanning Deepmind, Google, Amazon, Miro, Elise AI, IBM and Accern.

Position Summary

We are hiring for a hands‑on Head of SRE to establish, lead, and scale our Site Reliability Engineering function. This role combines strategic ownership with deep technical execution. You will be responsible for defining reliability standards, building operational processes, and ensuring production stability, while actively architecting infrastructure, improving automation, and embedding SRE best practices across the engineering organisation. There is significant scope to review, improve, and rebuild our systems, infrastructure and processes where necessary. You will be instrumental in designing, developing, and maintaining scalable backend systems, ensuring our AI products meet the highest standards. You will also become part of the product‑engineering leadership team, contribute to scaling the organization, and report directly to the CPTO.

Responsibilities
  • Own and lead all SRE‑related strategy, standards, and execution. Embed SRE culture and operational excellence across engineering teams.
  • Review the current infrastructure and operational model; redesign and rebuild where needed.
  • Architect, deploy, and maintain scalable, secure production environments.
  • Define and implement SLIs, SLOs, and uptime targets.
  • Establish robust monitoring, alerting, and observability practices.
  • Design and implement incident management, RCA and postmortem processes.
  • Build and manage sustainable on‑call frameworks and escalation models.
  • Automate the software delivery lifecycle to improve release predictability and safety.
  • Create reproducible environments and IaaC provisioning templates.
  • Improve system performance, availability, and reliability.
  • Support and productionise data platforms and ML workloads.
  • Partner closely with QA and Engineering leadership to improve release quality and stability.
  • Ensure infrastructure meets enterprise‑grade security and regulatory requirements.
  • Hire, manage, and mentor a team of SRE engineers.
Key Requirements
  • Proven hands‑on experience in Site Reliability Engineering, Production Engineering, or a similar role.
  • Strong hands‑on expertise in cloud infrastructure (AWS or Azure preferred), IaaC (Terraform) and Kubernetes.
  • Experience building or maturing SRE practices within an organisation.
  • Demonstrated ability to improve uptime, reliability, and operational processes.
  • Deep understanding of CI/CD, dev exp, infrastructure‑as‑code, and automation.
  • Experience designing on‑call processes and incident response frameworks.
  • Experience managing at least one team of SRE engineers.
  • Strong communication skills, with the ability to influence across teams.
  • Experience supporting data platforms and ML systems in production environments.
  • MLOps experience (model deployment, monitoring, retraining workflows).
Preferred Experience
  • Background in large‑scale global B2B/B2C products.
  • Background in enterprise environments with security and compliance requirements.
  • Expertise in ML, AI, LLMs.
  • Experience implementing regulatory controls within cloud infrastructure.
  • Experience evaluating and managing infrastructure vendors and tooling.
  • Experience scaling systems in high‑growth environments.
  • Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.
Personal Characteristics
  • Practical and hands‑on; willing to lead from the front.
  • Strong operational mindset with clear opinions on best practices.
  • Structured thinker who can build processes from ambiguity.
  • High ownership mentality and accountability.
  • Learning‑oriented with a continuous improvement mindset.
  • Excellent communication and interpersonal skills.
  • Continuous drive for improvement and innovation.
#J-18808-Ljbffr
Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Palo Alto, CA vacancy
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    1 day ago
  • $127.1k - $226k

     ...A leading technology company is seeking a Principal Kubernetes Software Engineer in Palo Alto, CA. The role involves developing and delivering features for Kubernetes and CNCF projects while collaborating with a global community. Candidates should have extensive experience... 
    Suggested

    Broadcom Corporation

    Palo Alto, CA
    1 day ago
  •  ...Site Reliability Engineer, Data Platform - USDS Responsibilities Engage in and improve the whole lifecycle of service, from inception and design, through to deployment, operation and refinement. Ensure reliable, fault-tolerant, efficiently scalable and cost-effective data... 
    Suggested

    Tik Tok

    Mountain View, CA
    22 hours ago
  •  ...Site Reliability Engineer There are NO limits to your career: come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software... 
    Suggested
    Immediate start
    Remote work
    Worldwide

    OutSystems

    Menlo Park, CA
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 
    Suggested
    Work at office

    Chase

    Palo Alto, CA
    4 days ago
  • $27.12 - $51.93 per hour

     ...About the Hiring Team What the Role Entails Role Summary We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business... 
    Hourly pay
    Full time
    Internship

    Tencent

    Palo Alto, CA
    3 hours ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation... 
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    20 hours ago
  • $222k - $300.5k

     ...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps...  ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency... 
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    3 days ago
  • $262k - $364k

     ...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with...  ...capacity and performance.Build creative engineering solutions to operations and...  ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software... 

    Google

    Mountain View, CA
    22 hours ago
  •  ...Overview We are seeking a highly motivated Systems Reliability Engineer (SRE) to lead the design and implementation of operational excellence...  ...company supporting sensitive and cleared workforces. The Site Reliability Engineer (SRE) - SecOps will embrace our commitment... 
    For contractors
    Work at office
    Flexible hours

    Arkenstone Defense

    Menlo Park, CA
    12 days ago
  •  ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently... 
    Shift work

    Hippocratic AI Inc.

    Menlo Park, CA
    3 days ago
  •  ...exceptional professionals for this role. JOB DESCRIPTION Elevate your engineering prowess to unprecedented levels by joining a team of...  ...and position yourself among the top echelon in site reliability. As a Senior Lead Site Reliability Engineer at JPMorgan Chase... 

    J.P. Morgan

    Palo Alto, CA
    2 days ago
  • $100k - $200k

    OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Full time

    OPPO

    Palo Alto, CA
    1 day ago
  • $200k - $260k

     ...Site Reliability Engineering Lead Glean is seeking a Site Reliability Engineering Lead to foster a culture of engineering excellence, drive technical strategy, and develop a high-performing, collaborative team. Your role is pivotal in ensuring our services meet stringent... 
    Work at office
    Home office

    Glean - Mountain View, CA, US

    Mountain View, CA
    3 days ago
  • $217.57k - $260k

     ...job description explicitly states otherwise, all roles are on-site five days per week at one of our offices in McLean, VA;...  ...which can be found here. Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale infrastructure... 
    Full time
    Temporary work
    Work at office
    Remote work
    Flexible hours
    Shift work

    ID.me

    Mountain View, CA
    4 days ago
  •  ...Lead Site Reliability Engineer Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within... 

    Chase

    Palo Alto, CA
    10 hours ago
  • $207k - $300k

    Lead a team of Software/Systems Engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...Science or Engineering.1 year of people management experience. Site Reliability Engineering (SRE) combines software and systems engineering... 

    Google

    Mountain View, CA
    4 days ago
  • $251k - $310k

     ...billions in simulation across 15+ U.S. states. Software Reliability Engineering : Waymo’s software reliability engineers (SRE) are responsible...  ...across functional groups, translating business needs into site reliability initiatives. Strong background in managing technical... 
    Full time
    Immediate start
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    22 hours ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    2 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Overview Title: Site Reliability Engineer SRE – ML platform Location: Austin, TX or Sunnyvale, CA Employment type: Full-time • Seniority: Mid-Senior level • ONLY W2 Responsibilities Continuous Deployment using GitHub Actions, Flux, Kustomize Design and implement cloud... 
    Full time

    Saransh

    Sunnyvale, CA
    22 hours ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    3 days ago
  •  ...Google is seeking a Senior Engineering Manager for Collaboration SRE to lead a multi-site engineering organization across Sunnyvale and Zurich. You will own the...  ..., and drive high-impact projects that improve reliability, performance, and scalability of #J-18808-Ljbffr

    Socket

    Sunnyvale, CA
    1 day ago
  • $150k - $195k

     ...customers worldwide. Our team is growing, and we are looking for engineers with passion for automation. You will help support the...  ...alongside engineering/operations teams to improve the scalability and reliability of internal processes. Participate in an on‑call rotation.... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    4 days ago
  • $145k - $175k

     ...straightforward communication and clinical domain expertise, Commence cuts straight to better care. Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data... 
    Full time
    Remote work

    GrabJobs

    Santa Clara, CA
    2 days ago
  • $276.1k - $311.4k

     ...Vehicle Software SRE team from the ground up — defining its charter, hiring its founding engineers, establishing the operating model, and creating the technical strategy that makes reliability a first-class property of the software running on our vehicles. You'll work in a... 
    Permanent employment
    Full time
    Work at office
    Work from home

    Lindus Health

    Sunnyvale, CA
    3 days ago
  • $145k - $165k

     ...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Work at office

    Bolt Graphics

    Sunnyvale, CA
    2 days ago
  • $160k - $240k

     ...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay...  ...come make a difference at Fiserv. Job Title Senior Site Reliability Engineer What does a successful Site Reliability Engineer do at... 

    Fiserv

    Sunnyvale, CA
    9 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!