Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

Wand AI

Build the Future Workforce

Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it's already operating at scale inside some of the world's largest organizations.

Wand built the world's first Agentic Labor Infrastructure enabling governments and global enterprises to create, manage, and scale digital workforces.

Our mission is to integrate agent ecosystems into the core of work and business, unlocking a generational leap in the global economy. We're building the infrastructure that lets humans and AI agents operate together safely, transparently, and at scale.

Join Wand in leading the Agentic Shift

Wand is building a high-performing global team who take full ownership of what they build. We lead by example, move fast, make data-aware decisions, and continuously push for more- always with a focus on delivering real value to customers.

You would be joining a world-class team that combines deep research expertise and real-world product execution, with experience spanning Deepmind, Google, Amazon, Miro, Elise AI, IBM and Accern.

Position Summary

We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices at scale. You will design and evolve resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI-driven products operate with high availability, performance, and security.

You will work across platform, product, data, and ML teams, helping us productionise models, absorb and standardise customer environments, strengthen Kubernetes-based architecture, and mature our CI/CD pipelines end-to-end.

You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology vision.

Responsibilities
  • Architect, deploy, and operate scalable, secure production environments (AWS preferred).
  • Lead reliability improvements across multiple engineering streams.
  • Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives.
  • Build and enforce strong Infrastructure-as-Code standards.
  • Define and operationalise SLIs, SLOs, and error budgets.
  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
  • Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights.
  • Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback.
  • Improve release safety, deployment frequency, and predictability of SLAs.
  • Lead incident response for complex cross-system failures and drive postmortems.
  • Reduce operational toil through automation and platform engineering improvements.
  • Design processes and tooling to absorb, standardise, and troubleshoot customer environments.
  • Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).
  • Ensure infrastructure aligns with enterprise-grade security and regulatory requirements.
  • Mentor engineers and raise the overall reliability bar across teams.
Key Requirements
  • Extensive hands-on experience in SRE or Production Engineering roles.
  • Demonstrated experience building or scaling SRE practices in high-growth or complex environments.
  • Deep expertise in AWS or Azure-based cloud infrastructure.
  • Strong experience with Kubernetes (including migration, scaling, and production hardening).
  • Advanced Infrastructure-as-Code experience (Terraform or equivalent).
  • End-to-end CI/CD pipeline design and optimisation experience.
  • Strong experience with observability tooling across distributed systems.
  • Experience troubleshooting complex multi-tenant or customer-hosted environments.
  • Experience supporting production data platforms and ML systems.
  • MLOps experience, including model deployment and monitoring.
  • Strong understanding of distributed systems, scalability, and fault tolerance.
  • Systems thinker who understands interactions across infrastructure, product, data, and ML.
  • Excellent communication skills and ability to work cross-functionally.
Preferred Experience
  • Experience in large-scale global B2B/B2C products.
  • Experience working with AI/ML systems, NLP, or LLM-based products.
  • Experience integrating product analytics and model performance metrics into operational monitoring.
  • Background in enterprise environments with strong security and compliance requirements.
  • Experience implementing regulatory controls within cloud infrastructure.
  • Experience scaling infrastructure during rapid growth phases.
  • Experience evaluating infrastructure tooling and vendors.
  • Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.
Personal Characteristics
  • Strong problem solver who anticipates failure modes.
  • High ownership mentality and accountability.
  • Comfortable working across streams and influencing without formal authority.
  • Learning-oriented with a drive for continuous improvement.
#J-18808-Ljbffr
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Palo Alto, CA vacancy
  • $165k - $280k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  • $262k - $364k

     ...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with...  ...capacity and performance.Build creative engineering solutions to operations and...  ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software... 
    Suggested

    Google

    Mountain View, CA
    3 days ago
  • $222k - $300.5k

     ...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps...  ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency... 
    Suggested
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    1 day ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    4 days ago
  •  ...Job Description Job Description Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes... 

    Amiri Recruiting

    Mountain View, CA
    3 days ago
  • $165k - $280k

     ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we're leveraging our experience in building rockets and spacecraft to deploy Starlink, the world's... 
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    2 days ago
  •  ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently... 
    Shift work

    AI Chopping Block

    Menlo Park, CA
    1 day ago
  • $150.2k - $283.5k

     ...components of the Android framework, enhancing the performance, reliability, and security of our IVI platform. The ideal candidate will...  ....Bachelor’s or Master’s degree in Computer Science, Software Engineering, or equivalent combination of relevant education and... 
    Immediate start
    Visa sponsorship
    Flexible hours

    Ford

    Palo Alto, CA
    4 days ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services—from designing fault... 
    Full time

    Zoox

    Foster, CA
    more than 2 months ago
  • $230k - $250k

     ...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change...  ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"... 
    Night shift

    Forward Networks

    Santa Clara, CA
    3 days ago
  • $101k - $161k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-...  ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-... 

    Arista Networks

    Santa Clara, CA
    1 day ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Full time

    Google

    Sunnyvale, CA
    14 hours ago
  • $168k - $270.25k

     ...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $170k - $200k

    We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,... 
    Full time
    Worldwide

    Fortinet

    Sunnyvale, CA
    14 hours ago
  • $152k - $241.5k

     ...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (...  ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,... 
    Full time

    Nvidia

    Santa Clara, CA
    14 hours ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 

    Google

    Sunnyvale, CA
    3 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Full time

    Nvidia

    Santa Clara, CA
    14 hours ago
  • $160k - $240k

     ...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit...  ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in... 
    Full time

    Fiserv

    Sunnyvale, CA
    1 day ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $145k - $165k

     ...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining... 
    Work at office
    Immediate start

    Bolt Graphics, Inc.

    Sunnyvale, CA
    4 days ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    1 day ago
  •  ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the... 
    Work at office

    Foxconn Industrial Internet - FII

    Sunnyvale, CA
    a month ago
  •  ...Position: Site Reliability Engineer-10+ Year exp required Location : Sunnyvale CA Job Summary We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep... 

    ReqRoute,Inc

    Sunnyvale, CA
    1 day ago
  •  ...of Huobi globe spanning infrastructure. •       Work with engineering teams to make sure new features and changes are deployed quickly...  .... •       Constantly improve our system performance and reliability through better tools, process and monitoring system. •... 
    Worldwide

    Cryptoware Technologies Inc

    Santa Clara, CA
    a month ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $262k - $364k

     ...recommendation models, novel model architectures, and optimize ML infrastructure to drive the growth of the Shorts ecosystem.Partner with engineering, product, data-science, and research teams to convert business goals into scalable technical solutions that grow the Shorts... 

    Google

    Mountain View, CA
    4 days ago
  • $150.2k - $283.5k

     ...world where every person is free to move and pursue their dreams.In this position... Ford Model E Platform Architecture Engineering is looking for a Staff Embedded Software Engineer to be part of the microcontroller which is designing the software stack from ground up to... 
    Work experience placement
    Immediate start
    Visa sponsorship
    Flexible hours

    Ford

    Palo Alto, CA
    3 days ago
  • $132.6k - $214.5k

     ..., you will collaborate closely with our engineering teams to develop innovative solutions that...  ...performance and health. As a Senior Staff SRE with the Cortex Observability team,...  ...operability of the product and ensure the reliability and availability of our services.... 
    Full time
    Work at office
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Santa Clara, CA
    1 day ago
  • $255.7k - $300k

    Lead a team of engineers to maintain service uptime while managing global on-call rotations...  ...improve operational practices to drive reliability, maintainability, and stakeholder alignment...  ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation... 
    Full time
    Work at office

    Google

    Sunnyvale, CA
    14 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!