Staff Site Reliability Engineer
Wand AI
Build the Future Workforce
Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it's already operating at scale inside some of the world's largest organizations.
Wand built the world's first Agentic Labor Infrastructure enabling governments and global enterprises to create, manage, and scale digital workforces.
Our mission is to integrate agent ecosystems into the core of work and business, unlocking a generational leap in the global economy. We're building the infrastructure that lets humans and AI agents operate together safely, transparently, and at scale.
Join Wand in leading the Agentic Shift
Wand is building a high-performing global team who take full ownership of what they build. We lead by example, move fast, make data-aware decisions, and continuously push for more- always with a focus on delivering real value to customers.
You would be joining a world-class team that combines deep research expertise and real-world product execution, with experience spanning Deepmind, Google, Amazon, Miro, Elise AI, IBM and Accern.
Position Summary
We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE practices at scale. You will design and evolve resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI-driven products operate with high availability, performance, and security.
You will work across platform, product, data, and ML teams, helping us productionise models, absorb and standardise customer environments, strengthen Kubernetes-based architecture, and mature our CI/CD pipelines end-to-end.
You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology vision.
Responsibilities
- Architect, deploy, and operate scalable, secure production environments (AWS preferred).
- Lead reliability improvements across multiple engineering streams.
- Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives.
- Build and enforce strong Infrastructure-as-Code standards.
- Define and operationalise SLIs, SLOs, and error budgets.
- Strengthen observability across applications, infrastructure, data pipelines, and ML systems.
- Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights.
- Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback.
- Improve release safety, deployment frequency, and predictability of SLAs.
- Lead incident response for complex cross-system failures and drive postmortems.
- Reduce operational toil through automation and platform engineering improvements.
- Design processes and tooling to absorb, standardise, and troubleshoot customer environments.
- Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).
- Ensure infrastructure aligns with enterprise-grade security and regulatory requirements.
- Mentor engineers and raise the overall reliability bar across teams.
Key Requirements
- Extensive hands-on experience in SRE or Production Engineering roles.
- Demonstrated experience building or scaling SRE practices in high-growth or complex environments.
- Deep expertise in AWS or Azure-based cloud infrastructure.
- Strong experience with Kubernetes (including migration, scaling, and production hardening).
- Advanced Infrastructure-as-Code experience (Terraform or equivalent).
- End-to-end CI/CD pipeline design and optimisation experience.
- Strong experience with observability tooling across distributed systems.
- Experience troubleshooting complex multi-tenant or customer-hosted environments.
- Experience supporting production data platforms and ML systems.
- MLOps experience, including model deployment and monitoring.
- Strong understanding of distributed systems, scalability, and fault tolerance.
- Systems thinker who understands interactions across infrastructure, product, data, and ML.
- Excellent communication skills and ability to work cross-functionally.
Preferred Experience
- Experience in large-scale global B2B/B2C products.
- Experience working with AI/ML systems, NLP, or LLM-based products.
- Experience integrating product analytics and model performance metrics into operational monitoring.
- Background in enterprise environments with strong security and compliance requirements.
- Experience implementing regulatory controls within cloud infrastructure.
- Experience scaling infrastructure during rapid growth phases.
- Experience evaluating infrastructure tooling and vendors.
- Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.
Personal Characteristics
- Strong problem solver who anticipates failure modes.
- High ownership mentality and accountability.
- Comfortable working across streams and influencing without formal authority.
- Learning-oriented with a drive for continuous improvement.
$165k - $280k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most...SuggestedPermanent employmentTemporary workWorldwideWeekend work$165k - $265k
...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield constellation...SuggestedPermanent employmentTemporary workImmediate startWeekend work$262k - $364k
...services within the AViD ecosystem have reliability and uptime appropriate to users' needs with... ...capacity and performance.Build creative engineering solutions to operations and... ...changing circumstances in a strategic way.Site Reliability Engineering (SRE) combines software...Suggested$222k - $300.5k
...possible.Job OverviewAbout the TeamIntuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps... ...hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency...SuggestedWorldwideShift work$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...SuggestedFull time- ...Job Description Job Description Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You’ll Do (Day-to-Day) Own and manage our cloud infrastructure (GCP or AWS, on-prem). Build, maintain, and optimize Kubernetes...
$165k - $280k
...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we're leveraging our experience in building rockets and spacecraft to deploy Starlink, the world's...Temporary workWorldwideWeekend work- ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently...Shift work
$150.2k - $283.5k
...components of the Android framework, enhancing the performance, reliability, and security of our IVI platform. The ideal candidate will... ....Bachelor’s or Master’s degree in Computer Science, Software Engineering, or equivalent combination of relevant education and...Immediate startVisa sponsorshipFlexible hours$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services—from designing fault...Full time$230k - $250k
...network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change... ...how things have always been done.Forward is looking for a Site Reliability EngineerAbout the Role This is not a "keep the lights on"...Night shift$101k - $161k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-... ...we do.Job DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s CloudVision-as-a-...$255.7k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation...Full time$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...Full time$170k - $200k
We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You will be responsible for building, maintaining, and troubleshooting cloud service/cluster, infrastructure, and monitoring systems to ensure high availability, performance,...Full timeWorldwide$152k - $241.5k
...infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (... ...languages such as Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through design reviews,...Full time$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...Full time$160k - $240k
...one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit... ...come make a difference at Fiserv.Job TitleSenior Site Reliability EngineerWhat does a successful Site Reliability Engineer do at Fiserv?You will join our global team in...Full time$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing, building, and operating large-scale production systems with exceptional efficiency, resilience, and availability. It combines software and systems engineering practices with...Full time$145k - $165k
...: Selflessly collaborate towards our shared purpose. About the role Bolt Graphics is seeking a highly experienced Site Reliability Engineer (SRE) to design, build, and operate highly reliable developer and production systems. This role is mission-critical to maintaining...Work at officeImmediate start$180k - $230k
...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...Work at officeLocal areaImmediate startRemote work3 days per week- ...design by customizing MES tool per business needs Education Requirements, Ideal Experience: Associate’s degree in Industrial Engineering or IT related field Minimum of 0-3 years’ relevant experience Experience in C#, Delphi desired Knowledge of the...Work at office
- ...Position: Site Reliability Engineer-10+ Year exp required Location : Sunnyvale CA Job Summary We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep...
- ...of Huobi globe spanning infrastructure. • Work with engineering teams to make sure new features and changes are deployed quickly... .... • Constantly improve our system performance and reliability through better tools, process and monitoring system. •...Worldwide
$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized...Full time$262k - $364k
...recommendation models, novel model architectures, and optimize ML infrastructure to drive the growth of the Shorts ecosystem.Partner with engineering, product, data-science, and research teams to convert business goals into scalable technical solutions that grow the Shorts...$150.2k - $283.5k
...world where every person is free to move and pursue their dreams.In this position... Ford Model E Platform Architecture Engineering is looking for a Staff Embedded Software Engineer to be part of the microcontroller which is designing the software stack from ground up to...Work experience placementImmediate startVisa sponsorshipFlexible hours$132.6k - $214.5k
..., you will collaborate closely with our engineering teams to develop innovative solutions that... ...performance and health. As a Senior Staff SRE with the Cortex Observability team,... ...operability of the product and ensure the reliability and availability of our services....Full timeWork at officeVisa sponsorshipWork visa$255.7k - $300k
Lead a team of engineers to maintain service uptime while managing global on-call rotations... ...improve operational practices to drive reliability, maintainability, and stakeholder alignment... ...or in a Manager, Software Engineer, Site Reliability Engineering-related occupation...Full timeWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- senior staff engineer Palo Alto, CA
- senior staff systems engineer Palo Alto, CA
- engineering aide Palo Alto, CA
- software engineer staff Palo Alto, CA
- assistant engineer Palo Alto, CA
- technology administrator Palo Alto, CA
- staff engineer Palo Alto, CA
- site reliability engineer sre Palo Alto, CA
- site reliability engineer Palo Alto, CA
- IT site lead Palo Alto, CA



