Staff Site Reliability Engineer
$120.6k - $150.9kWex Health
Staff Site Reliability Engineer (SRE)
We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX's platform reliability and operational excellence.
This is a particularly exciting time to be part of the SRE function at WEX. Our diverse product ecosystem supports a wide array of customer businesses and generates rich, complex telemetry across applications, infrastructure, and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and customer success.
As a Staff SRE, you will play a pivotal role in shaping the reliability engineering strategy at WEX. You'll architect and lead efforts that improve availability, performance, and efficiency at scale, driving initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization. You'll be hands-on in building foundational tooling and frameworks while also acting as a multiplier, mentoring engineers, aligning cross-functional teams, and influencing platform decisions with a strong reliability lens.
You'll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing security and operational guardrails so intelligent automation is trustworthy, measurable, and scalable. Our team embraces agile development, a strong product mindset, and modern engineering practices, including AI-assisted operations and intelligent automation.
You'll take on some of the most complex, high-impact challenges at WEX, supported by a team of highly skilled engineers and technical leaders invested in your success and growth.
If you're a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI-enabled operations, this is a fantastic opportunity for you.
What You'll Do
- Architect and oversee the implementation of mission-critical systems with a focus on availability, scalability, and operational excellence.
- Define and enforce SRE best practices and operational standards across engineering and platform teams.
- Lead cross-functional initiatives to enhance system reliability, performance, and efficiency at scale.
- Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk.
- Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks.
- Design self-healing and auto-recovery mechanisms that reduce manual intervention during failures.
- Drive cloud cost optimization and budgeting initiatives without compromising reliability.
- Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL.
- Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows.
- Apply AI security and governance controls, including least-privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries.
- Lead AI-enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes.
- Mentor engineers on production-grade agentic solutions and help embed AI into day-to-day reliability practices.
What You'll Bring
- 8+ years of experience with a focus on large-scale system reliability.
- Expertise in system architecture, cloud platforms, and automation frameworks.
- Deep knowledge of Kubernetes, service meshes, and distributed tracing.
- Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.).
- Knowledge of containerization and orchestration (Docker, Kubernetes).
- Experience designing high-availability, fault-tolerant architectures.
- Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures.
- Excellent incident command and crisis management skills.
- Hands-on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD).
- Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human-in-the-loop patterns.
- Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions.
- Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains.
Nice to Have
- Experience with multi-region and multi-cloud deployments.
- Deep expertise in scalable microservices and event-driven architectures.
- Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus).
- Leadership in driving large-scale SRE transformations.
- Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices.
- Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows.
- Ability to influence engineering culture and process improvements, including adoption of AI-assisted operations under change control, safety, and audit requirements.
The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non-sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well-being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section.
Pay Range: $120,600.00 - $150,900.00
- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems... ...make safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...SuggestedPermanent employmentWork experience placementWork at officeLocal area
$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SuggestedFull timeFor contractors$114.3k - $235.32k
...verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native platform built on AWS, Kubernetes/EKS, and ArgoCD-driven...SuggestedWork at officeLocal areaRelocationRelocation package$113.4k - $162k
...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at...SuggestedTemporary work$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...SuggestedFlexible hours$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...Permanent employmentLocal areaWorldwideFlexible hours$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,...Full timeWorldwideWeekend work- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...
$152.5k - $205k
...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical...Flexible hours- ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ..., working together to build scalable, reliable, and secure products that empower businesses... ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely...Temporary workLocal areaWorldwide
$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable...Local areaWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week- We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering...
$220k - $235k
...We are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...Full timeContract workWork at office$160k - $200k
Job Purpose:BTIG seeks a DevOps/Site Reliability Engineer to join our technology team. This role is central to improving developer velocity by handling production application support escalations, managing and evolving our infrastructure stack, and providing operational...Full time$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$174k - $239k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability...Work experience placementLocal areaWorldwideFlexible hours$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$195k - $257.5k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and operate the infrastructure that powers our blockchain platform at...Flexible hours$204k - $306k
...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from... ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$194k - $267k
...career-defining work. We're all in on this mission. If you are too, let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud...Local areaWorldwideFlexible hours$165k - $241.4k
...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale,...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'...
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...Local areaRemote workWorldwideFlexible hours$175k - $250k
...00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance of... ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build...Full timeRemote workRelocationRelocation package$260k - $300k
...Software Agents We are the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding... ...than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that...- The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating the core systems that power agentic AI at scale. Your mission: keep our ultra-low...
$157k - $239k
San Francisco, CA / Golden, COInfrastructure - Cloud Infrastructure /Full time /On-siteWanna join the adventure?As a Site Reliability Engineer with strong networking skills in our Cloud Infrastructure (SRE) team, you help the team own the networks that keep Loft running...Full timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- staff engineer San Francisco, CA
- assistant engineer San Francisco, CA
- research assistant engineering San Francisco, CA
- staff design engineer San Francisco, CA
- staff security engineer San Francisco, CA
- engineering aide San Francisco, CA
- senior staff engineer San Francisco, CA
- senior staff systems engineer San Francisco, CA
- assistant chief engineer San Francisco, CA
- assistant engineering manager San Francisco, CA

