Staff Site Reliability Engineer
$120.6k - $150.9kWEX
About the Role
We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX’s platform reliability and operational excellence.
This is a particularly exciting time to be part of the SRE function at WEX. Our diverse product ecosystem supports a wide array of customer businesses and generates rich, complex telemetry across applications, infrastructure, and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and customer success.
As a Staff SRE, you will play a pivotal role in shaping the reliability engineering strategy at WEX. You’ll architect and lead efforts that improve availability, performance, and efficiency at scale, driving initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization. You’ll be hands-on in building foundational tooling and frameworks while also acting as a multiplier, mentoring engineers, aligning cross-functional teams, and influencing platform decisions with a strong reliability lens.
You’ll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing security and operational guardrails so intelligent automation is trustworthy, measurable, and scalable. Our team embraces agile development, a strong product mindset, and modern engineering practices, including AI-assisted operations and intelligent automation.
You’ll take on some of the most complex, high-impact challenges at WEX, supported by a team of highly skilled engineers and technical leaders invested in your success and growth.
If you’re a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI-enabled operations, this is a fantastic opportunity for you.
What You’ll Do
Architect and oversee the implementation of mission-critical systems with a focus on availability, scalability, and operational excellence.
Define and enforce SRE best practices and operational standards across engineering and platform teams.
Lead cross-functional initiatives to enhance system reliability, performance, and efficiency at scale.
Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk.
Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks.
Design self-healing and auto-recovery mechanisms that reduce manual intervention during failures.
Drive cloud cost optimization and budgeting initiatives without compromising reliability.
Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL.
Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows.
Apply AI security and governance controls, including least-privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries.
Lead AI-enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes.
Mentor engineers on production-grade agentic solutions and help embed AI into day-to-day reliability practices.
What You’ll Bring
8+ years of experience with a focus on large-scale system reliability.
Expertise in system architecture, cloud platforms, and automation frameworks.
Deep knowledge of Kubernetes, service meshes, and distributed tracing.
Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.).
Knowledge of containerization and orchestration (Docker, Kubernetes).
Experience designing high-availability, fault-tolerant architectures.
Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures.
Excellent incident command and crisis management skills.
Hands-on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD).
Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human-in-the-loop patterns.
Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions.
Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains.
Nice to Have
Experience with multi-region and multi-cloud deployments.
Deep expertise in scalable microservices and event-driven architectures.
Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus).
Leadership in driving large-scale SRE transformations.
Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices.
Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows.
Ability to influence engineering culture and process improvements, including adoption of AI-assisted operations under change control, safety, and audit requirements.
$100.7k - $167.8k
Job SummaryThe Site Reliability Engineer III is a pivotal architect of stability for CME Clearing & Risk. You will engineer secure, scalable, and reliable technology solutions that safeguard the global marketplace. By bridging the gap between development and operations,...SuggestedFull timeWorldwide- ...to physicians, providing critical information about the right treatments for the right patients, at the right time.The Site Reliability Engineering team works with all departments and business units to provide dependable cloud infrastructure solutions, along with support...SuggestedFull time
$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...Suggested
- Qualifications: 8+ years of Software Engineering experience, or equivalent... ...and maintain scalable and reliable infrastructure on Google... ...the client, IT management and staff, and other groups in Information... ...resources Willingness to work on-site at stated location in the job...SuggestedContract workFor contractorsWork experience placement
$130k - $225k
...expectations, integrity, innovation and a willingness to challenge consensus.The Algorithmic Trading Team is looking for a Site Reliability Engineer for our Chicago office. The SRE team is critical to the success of our trading - ensuring that our production trading...Temporary workWork at officeFlexible hours- ...IL, United StatesIndustry: Trading FirmPosted: 2026-08-17Contact: Mike LaTulipEmail: ****@*****.*** Title: Site Reliability Engineer (Infrastructure & Systems)Location: Chicago, IL (Greater Metro Area)About the OpportunityJoin a premier financial technology...Local area
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours$130k - $180k
...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours- Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will use technology to solve...
$130k - $150k
...Cloud SolutionsInformation SecurityInformation Technology staff are based in the Boston, Chicago, London, Munich, New York,... ...technologies is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are...Work at officeWork from home3 days per week$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours$108.08k - $172.5k
Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts and applications to maintain services. Monitor metrics, system health and analyze reports. Provide on-call rotation support for production...Full timeRemote workWorldwide- ...consumers and companies, alikeKlover’s engineering team powers one of the fastest-growing... ...production-grade systems that prioritize reliability, security, and performance, and that... ...right candidateAbout the RoleAs a Senior/Staff Site Reliability Engineer, you will play a critical...Work at officeImmediate startRemote work
$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology team, you will solve complex and broad business problems with simple...
$132.1k - $220.1k
Staff Site Reliability Engineer (SRE) - Platform EngineeringNote: This position follows a hybrid work model, requiring 2 days per week on-site at our corporate office 20 S Wacker Dr, Chicago, IL 60606The first preference for this role is given to local candidates in the...Full timeWork at officeLocal areaWorldwide2 days per week$174k - $267k
...defining work. We're all in on this mission. If you are too, let's talk. The Team We are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud...Full timeLocal areaWorldwideFlexible hours$160k - $210k
...world. What you'll do:Join our Platform Engineering team, where you'll ensure the availability... ...and mentoring engineers across reliability initiativesAnalyze, troubleshoot, and remediate... ...need:8+ years of experience in DevOps, Site Reliability Engineering, or Platform Engineering...Work at officeWorldwideMonday to FridayFlexible hours$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$132.1k - $220.1k
...We're looking for a Staff Site Reliability Engineer to join our team, focusing on the core systems that power global financial markets. This isn't just about keeping the lights on; it's about pioneering the future of financial technology. As a member of our Clearing department...Full timeWork at officeWorldwide2 days per week$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$112.5k - $187.5k
...TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering... ...most complex and consequential work on the platform.As a Staff Site Reliability Engineer at TransUnion, you will serve as...Full timeTemporary workWork experience placementWork at officeFlexible hours2 days per week$118.3k - $219.8k
Are you excited to lead Site Reliability Engineering teams that keep mission-critical, 24/7 services running reliably and securely?Do you enjoy building... ...leading SRE Team/s, including rotating staff across products / platforms, mentoring and development, objective...Full timeLocal area$180k - $200k
...Company Name: tastytrade Role: Senior Site Reliability Engineer Location: Chicago, IL (Hybrid, 3 days/week in office) Role Summary Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options...Work at office3 days per week- ...with the investigation, Always validate if the team is following the SOPs or the process defined for an alerts/ issue. Contacting all the external vendors if in case their integrations fail, Measure the front-end metrics for the site with various tools available...
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...Local area
- ...As a Senior DevOps / SRE Engineer on contract, you will be embedded with the Central Technology AI enablement team, working alongside engineers from Direct, PitchBook, Retirement, and other business units. Your initial focus will be on the SRE and hosting side of our...Contract workImmediate start
- ...Senior Site Reliability Engineer We are looking for a Senior Reliability Engineer to join our Platform team. In this position, you will be responsible for maintaining, designing, implementing and upgrading our cloud infrastructure to support our microservices platforms...Temporary workFlexible hours
- ...Site Reliability Engineer II Play a key role in ensuring system reliability at one of the world's most iconic and largest financial institutions. As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- staff engineer Chicago, IL
- assistant engineer Chicago, IL
- engineering aide Chicago, IL
- assistant civil engineer Chicago, IL
- senior staff engineer Chicago, IL
- senior staff systems engineer Chicago, IL
- software engineer staff Chicago, IL
- technology administrator Chicago, IL
- project engineer assistant project manager Chicago, IL
- site reliability engineer remote Chicago, IL

