Staff Site Reliability Engineer
$120.6k - $150.9kWEX
About the RoleWe are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX’s platform reliability and operational excellence.This is a particularly exciting time to be part of the SRE function at WEX. Our diverse product ecosystem supports a wide array of customer businesses and generates rich, complex telemetry across applications, infrastructure, and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and customer success.As a Staff SRE, you will play a pivotal role in shaping the reliability engineering strategy at WEX. You’ll architect and lead efforts that improve availability, performance, and efficiency at scale, driving initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization. You’ll be hands-on in building foundational tooling and frameworks while also acting as a multiplier, mentoring engineers, aligning cross-functional teams, and influencing platform decisions with a strong reliability lens.You’ll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing security and operational guardrails so intelligent automation is trustworthy, measurable, and scalable. Our team embraces agile development, a strong product mindset, and modern engineering practices, including AI-assisted operations and intelligent automation.You’ll take on some of the most complex, high-impact challenges at WEX, supported by a team of highly skilled engineers and technical leaders invested in your success and growth.If you’re a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI-enabled operations, this is a fantastic opportunity for you.What You’ll DoArchitect and oversee the implementation of mission-critical systems with a focus on availability, scalability, and operational excellence.Define and enforce SRE best practices and operational standards across engineering and platform teams.Lead cross-functional initiatives to enhance system reliability, performance, and efficiency at scale.Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk.Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks.Design self-healing and auto-recovery mechanisms that reduce manual intervention during failures.Drive cloud cost optimization and budgeting initiatives without compromising reliability.Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL.Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows.Apply AI security and governance controls, including least-privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries.Lead AI-enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes.Mentor engineers on production-grade agentic solutions and help embed AI into day-to-day reliability practices.What You’ll Bring8+ years of experience with a focus on large-scale system reliability.Expertise in system architecture, cloud platforms, and automation frameworks.Deep knowledge of Kubernetes, service meshes, and distributed tracing.Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.).Knowledge of containerization and orchestration (Docker, Kubernetes).Experience designing high-availability, fault-tolerant architectures.Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures.Excellent incident command and crisis management skills.Hands-on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD).Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human-in-the-loop patterns.Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions.Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains.Nice to HaveExperience with multi-region and multi-cloud deployments.Deep expertise in scalable microservices and event-driven architectures.Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus).Leadership in driving large-scale SRE transformations.Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices.Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows.Ability to influence engineering culture and process improvements, including adoption of AI-assisted operations under change control, safety, and audit requirements.The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non-sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well-being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section.Pay Range: $120,600.00 - $150,900.00SummaryLocation: San Francisco, CA; Chicago, IL; Dallas, TXType: Full time
$128.6k - $184.9k
...global cloud platform. As a team of six engineers distributed across the US, Canada, and the... ...with a strong focus on automation, reliability, and operational excellence. We are one... ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure...SuggestedPermanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours$174k - $252k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical... ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you...Suggested$147k - $210k
...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style... ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations...Suggested$138.4k - $173k
...infrastructure as well as help improve the reliability, quality of services and overall... ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability... ...about our locations by visiting our site.Compensation & BenefitsThe base salary that...SuggestedFull timeFlexible hours$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SuggestedFull timeWork at officeLocal areaRemote workWork from home- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...Full time
- Qualifications: 8+ years of Software Engineering experience, or equivalent... ...and maintain scalable and reliable infrastructure on Google... ...the client, IT management and staff, and other groups in Information... ...resources Willingness to work on-site at stated location in the job...Contract workFor contractorsWork experience placement
- ...Senior Site Reliability Engineer (Permanent Role) Cleveland, OH, Pittsburgh, PA, or Dallas, TX Your future duties and responsibilities . Monitoring distribution systems and notifying them of any potential issues. . Assisting with troubleshooting on call....Permanent employmentFull timeLocal areaFlexible hoursShift workWeekend work
$48 per hour
...Site Reliability Engineer Trident Consulting is seeking a Site Reliability Engineer for one of our clients in Richardson, Texas or Scottsdale, Arizona. Job Title: Site Reliability Engineer Location: Richardson, Texas or Scottsdale, Arizona (Onsite) Length of Assignment...Contract work- ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market... ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications....
- ...Senior Site Reliability Engineer This role will require someone onsite at our client office in Cleveland, OH, Pittsburgh, PA, or Dallas, TX. Systemone is looking for a Site Reliability Engineer who will work within the Production support and Performance Management...Work at officeFlexible hoursShift workWeekend work
- ...ensure applications are highly available, reliable, and performant at a global scale.... ...Bachelor of Computer Science or related Engineering field required. Master's Degree preferred... ...Minimum of 1 year of lead experience of site reliability engineering team required....Contract workWork at office
- ...Site Reliability Engineer (SRE) The successful applicant may be performing work in FedRAMP High or IL-5 environments, and therefore, must be a U.S. Person (i.e. U.S. citizen, U.S. national, lawful permanent resident, asylee, or refugee). This position may also perform...Permanent employmentWorldwideShift work
- ...and continuously improving the platforms that power TI's digital integration, automation and DevOps capabilities. As an IT Site Reliability Engineer within the Enterprise Platforms team, you will serve as the primary technical platform owner for TI's Apigee Edge private...Local area
- ...Site Reliability Engineer TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital markets. With SEC approval and $275MM in funding, we are currently hiring a Site Reliability Engineer to help with a greenfield...Currently hiring
- ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology... ...AI-driven solutions that accelerate development and improve reliability. Your work will directly influence how GM Financial leverages...Full timeH1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week3 days per week
- ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability...Work at office
- ...Senior Site Reliability Engineer Our client is seeking a Senior Site Reliability Engineer for a month 6-month contract in Irving, TX. Will be working on an onsite schedule. Contract Duration: 6 Months Required Skills & Experience ~ Bachelors/4 Year Degree ~5+...Full timeContract workTemporary workWork experience placementFlexible hours
- ...Job Title: Site Reliability Engineer (SRE) Work Location: RichardsonTX 75082 **FULL ONSITE WORK** Interview Mode- In-person Interview Contract duration: 9 Months Job Details: Must Have Skills Python, Kubernetes, Terraform, GitLab CI/CD Nice to...Contract work
- Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering... ...best practices, and actively mentor and develop senior and staff-level engineers.Technology Evaluation & Adoption: Stay at...
$72.1k - $158.62k
...person, one family and one community at a time. Position Summary We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence. The...Hourly payFull timeTemporary workLocal area- Compliance EngineeringWe are Compliance Engineering, a global team of more than 500 engineers and scientists who work on the most complex... ...systems by pushing for changes that improve capacity and reliability.Practicing sustainable incident management in a blameless postmortem...
$113.1k - $232.3k
Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity...Work at officeLocal areaVisa sponsorshipFlexible hours3 days per week- ...a company that values diversity, integrity, and growth. Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the...
- Compliance Engineering, Site Reliability Engineering, Vice President, Dallas location_on Dallas, TX, United States We are Compliance Engineering, a global team of more than 300 engineers and scientists who work on the most complex, mission-critical problems. We build and...Full timeTemporary workWork at office
$207k - $300k
...projects.Collaborate closely with hardware, software, and system engineering teams to drive pre-silicon Firmware (FW)/Software (SW) co-... ...AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers...Worldwide- ...Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps‑driven environment...Long term contractContract workInternshipRemote work
$40 per hour
...A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work- ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,... ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
- ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- staff engineer Dallas, TX
- assistant engineer Dallas, TX
- engineering aide Dallas, TX
- senior staff engineer Dallas, TX
- senior staff systems engineer Dallas, TX
- assistant electrical engineer Dallas, TX
- assistant engineering manager Dallas, TX
- software engineer staff Dallas, TX
- technology administrator Dallas, TX
- assistant mechanical engineer Dallas, TX


