Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

$120.6k - $150.9k

WEX

About the RoleWe are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX’s platform reliability and operational excellence.This is a particularly exciting time to be part of the SRE function at WEX. Our diverse product ecosystem supports a wide array of customer businesses and generates rich, complex telemetry across applications, infrastructure, and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and customer success.As a Staff SRE, you will play a pivotal role in shaping the reliability engineering strategy at WEX. You’ll architect and lead efforts that improve availability, performance, and efficiency at scale, driving initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization. You’ll be hands-on in building foundational tooling and frameworks while also acting as a multiplier, mentoring engineers, aligning cross-functional teams, and influencing platform decisions with a strong reliability lens.You’ll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing security and operational guardrails so intelligent automation is trustworthy, measurable, and scalable. Our team embraces agile development, a strong product mindset, and modern engineering practices, including AI-assisted operations and intelligent automation.You’ll take on some of the most complex, high-impact challenges at WEX, supported by a team of highly skilled engineers and technical leaders invested in your success and growth.If you’re a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI-enabled operations, this is a fantastic opportunity for you.What You’ll DoArchitect and oversee the implementation of mission-critical systems with a focus on availability, scalability, and operational excellence.Define and enforce SRE best practices and operational standards across engineering and platform teams.Lead cross-functional initiatives to enhance system reliability, performance, and efficiency at scale.Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk.Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks.Design self-healing and auto-recovery mechanisms that reduce manual intervention during failures.Drive cloud cost optimization and budgeting initiatives without compromising reliability.Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL.Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows.Apply AI security and governance controls, including least-privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries.Lead AI-enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes.Mentor engineers on production-grade agentic solutions and help embed AI into day-to-day reliability practices.What You’ll Bring8+ years of experience with a focus on large-scale system reliability.Expertise in system architecture, cloud platforms, and automation frameworks.Deep knowledge of Kubernetes, service meshes, and distributed tracing.Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.).Knowledge of containerization and orchestration (Docker, Kubernetes).Experience designing high-availability, fault-tolerant architectures.Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures.Excellent incident command and crisis management skills.Hands-on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD).Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human-in-the-loop patterns.Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions.Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains.Nice to HaveExperience with multi-region and multi-cloud deployments.Deep expertise in scalable microservices and event-driven architectures.Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus).Leadership in driving large-scale SRE transformations.Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices.Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows.Ability to influence engineering culture and process improvements, including adoption of AI-assisted operations under change control, safety, and audit requirements.The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non-sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well-being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section.Pay Range: $120,600.00 - $150,900.00SummaryLocation: San Francisco, CA; Chicago, IL; Dallas, TXType: Full time

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Dallas, TX vacancy
  • $128.6k - $184.9k

     ...global cloud platform. As a team of six engineers distributed across the US, Canada, and the...  ...with a strong focus on automation, reliability, and operational excellence. We are one...  ...Qualifications7+ years of experience in Site Reliability Engineering, DevOps, Infrastructure... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    CISCO Systems

    Richardson, TX
    3 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Suggested

    Google

    Sunnyvale, TX
    4 days ago
  • $147k - $210k

     ...product or system development code.Review code developed by other engineers and provide feedback to ensure best practices (e.g., style...  ..., and troubleshooting large-scale distributed systems. Site Reliability Engineering (SRE) is what you get when you treat operations... 
    Suggested

    Google

    Sunnyvale, TX
    4 days ago
  • $138.4k - $173k

     ...infrastructure as well as help improve the reliability, quality of services and overall...  ...recovery. You’ll collaborate or embed with engineering teams, helping them to improve the reliability...  ...about our locations by visiting our site.Compensation & BenefitsThe base salary that... 
    Suggested
    Full time
    Flexible hours

    AppFolio

    Dallas, TX
    4 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Dallas, TX
    4 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Dallas, TX
    2 days ago
  • Qualifications: 8+ years of Software Engineering experience, or equivalent...  ...and maintain scalable and reliable infrastructure on Google...  ...the client, IT management and staff, and other groups in Information...  ...resources Willingness to work on-site at stated location in the job... 
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Dallas, TX
    4 days ago
  •  ...Senior Site Reliability Engineer (Permanent Role) Cleveland, OH, Pittsburgh, PA, or Dallas, TX Your future duties and responsibilities . Monitoring distribution systems and notifying them of any potential issues. . Assisting with troubleshooting on call.... 
    Permanent employment
    Full time
    Local area
    Flexible hours
    Shift work
    Weekend work

    System One

    Dallas, TX
    4 days ago
  • $48 per hour

     ...Site Reliability Engineer Trident Consulting is seeking a Site Reliability Engineer for one of our clients in Richardson, Texas or Scottsdale, Arizona. Job Title: Site Reliability Engineer Location: Richardson, Texas or Scottsdale, Arizona (Onsite) Length of Assignment... 
    Contract work

    Trident Consulting

    Richardson, TX
    4 days ago
  •  ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market...  ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications.... 

    Software Technology Inc

    Dallas, TX
    3 days ago
  •  ...Senior Site Reliability Engineer This role will require someone onsite at our client office in Cleveland, OH, Pittsburgh, PA, or Dallas, TX. Systemone is looking for a Site Reliability Engineer who will work within the Production support and Performance Management... 
    Work at office
    Flexible hours
    Shift work
    Weekend work

    System One Holdings, LLC

    Dallas, TX
    5 days ago
  •  ...ensure applications are highly available, reliable, and performant at a global scale....  ...Bachelor of Computer Science or related Engineering field required. Master's Degree preferred...  ...Minimum of 1 year of lead experience of site reliability engineering team required.... 
    Contract work
    Work at office

    3B Staffing LLC

    Irving, TX
    2 days ago
  •  ...Site Reliability Engineer (SRE) The successful applicant may be performing work in FedRAMP High or IL-5 environments, and therefore, must be a U.S. Person (i.e. U.S. citizen, U.S. national, lawful permanent resident, asylee, or refugee). This position may also perform... 
    Permanent employment
    Worldwide
    Shift work

    Webex Events (formerly Socio)

    Richardson, TX
    1 day ago
  •  ...and continuously improving the platforms that power TI's digital integration, automation and DevOps capabilities. As an IT Site Reliability Engineer within the Enterprise Platforms team, you will serve as the primary technical platform owner for TI's Apigee Edge private... 
    Local area

    Texas Instruments

    Dallas, TX
    5 hours ago
  •  ...Site Reliability Engineer TXSE is building the next-generation exchange infrastructure to support transparent, efficient, and resilient capital markets. With SEC approval and $275MM in funding, we are currently hiring a Site Reliability Engineer to help with a greenfield... 
    Currently hiring

    TXSE

    Dallas, TX
    4 days ago
  •  ...generative AI and cloud-native platforms to advanced release engineering practices, our teams are redefining how financial technology...  ...AI-driven solutions that accelerate development and improve reliability. Your work will directly influence how GM Financial leverages... 
    Full time
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    2 days per week
    3 days per week

    GMAC Financial Services

    Irving, TX
    4 days ago
  •  ...Site Reliability Engineer III There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability... 
    Work at office

    Hackajob

    Dallas, TX
    2 days ago
  •  ...Senior Site Reliability Engineer Our client is seeking a Senior Site Reliability Engineer for a month 6-month contract in Irving, TX. Will be working on an onsite schedule. Contract Duration: 6 Months Required Skills & Experience ~ Bachelors/4 Year Degree ~5+... 
    Full time
    Contract work
    Temporary work
    Work experience placement
    Flexible hours

    Motion Recruitment

    Irving, TX
    1 day ago
  •  ...Job Title: Site Reliability Engineer (SRE) Work Location: RichardsonTX 75082 **FULL ONSITE WORK** Interview Mode- In-person Interview Contract duration: 9 Months Job Details: Must Have Skills Python, Kubernetes, Terraform, GitLab CI/CD Nice to... 
    Contract work

    eTeam

    Richardson, TX
    1 day ago
  • Site Reliability Engineer - Vice PresidentSite Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering...  ...best practices, and actively mentor and develop senior and staff-level engineers.Technology Evaluation & Adoption: Stay at... 

    Goldman Sachs

    Dallas, TX
    1 day ago
  • $72.1k - $158.62k

     ...person, one family and one community at a time. Position Summary We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence. The... 
    Hourly pay
    Full time
    Temporary work
    Local area

    CVS Health

    Richardson, TX
    2 days ago
  • Compliance EngineeringWe are Compliance Engineering, a global team of more than 500 engineers and scientists who work on the most complex...  ...systems by pushing for changes that improve capacity and reliability.Practicing sustainable incident management in a blameless postmortem... 

    Goldman Sachs

    Dallas, TX
    5 days ago
  • $113.1k - $232.3k

    Position Summary Lead Applied AI Site Reliability Engineer II Role Overview: As a Lead Applied AI Site Reliability Engineer II, you will actively engage in your engineering craft, taking a hands-on approach to the reliability, performance, and operational integrity... 
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    3 days per week

    Deloitte

    Dallas, TX
    4 days ago
  •  ...a company that values diversity, integrity, and growth. Role Overview PDI Technologies is looking for a Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This role owns the... 

    PDI Technologies

    Dallas, TX
    1 day ago
  • Compliance Engineering, Site Reliability Engineering, Vice President, Dallas location_on Dallas, TX, United States We are Compliance Engineering, a global team of more than 300 engineers and scientists who work on the most complex, mission-critical problems. We build and... 
    Full time
    Temporary work
    Work at office

    Goldman Sachs Bank AG

    Dallas, TX
    4 days ago
  • $207k - $300k

     ...projects.Collaborate closely with hardware, software, and system engineering teams to drive pre-silicon Firmware (FW)/Software (SW) co-...  ...AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers... 
    Worldwide

    Google

    Sunnyvale, TX
    1 day ago
  •  ...Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps‑driven environment... 
    Long term contract
    Contract work
    Internship
    Remote work

    BayOne Solutions

    Richardson, TX
    1 day ago
  • $40 per hour

     ...A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants... 
    Long term contract
    Internship
    Remote work

    BayOne Solutions

    Richardson, TX
    1 day ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Dallas, TX
    4 days ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 

    AmNet Services

    Irving, TX
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!