SRE Architect, AI-Powered Reliability
$200.6k - $250.4kWex Health
About the Team & RoleWEX operates across multiple lines of business, Mobility, Benefits, and Travel, serving enterprise customers globally with payment and technology solutions that demand uncompromising reliability. These are mission-critical systems handling high-volume financial transactions where availability, transactional integrity, and low latency are non-negotiable. Our SRE practice is in its early stages, and the decisions made now will define how we build, operate, and continuously improve reliable systems for years to come.This person will define and enforce the reliability standards, operational practices, and architectural guardrails that every line of business at WEX must meet, and will use AI as a primary tool to establish, scale, and continuously improve those standards faster than traditional approaches alone can achieve.This is not a role embedded in a single business unit. It sits at the center of WEX engineering with a mandate that spans all LOBs. You will set the bar, and you will hold it , working with engineering leadership, platform teams, and LOB architects to make reliability a consistent, measurable, and continuously improving property of every system we operate.How you'll make an impactEnterprise Standards & GovernanceDefine, publish, and enforce enterprise-wide SRE best practices and operational standards covering observability, incident management, resilience, capacity planning, and reliability architecture, applicable across all WEX lines of business.Define and lead WEX’s AI-Powered Reliability Engineering strategy, driving adoption of SRE agents across the software lifecycle—from design and development through deployment and operations, to improve reliability, automation, and operational efficiency.Architect and oversee the implementation of mission-critical systems, ensuring that reliability, availability, and transactional integrity requirements are designed in from the start, not bolted on after the fact.Establish and govern SLO, SLI, and error budget frameworks across LOBs, partnering with engineering leadership to align reliability targets with business and commercial expectations. Own the production readiness review process, defining the criteria every service must meet before going live and driving accountability for remediation when gaps are found.Serve as the primary technical advisor to engineering leadership across WEX on matters of reliability, resilience architecture, and operational excellence.Observability Define the enterprise observability standard, what good looks like for metrics, distributed tracing, structured logging, and alerting, and hold all LOBs accountable to it.Use AI-powered tooling to move beyond static dashboards: deploy intelligent anomaly detection, dynamic baselining, and automated signal correlation to reduce noise and surface actionable signals at scale.Drive instrumentation practices that give engineering teams genuine insight into the health of high-availability, low-latency systems, including real-time payment flows and transaction pipelines where latency and consistency are critical.Lead the evaluation and adoption of AI-assisted observability platforms that reason across telemetry sources to accelerate detection and diagnosis.Incident ManagementEstablish the enterprise incident management framework: severity definitions, response playbooks, escalation paths, on-call standards, and cross-LOB communication protocols.Integrate AI into the full incident lifecycle, intelligent triage and automated runbook suggestions at detection, real-time signal correlation during active incidents, and AI-assisted timeline and impact summaries at resolution.Reduce cognitive burden on on-call engineers through tooling that surfaces relevant context, prior incidents, and likely remediation paths automatically during high-pressure situations.Define, track, and report on incident metrics (MTTD, MTTR, recurrence rate) across all LOBs, using trends to drive systemic improvement rather than one-off fixes.Resilience Engineering & Self-Healing SystemsLead cross-functional initiatives to enhance system resilience and performance across WEX, advocating for circuit breakers, bulkheads, graceful degradation, retry strategies, and fault isolation as enterprise standards.Design self-healing and auto-recovery mechanisms that allow systems to detect, respond to, and recover from common failure modes without human intervention, reducing toil and improving mean time to recovery.Build and operate chaos engineering programs appropriate for WEX's financial systems, running controlled failure experiments that expose resilience gaps safely and systematically before they manifest as production incidents.Use AI to proactively identify resilience risks: analyze production telemetry, deployment signals, and dependency graphs to surface systems most likely to fail under stress before incidents occur.Capacity Planning & Load TestingDevelop enterprise capacity planning strategies, establishing the models, tooling, and review cadences that ensure every LOB can anticipate and provision for demand growth without last-minute scrambles or over-provisioning.Define and enforce load testing standards as a gate in the software delivery lifecycle, ensuring that services can handle peak transactional load, including burst demand on payment and fleet systems, before they reach production.Apply AI-driven forecasting to capacity planning: model historical growth patterns, seasonal demand signals, and business pipeline data to produce reliable capacity outlooks across LOBs.Cloud Cost OptimizationDrive cloud cost optimization and budgeting initiatives across WEX engineering, establishing the frameworks, tooling, and governance processes that ensure cloud spend is rationalized against reliability and performance outcomes.Identify and remediate cost inefficiencies without compromising availability: right-sizing, reserved capacity strategy, workload scheduling, and architecture patterns that reduce waste in high-availability deployments.Partner with LOB engineering and finance leadership to produce credible cloud cost forecasts, and hold teams accountable to efficiency targets.Blameless Postmortem CultureDesign and champion the enterprise blameless postmortem process, creating templates, facilitation standards, and review cadences that make postmortems genuinely useful and consistently practiced across all LOBs.Use AI to accelerate postmortem quality: generate draft timelines from incident telemetry, surface contributing factors from logs and traces, and identify systemic patterns across multiple incidents over time.Build a postmortem knowledge base that is searchable and actionable, so lessons from past incidents actively inform future architectural decisions and operational practices.Close the loop on postmortem action items, tracking completion rates across LOBs and escalating chronic non-compliance to engineering leadership.Technical Advisory & Cross-LOB EnablementServe as a technical advisor to engineering leaders and architects across WEX, reviewing system designs for reliability risk, providing guidance on high-availability and low-latency architecture patterns, and advising on operational tradeoffs.Lead cross-functional initiatives that span LOBs, bringing together engineering teams to solve shared reliability challenges, establish common tooling, and align on enterprise standards.Create and deliver internal enablement programs, workshops, documentation, office hours, and design review forums, that build SRE capability across WEX engineering without requiring headcount growth in every team.Communicate clearly and influentially to senior leadership: produce written strategy documents, present reliability trends and investment recommendations, and maintain executive visibility into the state of reliability across the enterprise.Experience you'll bring Required12+ years in SRE, platform engineering, or distributed systems, with a hands-on track record of operating mission-critical systems at scale.Deep practical expertise across observability, incident management, resilience engineering, and capacity planning, not just familiarity, but proven delivery in production environments.Experience with high-availability, low-latency systems where transactional integrity and consistency are critical requirements, payment processing, financial platforms, or equivalent.Demonstrated experience using AI tools to solve real reliability problems: anomaly detection, incident triage, noise reduction, postmortem acceleration, capacity forecasting, or auto-remediation.Proven ability to define and enforce technical standards across multiple engineering teams or business units without direct managerial authority.Experience designing self-healing and auto-recovery mechanisms in production distributed systems.Strong background in cloud cost optimization, architecture patterns, governance frameworks, and tooling for managing cloud spend at scale (AWS, GCP, or Azure).Excellent written and verbal communication skills, able to produce authoritative strategy documents, lead cross-LOB forums, and advise VP and C-level engineering leaders.PreferredExperience in payments, fintech, fleet technology, or benefits administration, familiarity with the reliability and compliance demands of financial transaction systems.Experience building or maturing an SRE practice from an early stage across a multi-product or multi-LOB organization.Familiarity with AI-native observability or AIOps platforms (Dynatrace, Honeycomb, Coralogix, or similar).Background in chaos engineering (Gremlin, LitmusChaos, AWS Fault Injection Simulator) and controlled failure experimentation in regulated or financial environments.Experience with systems requiring strict transactional consistency, distributed databases, event-driven architectures, or payment settlement pipelines.Proficiency with Kubernetes, service mesh (Istio/Linkerd), and OpenTelemetry-based observability stacks.BS/MS in Computer Science, Engineering, or equivalent practical experience.The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non-sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well-being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section.Pay Range: $200,600.00 - $250,400.00SummaryLocation: Portland, ME; Chicago, IL; Boston, MA; Seattle, WA; Dallas, TXType: Full time
- NewRocket is a leading AI-focused consultancy partnering with Anthropic to deliver Claude-powered enterprise solutions. We’re hiring an Agentic AI Architect to design scalable AI systems, integrate with ServiceNow, and lead technical delivery for global clients. Expect...Suggested
$134.9k - $237.3k
...Be the one building AI-powered experiences where they matter most. At Genesys, we help... ...day. Role Overview: The Senior AI Architect, Presales will shape how global enterprises... ...that turn advanced capabilities into reliable, scalable business outcomes. At Genesys,...SuggestedWork from homeWorldwideFlexible hours- What You Will Be Doing AI/ML Architecture & Solution Delivery Agentic AI Architect-Anthropic Partnership Why Us NewRocket is... ...adopt and operationalize Claude-powered AI solutions. Through this relationship... ...business needs into secure, reliable, and production‑ready AI...Suggested
$88.54k - $207.4k
...we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential... ...About the job you’re considering Capgemini is hiring an AI-Powered SDLC Architect to design and run the AI-Powered software...SuggestedPermanent employmentFull timeLocal area- ...generation computing experiences—from AI and data centers, to PCs,... ...an experienced Platform Architect to help shape the future of AI... ...rack-scale systems that power hyperscale and cloud computing... ...optimize performance, scalability, reliability, cost, and customer value....Suggested
$180k - $210k
...companies modernize through shared AI expertise and operational... ...are hiring our first Lead AI Architect to design, build, and run the... ...centralized data and AI platform that powers Banyan. The cornerstone of the... ...that keep the platform reliable and improving as adoption grows...Permanent employmentFull timeFor contractorsStart working todayWork at officeWorldwide3 days per week- TikTok USDS Joint Venture is seeking an experienced Site Reliability Engineer Leader to oversee a team building a globally distributed, observable... ...in a fast-paced environment. The role requires deep SRE expertise, leadership skills, and hands-on coding in Go/Python/...
$110k - $145k
BCG X is seeking a Growth Architect to join their Seattle office. This role focuses on bringing innovative thinking and strategies to drive... ...growth for digital businesses. Candidates should have a strong AI mindset, experience in marketing automation, and a degree in a relevant...Work at office$135.5k - $203.3k
...scale. We need deep knowledge of modern AI technologies, including machine learning... ..., observability, performance, reliability, and cost efficiency. We design AI solutions... ...outcomes. Technologies: AI AWS Architect Azure Cloud Databricks GIS...Full timeTemporary work- ...Azure DevOps & GitHub Engineer/Architect at the Manager level. In this... ...help clients define their AI-enabled engineering vision, translate... ...gets passionate about the power of digital to transform... ...practices.Continuously improve reliability, security, governance, compliance...Local area
$169.62k - $237.47k
...bold, experienced engineers who are ready to push the boundaries of what's possible in space power systems. We are seeking a Principal Electrical Power Subsystem (EPS) Architect to join our growing BNS team. In this role, you will serve as a technical authority and key...Permanent employmentTemporary workLocal areaRelocationRelocation package- Stripe is seeking a Forward Deployed AI Accelerator to embed with a group of ~20 marketers, building AI-powered tools and workflows that transform everyday work. You will identify high-leverage opportunities, coach others, and scale solutions across the cohort. This role...
$224k - $356.5k
SoC Product Architect Low Power SoC page is loaded## SoC Product Architect Low Power SoClocations: US, CA, Santa Clara: US, WA, Seattletime type... ...people.Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU...$122k - $240.5k
Position Summary Google AI Architect/AI and EngineeringJoin our AI & Engineering team... ...infrastructure. These solutions are powered by engineering for business advantage,... ...and Gemini; optimize for scalability, reliability, security, and cost.Design, fine-tune,...Local areaVisa sponsorshipFlexible hours- YO AI Labs is seeking an experienced Business Intelligence & Analytics Expert to evaluate and improve AI-powered analytics workflows. You will review dashboards, KPIs, reports, and data visualizations to ensure accuracy and actionable insights in support of AI training...Remote jobFor contractors
$135.5k - $203.3k
...excellence, and responsible stewardship. As Principal AI Architect, you will define and evangelize AI architectures that power Weyerhaeuser’s digital transformation.... ...governance, transparency, observability, performance, reliability, and cost efficiency. Ensure AI solutions are...Full timeTemporary work- Blue Origin is seeking a Principal Electrical Power Subsystem Architect to lead EPS design for LEO satellite programs and new space technologies. You will shape power architecture, budgets, and interfaces, coordinating with systems, thermal, avionics and software teams...
- ...engineering fundamentals with hands‑on Generative AI expertise and exceptional customer‑facing... ...RAG systems, tool‑calling workflows, LLM‑powered applications, or AI workflow automation.... ...Engineer, Solutions Engineer, Solutions Architect, Technical Consultant, or customer‑facing...Full timeWork at office2 days per week3 days per week
- Genesys is seeking an AI Architect, Presales to shape how global enterprises apply AI to customer experience, designing architectures that deliver reliable, scalable outcomes. You will guide technical direction and connect Genesys Cloud AI with enterprise data, workflows...Remote jobWork at office
$152.07k - $202.76k
...the trusted network for the AI‑powered world, connecting people, data... ...strategic and technical Principal Architect to design the new... ...with a focus on scalability, reliability, and performance. Work with... ...Collaborate with DevOps and SRE teams to embed reliability and...Full timeTemporary workRemote work- Salesforce is seeking an Agentforce Success Architect to drive AI-driven agent design and autonomous data-powered solutions for strategic customers. You will influence technical architecture, orchestration, and deployment of Agentforce across sales, service, and marketing...
$142.6k - $261.5k
...help to build a better working world. ServiceNow- ServiceNow AI Architect Manager In the digital economy, it takes more than good ideas... ...Skills Kit, including designing, training, and optimising AI-powered virtual agents for tailored user experiences. Proven ability...Summer holidayWorldwideFlexible hours- ...integration, automation, and emerging AI technologies.Our teams work... ...enterprise integration to AI-powered and agentic architectures, we... ...discussions, and support architects during architecture and... ...handling, and performance or reliability troubleshooting.Minimum of 4...Full timeWork experience placementLive inWork at officeLocal areaFlexible hours
- Docusign seeks a strategic leader to design and govern an AI-powered prospecting program that spans 4 regions and 12 languages. You will define agent capabilities, architect signal-to-message workflows, and oversee a growing portfolio of experiments driving pipeline and...
- ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ...deep technical skills of Site Reliability Engineering with the consultative... ..., you'll lead the Solutions Architects who own our most complex... ...team (Solutions Architecture, SRE, TAM, professional services, or...Permanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours
$110k - $176k
...Description Job Description Shield AI is a venture-backed defense-... ...help deliver scalable and reliable site infrastructure, including... ...layouts, structured cabling, power and cooling capacity, MDF/IDF... ...of work. * Coordinate with architects, contractors, colocation...Full timeTemporary workPart timeFor contractorsWork at officeWorldwideRelocation$90 - $110 per hour
...seeking an experienced Integration & API Architect to provide architectural leadership for... ...integration solutions meet scalability, reliability, performance, and maintainability requirements... ...origin, age, disability, veteran status. Powered by JazzHR Grwh7IbgRT...Contract workPart timeFor contractorsRemote work$236k - $275k
...Principal Architect Hybrid Seattle preferred Full-Time We... ...work gets done in the age of AI. Boundless is a Series C legal... ...navigate immigration. We build tech-powered immigration solutions for two... ..., at the scale required to reliably serve thousands of people...Full timeTemporary workWork at officeImmediate startRemote workShift work2 days per week- ...Weyerhaeuser is seeking a Principal AI Architect to lead enterprise AI strategy across our sustainable manufacturing and forestry operations. You will design scalable, secure AI and ML architectures that optimize mill performance, supply chains, and forest resource management...
$187.6k - $220.7k
Are you ready to make an impact?Agentic AI Architect-Google Cloud West Monroe is seeking an experienced Agentic AI Architect to join our... ...Knowledge of cloud networking, IAM, security, observability, reliability engineering, and enterprise platform governance. Experience...Local areaImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE Architect, AI-Powered Reliability. Be the first to apply!




