Senior AI Site Reliability Engineer, AI.x
The Charles Schwab Corporation
Your OpportunityAt Schwab, you will build a rewarding career while making a difference in the lives of our millions of clients. Here, innovative thinking meets creative problem solving as we work together to challenge the status quo. Joining Schwab means joining a company committed to transforming the financial industry and putting clients at the center of everything we do.Schwab’s AI Strategy & Transformation team, known as AI.x, is the central hub for Artificial Intelligence at Schwab. We are an integrated product, engineering, strategy and risk team, all based in San Francisco. We help set the enterprise vision for AI, invest in the most promising opportunities, and accelerate delivery across the company. We also build the core platform that powers AI at scale and explore next-generation GenAI efforts that will redefine how we serve our clients. As a Senior Engineer on AI.x, you will play a key role in bringing these priorities to life by designing and delivering innovative AI solutions.This role is an opportunity to join a high-profile team shaping Schwab’s future with AI, to build solutions that matter to millions of clients, and to grow your career in one of the most exciting areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts for cutting-edge GenAI applications that enhance the client experience and create value. You will work closely with architects and engineers to ensure scalability, reliability and security of solutions that build towards an enterprise strategy. You will lead automation-first initiatives, build robust CI/CD pipelines for one-touch deployments, and implement comprehensive observability frameworks to minimize MTTD and MTTR. This role requires participation in on-call rotations to ensure 24/7 reliability of critical AI systems. Above all, you will apply the rigor, discipline, and technical depth to help shape the next generation of AI at Schwab.Roles & Responsibilities:Lead automation-first initiatives to eliminate toil and manual interventions, defining and executing the strategic roadmap for reliability, observability, and self-healing systems across AI.x platformsDesign and implement robust CI/CD pipelines enabling one-touch deployments with automated testing, validation, and rollback capabilities to accelerate delivery velocity and reduce deployment riskImplement comprehensive observability frameworks for real-time monitoring of AI services, including metrics, logs, and traces, with intelligent alerting and automated diagnostics to minimize MTTD and MTTRParticipate in on-call rotation providing 24/7 support for production AI systems, ensuring rapid incident response, root cause analysis, and resolution with measurable SLO targetsEstablish and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and incident response runbooks to drive continuous reliability improvementsChampion Infrastructure-as-Code (IaC) practices and automate environment provisioning, configuration management, and deployment processes to ensure consistency, repeatability, and operational efficiencyCollaborate seamlessly with AI Engineering teams to integrate SRE practices early in the development lifecycle, promoting a culture of reliability and shared responsibilityProactively identify and resolve reliability, performance, and scalability issues through data-driven analysis, capacity planning, and system optimizationImplement and maintain monitoring, alerting, and incident response frameworks to ensure system health and reliability, maximizing production availabilityChampion reliability, monitoring, observability, and operational best practices for AI systems and data pipelines, establishing patterns and standards for the organizationWhat you haveRequired Qualifications8+ years of software engineering experience, with 4+ years as a hands-on Site Reliability Engineer in startups and/or large organizations.Bachelor’s degree in Computer Science or related field, or equivalent experience.5+ years building complex products from scratch, running them in production, and ensuring operational reliability.3+ years working with containers and cloud-native applications, operationalizing them in the public cloud with infrastructure as code and CI/CD pipelines.3+ years of experience working in high-availability hybrid-cloud environments.Preferred QualificationsStrong computer science fundamentals and experience across the tech stack.Experience with proprietary or open-source LLMs (e.g., Gemini, Claude, OpenAI), deploying LLM-powered applications to production and maintaining availability.Strong written and verbal communication skills to clearly convey ideas and feedback.Strong understanding of observability, incident management and reliability engineering principles.Mindset of continuous learning and improvement, adept at both giving and receiving feedback.Ability to troubleshoot complex problems with ambiguous or incomplete data in distributed systems.Curiosity about new technologies and processes, proactively sharing knowledge and seeking improvement.Experience with Terraform and Google Cloud Platform.In addition to the salary range, this role is eligible for bonus or incentive opportunities.Job SummaryRequisition ID: 2026-122527Posted Date: 19 hours ago(9/15/2026 10:54 AM)Category: Engineering & Software DevelopmentSalary Range: USD $170,000.00 - $220,000.00 / YearApplication deadline: 9/29/2026Position Type: Full time
- ...work better. Our software solutions harness the power of AI and shape the future of digitalization.We believe that our... ...thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the...SeniorTemporary workCasual workWorldwide
- ...the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play... ...Innovation & Emerging Capabilities Explore the use of AI and automation to improve incident detection,...SeniorFull timeWork at office
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of... ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... ...redefined the data platform for the AI era, enabling builders to create, transform...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure... ...role in engineering the reliable, globally connected, multi-... ...are seeking a talented Senior Site Reliability Engineer (SRE) with... ...redefined the data platform for the AI era, enabling builders to...SeniorLocal areaRemote workWorldwideFlexible hours$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our Compute Farm... ...You’ll harness the power of AI to deliver groundbreaking... ...lifecycle management, fleet reliability/auto-healing, E2E observability... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...SeniorFull time- ...TX** Our Opportunity: We are looking for a skilled engineer with disciplines that incorporate aspects of software... ...of managing and operating applications — including AI/ML-driven approaches to observability and reliability. What you’ll do: • Evangelize SRE mindset and solve...Senior
- ...branded shopping experiences across every AI channel. We work with enterprise... .... The role We\'re looking for a Senior SRE to own the reliability, scalability, and operational posture... ...development workflows Partner closely with engineering on reliability reviews and...Senior
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As... ...of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries...SeniorLocal areaRemote workWorldwideFlexible hours$127k - $249k
...zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support,... ...Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure... ...redefined the data platform for the AI era, enabling builders to create, transform...SeniorLocal areaRemote workWorldwideFlexible hours$192.4k - $275.8k
...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service... ...for you Your ImpactYou will be the most senior technical individual contributor on the... ...and protect organizations in the AI era - and beyond. We’ve been innovating...SeniorFull timeTemporary workLocal areaFlexible hours$136.2k - $214.01k
...agent-centric cybersecurity. We protect how people, data, and AI agents connect across email, cloud, and collaboration... ...Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the...SeniorFull timeFlexible hours- ...British ColumbiaTechnology – Engineering /Full-time - Permanent /RemoteAbout... ...you to apply.The RoleAs a Senior Platform Engineer, you are a... ...Be DoingImproving production reliability and system resilience within... ...use artificial intelligence (AI) tools to support parts of...SeniorPermanent employmentFull timeRemote workFlexible hours
$152k - $241.5k
...seeking a highly experienced and passionate Senior Software Engineer to join a team building large-scale... ...computing problems, helping organizations adopt AI technologies that enable faster, smarter... ...on GPU infrastructureSupport CUDA-X Libraries and integrations used in...SeniorFull timeRemote work- ...per week. The RoleWe are looking for a Senior Software Engineer to join the team that is revolutionizing... ...core focus of this role is Automation, AI Enablement, and Modernization — you will... ...systemExpertise in accessibility programming (WCAG 2.x compliance) to meet regulatory standards...SeniorFull timeLocal areaWork from homeRelocation package
- ...-15382** Onsite in Austin or Southlake 4 x weekly** Your Opportunity At Client, you’... ...our core technology infrastructure. As a Senior AI Developer, this role will be a leader in... ...coding assistants used across day-to-day engineering workflows in the SDLC—implementation, refactoring...Senior
$140k - $215k
...with the world’s most advanced AI-native platform. We work on... ...our Core Platform and Embedded Reliability charters: building the... ...embedding directly with product engineering teams and their leadership to... ...deployment processes.At the Senior Engineer level, your influence...SeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week- ...Senior Cloud Platform Engineer Location: Austin, Texas (Partial Remote... ...and a product running reliably on it in production. You... ...engineering with Site Reliability Engineering... ...platform like Apigee X from the workload side... ...agree to receive calls, AI-generated calls, text...SeniorRemote work
$168k - $270.25k
...tapping into the unlimited potential of AI to define the next era of computing.... ...Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer. This person... ...like InfiniBand, NVLink, and Spectrum-X that link GPUs and AI compute infrastructure...SeniorFull timeWeekend work$140k - $224.25k
The NVIDIA Experience (NVEX) Solutions Engineering team is looking for a senior Computer or Software Engineer who is... ...-breaking network technology used in AI clusters. Our team of software... ...for InfiniBand, NVLink, and Spectrum-X network systems that interconnect GPUs...SeniorFull timeWeekend work- ...leadership in cloud, data and AI with unmatched industry... ..., Operations, Industry X and Song, together with... ...applied AI and data engineering. We help the world’s... ....You Are:You are a Senior AI engineering Leader who... ...for LLMOps, security, reliability, and cost governance across...SeniorFull timeWork experience placementLive inWork at officeLocal area
$98.58k - $138.02k
...Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant36... ..., TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting... ....We may use artificial intelligence (AI) tools to support parts of the hiring...Full timeWork at office- ...the elite technical and product engine of the Accenture Google... ...technology: the move to Agentic AI and Product-Led Operating Models... ...Platform (GCP) Agentic AI Delivery Senior Engineer, you are a new breed of... ...Technology, Operations, Industry X and Song, together with our...SeniorFull timeWork experience placementLive inWork at officeLocal areaShift work
$198.24k - $272.58k
We’re looking for a Principal Site Reliability Engineer to join Procore’s Compute Division to work on our FedRAMP initiative. In this role, you’ll help build Procore’s next-generation construction compute platform for others to build upon, including Procore developers,...Full timeContract workWork at officeLocal areaImmediate start- ...engagement. A Cybersecurity Forward Deployed Engineer is a production engineer who works... ...their security and engineering teams—to make AI systems secure, governed, and resilient in... ...Consulting, Technology, Operations, Industry X and Song, together with our culture of...SeniorFull timeWork experience placementLive inWork at officeLocal area
$152k - $241.5k
...tapping into the unlimited potential of AI to define the next era of computing.... ...Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity... ...Chips Simulation as a trusted and reliable virtual platform.What you will be doing...SeniorFull time- ...governance, BOM/routing ownership models, and engineering change management (ECM) processesHelp... ...technology and leadership in cloud, data and AI with unmatched industry experience,... ...Consulting, Technology, Operations, Industry X and Song, together with our culture of shared...SeniorFull timeLive inWork at officeLocal areaShift work
- ...candidate for this role to work on site in the specified location(s).... ...responsible for ensuring the reliability, scalability, and operational... ...clock. As a Site Reliability Engineer, you will partner across... ...monitoring, CI/CD automation, and AI-assisted operational capabilities...Work at office
- ...Job Summary We are seeking a Site Reliability Engineer to support and administer the agency\'s Microsoft Power Platform environment, with a primary... ...and the ability to identify opportunities for automation, AI adoption, governance improvements, and platform optimization...
$106.9k - $176.5k
...build a better working world. The opportunity We are seeking AI Systems Engineers to own the security and trust fabric of EY’s AI-native... ...Keycloak/Entra ID (IAM), OpenBao (secrets store), cert-manager (X.509 lifecycle), PKI issuers/roots, and transit encryption — propagated...SeniorWork experience placementSummer holidayRemote workFlexible hours- ...Site Reliability Engineer Department: Infrastructure Employment Type: Full Time Location: Austin Reporting To: SRE Manager Description... ...you an SRE ready to grow your impact at the world's largest AI cloud-native physical security company? Following the merger...Full timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Site Reliability Engineer, AI.x. Be the first to apply!
- ai engineer Austin, TX
- ai engineer remote Austin, TX
- ai prompt engineer Austin, TX
- ai developer Austin, TX
- senior ai engineer Austin, TX
- machine learning ai engineer Austin, TX
- ai ml engineer Austin, TX
- site reliability engineer sre Austin, TX
- site reliability engineer Austin, TX
- site reliability engineer remote Austin, TX

