Senior Site Reliability Engineer
Charles Schwab
Senior Site Reliability Engineer
At Schwab, you're empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together.
We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).
As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the reliability, scalability, and performance of Schwab's mobile and digital platforms. You will lead efforts to elevate production operations through modern Site Reliability Engineering practices, shaping how engineering teams design, build, and operate resilient systems at scale.
In this role, you will drive measurable improvements in service health and client experience by defining and executing strategies that enhance observability, automation, and system resilience. You will partner cross-functionally with engineering, architecture, infrastructure, and product teams to embed reliability, scalability, and operational excellence into the full software development lifecycle.
Success in this role requires strong problem-solving and decision-making, particularly in complex, high-scale distributed environments. You will influence technical direction, introduce best practices such as service level objectives and error budgets, and guide teams in reducing operational toil through automation and tooling innovation. Your leadership will ensure teams are aligned on reliability goals, respond effectively to production challenges, and continuously improve systems through learning and adaptation.
You will also play a key role in evolving operational maturity by strengthening on-call practices, enabling faster detection and resolution of issues, and fostering a culture of accountability, collaboration, and continuous improvement. This is an opportunity to shape enterprise-wide engineering standards while developing high-performing teams and advancing modern reliability engineering capabilities.
Key Responsibilities:
Production Operations & Incident Management
- Respond to system alerts and production incident escalations
- Lead or support incident triage, resolution, and root cause analysis
- Drive and contribute to post-incident reviews and continuous improvement actions
- Participate in an on-call rotation to support high-availability systems
Observability & Monitoring
- Ensure comprehensive monitoring coverage and effective alerting strategies across systems
- Continuously improve visibility into system performance, reliability, and health
- Define and evolve observability best practices, including telemetry, dashboards, and alert thresholds
Automation & Engineering Excellence
- Design and build automation solutions to reduce operational toil and improve resiliency
- Develop scripts and tooling using Python and shell scripting for system maintenance and performance optimization
- Contribute to CI/CD and deployment pipeline improvements
- Automate processes such as service recovery, system maintenance, and certificate management
Collaboration & Partnership
- Partner with development teams to understand system changes and ensure production readiness
- Establish guardrails for monitoring, alerting, and escalation procedures
- Embed reliability practices into the software development lifecycle
Reliability Engineering & Continuous Improvement
- Proactively identify system weaknesses, risks, and performance gaps
- Drive improvements in system reliability, scalability, and resilience
- Implement and evolve SRE best practices (SLOs, error budgets, incident reduction strategies)
Innovation & Emerging Capabilities
- Explore the use of AI and automation to improve incident detection, triage, and response
- Identify opportunities to enhance response times and reduce manual intervention
Technical Influence & Mentorship
- Mentor and support junior engineers in SRE best practices and automation techniques
- Influence engineering teams to adopt proactive reliability and observability practices
- Promote a culture of curiosity, ownership, and continuous improvement
What You Have
To ensure that we fulfill our promise of "challenging the status quo," this role has specific qualifications that successful candidates should have:
Required Qualifications:
- Bachelor of Science degree in Computer Science or a related field
- 10+ years of experience in software development and site reliability engineering, including work with cloud-native architectures and distributed systems
- 8+ years of experience in DevOps and/or site reliability engineering, with a focus on production operations, automation, and system reliability at scale
- 8+ years of experience with CI/CD pipelines, observability, and monitoring/telemetry platforms
- 5+ years of experience leading the implementation and scaling of reliability engineering practices such as service level objectives, monitoring strategies, incident reviews, and automation-driven improvements
- Demonstrated ability to design, develop, and maintain production-grade systems, automation frameworks, and reliability tooling across the software development lifecycle
- Experience supporting high-availability, distributed systems at scale
- Deep experience with monitoring, observability, and incident management practices
- Strong experience with automation, scripting, and operational tooling
Preferred Qualifications:
- Strong programming and automation experience using languages such as Python or Java, including building scalable services and APIs
- Experience with application performance monitoring tools (e.g., Splunk preferred)
- Experience with Kubernetes and container orchestration platforms
- Experience with infrastructure-as-code tools such as Terraform or similar technologies
- Familiarity with public cloud platforms (AWS, GCP, or Azure)
- Understanding of cloud infrastructure components including compute, storage, networking, load balancing, DNS, and security architectures
- Proven ability to lead in fast-paced environments, influence cross-functional teams, and drive alignment on technical strategy
- Strong communication skills with the ability to translate complex technical concepts to a variety of audiences
What Sets You Apart:
- Proven ability to operate effectively in complex, large-scale systems with evolving documentation
- Strong analytical and problem-solving skills with a proactive mindset
- Curiosity and initiative to deeply understand systems and identify improvement opportunities
- Ability to automate troubleshooting, enhance observability, and reduce manual intervention
- Strong communication skills, especially during incidents and cross-team coordination
Applicants must be currently authorized to work in the United States on a full-time basis without employer sponsorship.
In addition to the salary range, this role is eligible for bonus or incentive opportunities.
- Job Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering...SeniorFull time
- ...the team takes that seriously!The RoleThe Senior SRE at 2K is a hands-on technical leader... ...regions while partnering with network engineers, systems architects, and game studio developers... ...technical direction, influencing reliability from architecture review through production...Senior
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours- ...the rapidly evolving state of the art to engineer scalable, innovative, and research... ...business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including...SeniorFull timeLocal area
- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical...SeniorFull timeWork at office
$152k - $241.5k
...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and... ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-... ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through...SeniorFull time$127k - $249k
...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background....SeniorLocal areaRemote workWorldwideFlexible hours$210.6k - $305.1k
...Minimum Qualifications: You have led a distributed team of 5+ engineers, can demonstrate strong technical vision for your team, and ensure... ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible...SeniorFull timeTemporary workLocal areaFlexible hours- ...Schwab. We are an integrated product, engineering, strategy and risk team, all based in San... ...how we serve our clients. As a Senior Engineer on AI.x, you will play a key role... ...areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts...SeniorFull time
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$152k - $195k
...Senior Site Reliability Engineer Austin, TX (Hybrid) SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Alex Yampolskiy and...Senior- ...commercialization, and mass production to change the world for the better. JOB SUMMARY We are seeking an experienced Site Reliability Engineer to own and maintain the deployment of our cloud-based infrastructure to customer sites. In this role, you will work...SeniorFull timeLocal area
$168k - $200k
...that is passionate about creating transformative change in healthcare. What We’re Looking For We’re looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You’ll be at the forefront of building and operating a resilient, observable,...Senior- ...the layer where data becomes decisions, and decisions make the advantage. About the Role Gallatin is looking for a Site Reliability Engineer to keep our production systems running with the reliability our national security customers require. You'll work at the...SeniorFull timeLocal area
$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling...SeniorWork experience placementWork at office$110.7k - $171.8k
...components Participation in on-call rotation as a platform reliability escalation point Incident response, post-incident reviews,... ..., and internal control requirements. Collaborate with engineering teams across the organization to influence platform adoption,...SeniorWork experience placementWork at officeLocal area$185k - $227k
...professionals. If the opportunity to build your career is compelling, read on for more details. ROLE AND RESPONSIBILITIES: A Senior Site Reliability Engineer (SRE) is expected to own the operational stability and performance ofJuul’s hybrid cloud infrastructure (Nutanix, AWS/...SeniorRemote work- Recognized as the No. 1 site trusted by real estate professionals, Realtor.com has been at the forefront of online real estate... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization...SeniorWork at officeLocal area
- ...in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve Schwab's enterprise streaming platform ecosystem...SeniorFull timeWork at office
- ...Job Description Job Description Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce... ...logistics. Position Overview: We’re seeking a Senior Site Reliability Engineer to join our fast-paced Engineering...SeniorFull timeWork at office
$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...eliminated manual deployment processes.At the Senior Engineer level, your influence is...SeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week- Selby Jennings is seeking a Senior Site Reliability Engineer to scale and support critical workflow orchestration and automation platforms across the organization. The role sits in Platform Engineering, delivering highly available and resilient infrastructure for business...Senior
$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$109.65k - $182.76k
...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week$196k - $269.5k
Senior Principal AI Agent EngineerThe Software Engineering team delivers next-generation software application enhancements and new products for a changing world. Working at the cutting edge, we design and develop software for platforms, peripherals, applications and diagnostics...Senior- ...Principal Site Reliability Engineer About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years of experience helping merchants deliver better checkout experiences. Founded in 2009, we power shipping logic and checkout optimization...Full timeWork at office
$152k - $241.5k
...the world.Join the Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on... ...development by enabling Chips Simulation as a trusted and reliable virtual platform.What you will be doing:Drive early...SeniorFull time$167.18k - $203.61k
...remotely part of the weekTravel %NoWork ShiftJob DescriptionCox Automotive Corporate Services, LLCLEAD SITE RELIABILITY ENGINEERJob Description: Lead Site Reliability Engineer positions offered by Cox Automotive Corporate Services, LLC (Austin, Texas). Lead the...Full timeWork at officeRemote workFlexible hours- ...You will provide cloud operations for Oracle National Security Realms. Responsibilities Escalation points for junior site reliability engineers during complex or high-impact incidents. Manage and execute complex manual Change Management tickets, by working...Temporary workWork experience placementFlexible hoursNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Austin, TX
- site reliability engineer remote Austin, TX
- site reliability engineer sre Austin, TX
- senior maintenance supervisor Austin, TX
- senior lead project manager Austin, TX
- senior robotics software engineer Austin, TX
- senior firewall engineer Austin, TX
- senior devops engineer remote Austin, TX
- senior sas administrator Austin, TX
- senior IT manager Austin, TX


