Principal Site Reliability Engineer
$173k - $230kEarly Warning Services, LLC
At Early Warning, we've powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze℠, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Role Summary
The Principal Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle.
The role operates at enterprise scope, establishing technical direction and applying evidence-driven engineering, technical rigor, sound judgment, automation, and broad systems expertise across organizational boundaries.
Core Responsibilities
- Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed, observed, operated, and recovered.
- Use data, evidence, experimentation, and rigorous engineering analysis appropriate to the level to identify reliability risks, test assumptions, and guide technical decisions.
- Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.
- Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.
- Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
- Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.
- Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
- Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning appropriate to the level.
- Provides enterprise-level technical leadership for critical production incidents and establishes or influences engineering practices that improve incident response, escalation, service restoration and sustainable on-call operations across the organization.
- Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.
Principal represents enterprise domain-level technical leadership and organizational impact. Deep individual expertise is expected, but Principal-level impact comes from identifying systemic risk, establishing technical direction, influencing engineering practices and architecture, and multiplying the capability of the broader engineering organization.
Level Expectations
- Acts as an enterprise force multiplier, raising the effectiveness and technical capability of engineers and teams across the organization while building sustainable organizational capability rather than individual dependency.
- Demonstrates software engineering, systems thinking, troubleshooting, and production reliability capabilities appropriate to the level.
- Applies evidence-driven reasoning and technical rigor to distinguish observed facts from assumptions and make defensible engineering recommendations.
- Shares knowledge and contributes to sustainable engineering capability rather than creating dependency on individual expertise.
- Operates with significant autonomy across the organization's most consequential reliability challenges.
- Establishes enterprise technical direction, develops senior technical leaders, and demonstrates impact well beyond systems personally touched.
- Typically 15+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.
- Experience with software development or scripting using one or more modern programming languages.
- Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level.
- Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level.
- Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role.
- Hands-on experience with AWS is preferred, or comparable experience with another major cloud platform such as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
- Experience developing, deploying, operating, or improving highly available production software or distributed systems.
- Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation.
- Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness appropriate to the level.
- Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness.
- Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.
Phoenix, AZ in USD per year is: $173,000 - $230,000.
San Francisco, CA in USD per year is: $207,000 - $276,000.
Additionally, candidates are eligible for a discretionary incentive plan and benefits. This pay scale is subject to change and is not necessarily reflective of actual compensation that may be earned, nor a promise of any specific pay for any specific candidate, which is always dependent on legitimate factors considered at the time of job offer. Early Warning Services takes into consideration a variety of factors when determining a competitive salary offer, including, but not limited to, the job scope, market rates and geographic location of a position, candidate's education, experience, training, and specialized skills or certification(s) in relation to the job requirements and compared with internal equity (peers). The business actively supports and reviews wage equity to ensure that pay decisions are not based on gender, race, national origin, or any other protected classes. Physical Requirements
Early Warning works together in a highly collaborative office environment. Working conditions consist of a normal office environment. Work is primarily sedentary and requires extensive use of a computer and involves sitting for periods of approximately four hours. Work may require occasional standing, walking, kneeling, and reaching. Must be able to lift 10 pounds occasionally and/or negligible amount of force frequently. Requires visual acuity and dexterity to view, prepare, and manipulate documents and office equipment including personal computers. Requires the ability to communicate with internal and/or external customers.
Employee must be able to perform essential functions and physical requirements of position with or without reasonable accommodation.
Candidates responding to this posting must independently possess the eligibility to work in the United States at the date of hire.
Some of the Ways We Prioritize Your Health and Happiness
- Healthcare Coverage - Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.
- 401(k) Retirement Plan - Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.
- Paid Time Off - Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.
- 12 weeks of Paid Parental Leave
- Maven Family Planning - provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in San Francisco, CA vacancy
$250k
...Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves working...SuggestedFull timeRemote work- ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...Suggested
$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly...SuggestedLocal areaWorldwideFlexible hours- ...SingleStore engineers build the real-time data platform powering some of the world’s most... ...Position Summary We are seeking a Senior/Principal Software Engineer to join the... ...Demonstrated ability to design and build highly reliable, high-performance system software. ~ Experience...PrincipalFull time
$185.5k - $232k
...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating...SuggestedWork experience placementWork at officeLocal areaRelocation3 days per week$117k - $209.33k
...Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...Full timeFor contractors- ...long-term maintainability. # Own core runtime foundations: distributed control, state management, fault handling, and reliability. # Drive engineering rigor: testability, code quality, review standards, performance regression prevention, and release processes. #...PrincipalFull timeRemote work
$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,...Full timeWorldwideWeekend work- ...hyperscaler with authoritative knowledge of the Linux kernel and high-performance networking. Must possess a degree in Computer Science or Engineering and a proven track record of shipping production-grade infrastructure code. Key Skills: Linux Kernel Virtualization KVM QEMU...PrincipalTemporary work
$150k
...Job Description Job Description About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and...$61k - $101k
...Salary: $61,000 - 101,000 per year Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong familiarity with SRE culture and the practical application of reliability principles...Full time- ...guarantees and certifications. We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help...Remote workFlexible hours
- ...design of information and operational support systems. Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted. Proven working experience in installing, configuring, and troubleshooting...Full timeWork experience placementRemote workFlexible hours
$110k - $160k
...your skills and experience — talk with your recruiter to learn more. Base pay range $110,000.00/yr - $160,000.00/yr Site Reliability Engineer Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human decision in...Hourly payFull timeLocal areaRelocation packageMonday to Friday$349k - $431k
...allowing ML practitioners like you to develop multi-modal models and techniques at scale. You will report to our Director of Engineering in our Perception Organization. You will: Develop sensor-fusion foundation models for the Waymo Driver which can run on...PrincipalFull timeRemote work- ...mission. If you’re ready to shape the future of healthcare, we’d love to have you on our team! We're seeking a visionary Principal Engineer to lead our cutting-edge reporting product – a cornerstone initiative that's shaping the future of Rad AI. This highly...PrincipalFull timeWork at officeRemote workFlexible hours
- ...defenders understand and respond to cyber threats while improving the safety and reliability of frontier models in security-sensitive settings. The team works across product engineering, model training, evaluations, safeguards, and deployment to make advanced cyber capabilities...PrincipalFull time
$200k - $240k
...Senior Site Reliability Engineer (SRE) Location: San Francisco, CA Work Model: Onsite Industry: Renewable Energy Comp: $200,000 - $240,000 We're partnering with a fast-growing energy technology company looking for a Senior Site Reliability Engineer...- ...amplify your impact and a culture that backs your ambition, you won’t just contribute. You’ll make things happen–fast. As a Principal Engineer on our Advertising, Company Intelligence, and Intent team, you’ll help design and implement the core systems that power...PrincipalFull timeRemote work
- ...millions of daily users while enabling our engineering teams to ship fast. You'll own the... ...building automation and tooling that improves reliability and partnering with engineering to... ...services What you'll bring ~5+ years in Site Reliability Engineering, DevOps, or...Work at officeWork from home
$200k - $300k
...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round — Healthcare, AI Office Type: Onsite Salary: $200K–$300K Company Description We're representing a dynamic company...Work at office- ...professional like you is headed and only reach out when we genuinely believe there's a fit worth exploring Position Title: Site Reliability Engineer Location: Remote Duration: 12 Months Overview: We are seeking a highly motivated Site Reliability...Remote work
- ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...Relocation package
- ...globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You...Temporary workWorldwide
- ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating the core systems that power agentic AI at scale. Your mission: keep...
$260k - $300k
...software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding... ...faster than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that...- ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will help Shipt understand where we can improve stability and reliability. There will be a focus on the intersection of systems...
- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical...Remote work
$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...Work at office- ...functional and high leverage: researchers depend on it to run reliably, platform teams depend on it to evolve cleanly, and... ...and correctness. About the Role We're looking for a Principal Software Engineer to lead the architecture and evolution of the Agent Harness...PrincipalFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
Related searches
- senior chief engineer San Francisco, CA
- general engineer San Francisco, CA
- project engineer assistant project manager San Francisco, CA
- chief design engineer San Francisco, CA
- principal infrastructure engineer San Francisco, CA
- principal cloud engineer San Francisco, CA
- assistant chief engineer San Francisco, CA
- chief engineer San Francisco, CA
- principal developer San Francisco, CA
- senior principal engineer San Francisco, CA



