Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director, Site Reliability Engineering - Paze

$173k - $230k

Early Warning

At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle, Paze, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.Early Warning follows a hybrid work model to allow for a more collaborative working environment.Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role SummaryThe Director, Site Reliability Engineering leads the SRE function for an assigned product, platform, pillar or business domain and is accountable for the reliability, scalability, performance and operability of its business-critical production services. The Director develops engineering talent, establishes domain reliability strategy and priorities, and partners with Product, Software Engineering, Platform, Infrastructure, Security and Risk teams relevant to the domain.The role translates business and product priorities into measurable reliability outcomes and ensures teams have the engineering practices, capabilities and operating mechanisms needed to achieve them.Core SRE Responsibilities1. Reliability Objectives and Measurement• Establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) aligned with customer, product and business outcomes.• Use quantitative production data to measure performance and partner with domain teams to set reliability targets appropriate to business criticality, architecture, customer expectations and cost.2. Reliability Risk and Error Budgets• Use SLO performance, error budgets, failure data and production evidence to identify and prioritize reliability risk.• Ensure persistent risks have accountable engineering plans, appropriate escalation and sustained follow-through.3. Software Engineering and Automation• Champion reusable software, tooling and automation that improve reliability, scalability and operability while reducing repetitive operational work.• Ensure SRE-developed solutions follow sound engineering practices, including source control, code review, testing, maintainability and secure development.• The Director is expected to remain hands-on and maintain sufficient technical depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.4. Observability Engineering• Establish expectations for metrics, logs, traces, dashboards and telemetry sufficient to understand service behavior, customer impact and systemic risk.• Drive actionable alerting and observability that support rapid diagnosis, performance analysis, capacity planning and continuous improvement.5. Incident Management and Service Restoration• Provide domain leadership and accountability for disciplined incident response, service restoration and stakeholder communication.• Improve response through automation, runbooks, training, exercises and analysis of recurring patterns.6. Learning From Failure• Promote blameless post-incident review focused on systemic learning, prevention and measurable follow-through.• Use patterns across incidents and near misses to drive architectural, operational and domain-level improvements.7. Production Readiness, Resilience and Capacity• Ensure critical services meet appropriate production-readiness, resilience, recovery, capacity and scalability expectations.• Partner with engineering teams to address systemic failure modes, dependency risks, capacity constraints and recovery gaps before they affect customers.8. Operational Toil and Sustainable Engineering• Measure and reduce repetitive, manual and low-value operational work through engineering, automation and simplification.• Ensure operational responsibilities inform engineering priorities without becoming the primary definition of the SRE role or creating unsustainable team load.9. Security, Risk and Compliance• Partner with Security, Risk, Compliance and engineering teams relevant to the domain to meet EWS control, resilience and regulatory obligations.• Ensure material reliability risks and control gaps are visible and addressed within the domain or escalated when they exceed its authority.The Director operates within an assigned domain and delivers outcomes through its SRE teams. Impact is demonstrated by building strong teams and leaders, setting direction, resolving barriers within the domain and with direct dependencies, and creating durable engineering mechanisms. The Director leads through others while remaining sufficiently hands-on to contribute directly when appropriate. Success is measured primarily by the reliability outcomes and capabilities of the teams the Director leads.Leadership Responsibilities• Own SRE strategy, execution and measurable reliability outcomes for the assigned domain.• Build, develop and retain high-performing SRE teams with clear accountability, career development and succession.• Translate business and product priorities into reliability investments and an executable multi-quarter roadmap.• Set priorities and make evidence-based tradeoffs across reliability, delivery, operational risk, capacity and business needs.• Develop capable managers and technical leaders, ensuring decisions are made at the appropriate level.• Partner with engineering, product, infrastructure, security and risk teams relevant to the domain to embed reliability into engineering decisions.• Make reliability risk, SLO performance, operational load and improvement progress visible through effective metrics and operating reviews.• Provide accountable leadership during significant incidents and ensure systemic corrective actions are completed.Required Qualifications• Typically 12 + years of relevant software engineering, site reliability engineering, production engineering, platform engineering or closely related experience, including significant technical leadership.• 5+ years of people leadership experience, with demonstrated success leading engineering teams and developing managers and/or senior technical leaders.• Experience operating highly available, business-critical distributed systems and leading reliability improvement across a major product, platform, pillar or business domain.• Strong understanding of SRE practices, including SLOs/SLIs, error budgets, observability, incident management, automation, capacity, resilience and production readiness.• Demonstrated hands-on technical capability and sufficient depth to develop code, contribute to automation and engineering solutions, and work directly with engineers when appropriate.• Ability to communicate technical risk, tradeoffs and investment needs clearly to engineering, product and senior business stakeholders.• Demonstrated ability to build inclusive, accountable and high-performing engineering teams.Preferred Qualifications• Experience in payments, financial services or another highly regulated, high-availability environment.• Experience leading SRE or production engineering across multiple teams or a complex product or platform ecosystem.• Experience with cloud platforms, distributed systems, modern observability, infrastructure automation and software delivery at scale.• Experience establishing reliability metrics, governance and operating reviews across teams within a defined domain.The base pay scale for this position in:Phoenix, AZ/ Chicago, IL in USD per year is: $173,000 - $230,000.San Francisco, CA in USD per year is: $207,000 - $276,000.Additionally, candidates are eligible for a discretionary incentive plan and benefits.Some of the Ways We Prioritize Your Health and Happiness Healthcare Coverage – Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.401(k) Retirement Plan – Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.Paid Time Off – Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.12 weeks of Paid Parental Leave Maven Family Planning – provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.And SO much more! We continue to enhance our program, so be sure to check our Benefits page here for the latest. Our team can share more during the interview process!Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Early Warning Services, LLC (“Early Warning”) considers for employment, hires, retains and promotes qualified candidates on the basis of ability, potential, and valid qualifications without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote equal employment opportunity and affirmative action, in accordance with all applicable federal, state, and municipal laws. The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our employees. SummaryLocation: Scottsdale; Chicago; San FranciscoType: Full time

Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Director, Site Reliability Engineering - Paze in San Francisco, CA vacancy
  • $106k - $130k

     ...over thirty years with cutting-edge solutions like Zelle, Paze, and so much more. As a trusted name in payments, we partner...  ...ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering... 
    Suggested
    Hourly pay
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning

    San Francisco, CA
    4 days ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    2 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale...  ...will be instrumental in advancing the reliability, scalability, automation, observability,...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Suggested
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Suggested
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    9 hours ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Suggested

    Alembic

    San Francisco, CA
    3 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    2 days ago
  • $190.8k - $267.1k

     ...while helping Reddit grow its business. The reliability of our Ads systems directly impacts...  ...Reliability team partners closely with Ads Engineering to improve reliability, scalability,...  ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB...  ...plays a pivotal role in engineering the reliable, globally connected, multi-cloud network...  ...are seeking a talented Senior Site Reliability Engineer (SRE) with a strong... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  • $117k - $209.33k

     ...Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...principles to see it in full.About the teamThe Engineering team at Airwallex is a diverse group of...  ..., working together to build scalable, reliable, and secure products that empower...  ...Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely... 
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    3 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad business... 

    JP Morgan Chase

    San Francisco, CA
    9 hours ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $194k - $267k

     ...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $181k - $263k

     ...supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a senior... 
    Worldwide

    LiveRamp

    San Francisco, CA
    3 days ago
  • $210.38k - $243.21k

    Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the RoleWe are seeking an experienced and hands-on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring the reliability... 

    Twist Bioscience

    San Francisco, CA
    1 day ago
  • $150k - $220k

     ...companies, teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a...  ...engineers can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for... 
    Local area

    Forge Global

    San Francisco, CA
    3 days ago
  • $195k - $257.5k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and operate the infrastructure that powers our blockchain platform at... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $204k - $306k

     ...all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity, from...  ...in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions millions... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    3 days ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Full time
    H1b

    PwC

    San Francisco, CA
    3 days ago
  •  ...Bay Area · Global Offices or Remote Available What we're looking for As an SRE at Wordbricks, you will keep our systems fast, reliable, and boring. You'll own the infrastructure and operations behind our products so the rest of the team can ship without worrying... 
    Remote work
    Flexible hours

    Wordbricks, Inc.

    San Francisco, CA
    1 day ago
  •  ...manifesto. About the Role We're looking for an Infrastructure Engineer to take the lead on scaling our operational resilience as we...  ...This is a high-impact, high-trust role where you’ll shape how reliability is done - reducing incident load, building internal tooling,... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    3 days ago
  •  ...streamline both the TypeScript and Python/ML deployment pipelines to support high‑velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to streamline developer workflows, enhance tool proficiency, and... 
    Work at office
    Flexible hours

    Latent Space

    San Francisco, CA
    2 days ago
  •  ...Competitive salary Competitive salary Plus meaningful equity All roles San Francisco, CA Site Reliability Engineer San Francisco, CAFull-timeMid to SeniorOn-site Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for... 
    Full time

    Zof AI, Inc.

    San Francisco, CA
    3 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Worldwide
    Home office
    Flexible hours

    Coda

    San Francisco, CA
    3 days ago
  • $7.3 per hour

     ...The role We’re looking for a world‑class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You’ll be building and operating the core systems that power agentic AI at scale. Your mission: keep our... 

    Blaxel (YC X25)

    San Francisco, CA
    3 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $24... 
    Full time

    Alembic Technologies

    San Francisco, CA
    4 days ago
  •  ...millions of daily users while enabling our engineering teams to ship fast. You'll own the...  ...building automation and tooling that improves reliability and partnering with engineering to...  ...services What you'll bring ~5+ years in Site Reliability Engineering, DevOps, or... 
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director, Site Reliability Engineering - Paze. Be the first to apply!