Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Federal Reserve System

Company
Federal Reserve Bank of San Francisco

When you join the Federal Reserve-the nation's central bank-you'll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we're building a dynamic and diverse team for our future.


The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.


We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities

System Reliability & Performance

• Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure

• Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance

• Lead incident response, conduct root cause analysis, and implement preventive measures

• Develop and maintain disaster recovery and business continuity plans

Infrastructure & Automation

• Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)

• Automate deployment pipelines, monitoring, and operational workflows

• Optimize cloud resource utilization and cost management

Engineering & Development

• Build and maintain internal tools and services to improve operational efficiency

• Collaborate with development teams to implement reliability best practices

• Conduct code reviews and provide technical guidance on system design

• Develop monitoring solutions, alerting systems, and observability frameworks

Security & Compliance

• Integrate security practices into CI/CD pipelines (SAST/DAST)

• Implement and maintain security controls across infrastructure and applications

• Ensure compliance with industry standards and regulatory requirements

• Conduct security assessments and vulnerability management

Leadership & Collaboration

• Mentor junior SRE team members and promote SRE culture across the organization

• Partner with software engineering teams to improve system reliability

• Drive technical initiatives and contribute to architectural decisions

• Document processes, runbooks, and technical specifications

Software Engineering:
  • Strong proficiency in Java , Python , and Node.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
  • Extensive experience with AWS services including:
    • Compute: Lambda, ECS, EC2, Fargate
    • Storage: S3, EBS, EFS
    • Database: RDS, DynamoDB, Aurora
    • Networking: VPC, Route53, CloudFront, API Gateway
    • Monitoring: CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
  • Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
  • Advanced Terraform skills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools
Security:
  • Hands-on experience with SAST (Static Application Security Testing) tools
  • Knowledge of DAST (Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray
Qualifications
  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
  • Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Screening:

Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.

Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.

Base Salary Range: Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco)

The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.

The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io.

Full Time / Part Time
Full time

Regular / Temporary
Regular

Job Exempt (Yes / No)
Yes

Job Category
Information Technology Family Group

Work Shift
First (United States of America)

The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.

Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.

Privacy Notice
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in San Francisco, CA vacancy
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Suggested

    Alembic

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re...  ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Suggested
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure...  ...performance engineeringTroubleshoots incidents/problems; leads blameless post‑mortems and drives non‑recurrence... 
    Suggested

    JP Morgan Chase

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

    Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open...  ...a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate... 
    Suggested
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  •  ...it in full.About the teamThe Engineering team at Airwallex is a diverse...  ...together to build scalable, reliable, and secure products that empower...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll...  ...to the roadmap, you’ll lead on infrastructure design and... 
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    2 days ago
  • $190.8k - $267.1k

     ...Reddit grow its business. The reliability of our Ads systems directly...  ...team partners closely with Ads Engineering to improve reliability,...  .... We’re looking for a Senior Site Reliability Engineer to build...  ...Participate in on-call rotations and lead incident response efforts for... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    2 days ago
  • $139.76k - $287.75k

     ...their business.We are seeking a Senior Site ReliabilityEngineer to help operate, scale...  ...will be instrumental in advancing the reliability, scalability, automation, observability,...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  • $106k - $130k

     ...ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering...  ...throughout the development lifecycle. Participate in or lead incident response, troubleshooting, service restoration, and... 
    Hourly pay
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning

    San Francisco, CA
    3 days ago
  • $152.5k - $205k

    Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open...  ...is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries... 
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    22 hours ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB...  ...plays a pivotal role in engineering the reliable, globally connected, multi-cloud network...  ...are seeking a talented Senior Site Reliability Engineer (SRE) with a strong... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    3 days ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...and implementing operational automation at scaleExperience leading or participating in Gamedays, disaster recovery exercises,... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $195k - $257.5k

    Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open...  ...is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and... 
    Flexible hours

    Circle

    San Francisco, CA
    22 hours ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    2 days ago
  • $181k - $263k

     ...future of responsible data collaboration between the world’s leading brands, retailers, financial services providers, and...  ...line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering... 
    Worldwide

    LiveRamp

    San Francisco, CA
    2 days ago
  • $210.38k - $243.21k

    Manager, Site Reliability Engineer (Hybrid in South San Francisco)About the RoleWe are seeking an experienced and hands-on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring the reliability... 

    Twist Bioscience

    San Francisco, CA
    22 hours ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $55k - $151.47k

     ...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in...  ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational... 
    Full time
    H1b

    PwC

    San Francisco, CA
    2 days ago
  • $150k - $220k

     ...companies, teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a...  ...achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems... 
    Local area

    Forge Global

    San Francisco, CA
    2 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Worldwide
    Home office
    Flexible hours

    Coda

    San Francisco, CA
    2 days ago
  • $300 per month

     ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Site Reliability Engineer at Baseten, you will define and...  ...Experience writing and improving runbooks, leading incident response, and conducting post‑mortem... 
    Flexible hours

    Baseten

    San Francisco, CA
    2 days ago
  • $189k - $283.6k

     ...proactively and reactively improve the reliability of Block's platform and critical infrastructure...  ...0) services. In this role, you will lead incident command, coordinate mitigation,...  ...strong desire to perform and grow as an engineer ~5+ years of software development... 
    Full time
    Relocation package
    Flexible hours
    Shift work

    Block Inc

    San Francisco, CA
    22 hours ago
  •  ...a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces. The Role We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer... 
    Remote work

    Specter Services LLC

    San Francisco, CA
    22 hours ago
  •  ...daily users while enabling our engineering teams to ship fast. You'll...  ...automation and tooling that improves reliability and partnering with...  ...prioritize stability. You'll lead incident response, drive systemic...  ...you'll bring ~5+ years in Site Reliability Engineering, DevOps... 
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    4 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $24... 
    Full time

    Alembic Technologies

    San Francisco, CA
    3 days ago
  • $7.3 per hour

     ...The role We’re looking for a world‑class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...system surfaces to maintain world‑class reliability. Lead incident response with rigor: root cause analysis, post‑... 

    Blaxel (YC X25)

    San Francisco, CA
    2 days ago
  •  ...Competitive salary Competitive salary Plus meaningful equity All roles San Francisco, CA Site Reliability Engineer San Francisco, CAFull-timeMid to SeniorOn-site Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for... 
    Full time

    Zof AI, Inc.

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!