Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$146.7k - $234.3k
Full-time

Federal Reserve System

Company Federal Reserve Bank of San FranciscoWhen you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.

We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities

System Reliability & Performance

• Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure

• Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance

• Lead incident response, conduct root cause analysis, and implement preventive measures

• Develop and maintain disaster recovery and business continuity plans

Infrastructure & Automation

• Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)

• Automate deployment pipelines, monitoring, and operational workflows

• Optimize cloud resource utilization and cost management

Engineering & Development

• Build and maintain internal tools and services to improve operational efficiency

• Collaborate with development teams to implement reliability best practices

• Conduct code reviews and provide technical guidance on system design

• Develop monitoring solutions, alerting systems, and observability frameworks

Security & Compliance

• Integrate security practices into CI/CD pipelines (SAST/DAST)

• Implement and maintain security controls across infrastructure and applications

• Ensure compliance with industry standards and regulatory requirements

• Conduct security assessments and vulnerability management

Leadership & Collaboration

• Mentor junior SRE team members and promote SRE culture across the organization

• Partner with software engineering teams to improve system reliability

• Drive technical initiatives and contribute to architectural decisions

• Document processes, runbooks, and technical specifications

Software Engineering:

  • Strong proficiency in Java , Python , and Node.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code

Cloud Infrastructure (AWS):

  • Extensive experience with AWS services including:
    • Compute: Lambda, ECS, EC2, Fargate
    • Storage: S3, EBS, EFS
    • Database: RDS, DynamoDB, Aurora
    • Networking: VPC, Route53, CloudFront, API Gateway
    • Monitoring: CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred

DevOps & CI/CD:

  • Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
  • Advanced Terraform skills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools

Security:

  • Hands-on experience with SAST (Static Application Security Testing) tools
  • Knowledge of DAST (Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)

Monitoring & Observability:

  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments

The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.

  • Eligible Locations for Hire: Richmond, VA, San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Base Salary Range: Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco)

The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate’s qualifications, internal alignment considerations, district assignment, and geographic location.

The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on aiapply.co .

Full Time / Part Time

Full time

Regular / Temporary

Regular

Job Exempt (Yes / No)

Yes

Job Category

Information Technology Family Group

Work Shift

First (United States of America)

The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.

Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.

Privacy Notice

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in San Francisco, CA vacancy
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    6 hours ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...and implementing operational automation at scaleExperience leading or participating in Gamedays, disaster recovery exercises,... 
    Suggested
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    2 days ago
  • $114.3k - $235.32k

     ...verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native platform built on AWS, Kubernetes/EKS, and ArgoCD-driven... 
    Suggested
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Suggested
    Temporary work

    TextNow

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

    Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open...  ...is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $147k - $227k

     ...mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group...  ...systems subject to SLA Experience leading incident response and driving operational improvements... 
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  •  ...billion and backed by world-leading investors including T. Rowe Price...  ...’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...together to build scalable, reliable, and secure products that empower...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

    Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open...  ...a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate... 
    Flexible hours

    Circle

    San Francisco, CA
    6 hours ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 

    Alembic

    San Francisco, CA
    1 day ago
  • $148.5k - $223.9k

     ...it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re...  ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    6 hours ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    3 days ago
  •  ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    1 day ago
  • $167.7k - $245.2k

     ...assurance insights within Cisco’s leading Networking, Security,...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    6 hours ago
  • $220k - $235k

    Ironclad is the leading AI contracting platform that transforms agreements into assets...  ...of our cloud platform and champion engineering excellence across Ironclad. In this role...  ...and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    6 hours ago
  • $150k - $220k

     ...and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative...  ...achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems... 
    Local area

    Forge Global

    San Francisco, CA
    2 hours ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $195k - $257.5k

    Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open...  ...is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $204k - $306k

     ...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,...  ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and...  ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    1 day ago
  • $194k - $267k

     ...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is...  ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $217k - $303.9k

     ...Reddit grow its business. The reliability of our Ads systems directly...  ...team partners closely with Ads Engineering teams to improve reliability,...  ....We're looking for a Staff Site Reliability Engineer who will...  ...infrastructure at Reddit.What you’ll do:Lead reliability initiatives... 
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    San Francisco, CA
    2 hours ago
  • $174k - $239k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    6 hours ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA....  ...Remote unavailable. Modality: On-Site only. Must live within...  ...this role, you will take the lead on designing, deploying, and...  ...scalability, performance, and reliability across environments. What You... 
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    2 days ago
  •  ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'... 

    Alembic

    San Francisco, CA
    3 days ago
  • $167.7k - $245.2k

     ...Seattle, Austin or New York.Meet the TeamCisco ThousandEyes is a leading Digital Experience Assurance platform that empowers...  ...Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure...  ...Responsibilities:Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...CD, ArgoCD). Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis.... 
    Flexible hours

    Baseten

    San Francisco, CA
    5 days ago
  • The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...system surfaces to maintain world-class reliability. Lead incident response with rigor: root cause analysis, post-mortems... 

    Blaxel, Inc

    San Francisco, CA
    3 days ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer. Our team is extremely talent-dense....  ...expects. You will own both the production reliability of our user-facing products and the...  ...times. Incident Response and On-Call: Lead incident response with speed and clarity... 

    Cognition AI

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!