Lead Site Reliability Engineer
Federal Reserve Bank of San Francisco
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.
Responsibilities
System Reliability & Performance
* Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
* Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
* Lead incident response, conduct root cause analysis, and implement preventive measures
* Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
* Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
* Automate deployment pipelines, monitoring, and operational workflows
* Optimize cloud resource utilization and cost management
Engineering & Development
* Build and maintain internal tools and services to improve operational efficiency
* Collaborate with development teams to implement reliability best practices
* Conduct code reviews and provide technical guidance on system design
* Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
* Integrate security practices into CI/CD pipelines (SAST/DAST)
* Implement and maintain security controls across infrastructure and applications
* Ensure compliance with industry standards and regulatory requirements
* Conduct security assessments and vulnerability management
Leadership & Collaboration
* Mentor junior SRE team members and promote SRE culture across the organization
* Partner with software engineering teams to improve system reliability
* Drive technical initiatives and contribute to architectural decisions
* Document processes, runbooks, and technical specifications
Software Engineering:
- Strong proficiency in Java , Python , and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
- Extensive experience with AWS services including:
- Compute: Lambda, ECS, EC2, Fargate
- Storage: S3, EBS, EFS
- Database: RDS, DynamoDB, Aurora
- Networking: VPC, Route53, CloudFront, API Gateway
- Monitoring: CloudWatch, X-Ray
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
- Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
- Proficiency with configuration management tools
Security:
- Hands-on experience with SAST (Static Application Security Testing) tools
- Knowledge of DAST (Dynamic Application Security Testing) methodologies
- Understanding of security best practices, OWASP Top 10, and compliance frameworks
- Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with aws X-Ray
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
- Eligible Locations for Hire: Richmond, VA, San Francisco, CA
- The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA
Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)
The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.
The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io .
Full Time / Part TimeFull time Regular / TemporaryRegular Job Exempt (Yes / No)Yes Job CategoryInformation Technology Family Group Work ShiftFirst (United States of America)The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.
Privacy Notice
- ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ...system surfaces to maintain world-class reliability. Lead incident response with rigor: root cause analysis, post-...Suggested
$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense.... ...expects. You will own both the production reliability of our user-facing products and the... ...times. Incident Response and On-Call: Lead incident response with speed and clarity...Suggested$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering, this role will help...SuggestedPermanent employmentLocal areaWorldwideFlexible hours- ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents... ...engineering candidate to join the Site Reliability organization in San Francisco. Working... ...operational efficiency. Incident Management: Lead the coordinated response to incidents...SuggestedWorldwideWeekend work
- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...CD, ArgoCD). Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis....SuggestedFlexible hours
- ...culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software... ...here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets;...Immediate startRemote workWorldwide
$170k - $220k
...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes. Our... ...process. You'll work closely with engineers, product leads, and company leadership to ensure uptime, speed, and confidence...Work at officeRemote workFlexible hours- ...runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging workflows...WorldwideShift work
- ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate...
$117k - $209.33k
...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ...observability capabilities across supported services Lead and participate in incident response, troubleshooting, and...Full timeFor contractors$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...Temporary workImmediate startFlexible hoursShift work- ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online... ...foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at...Permanent employmentShift work
- ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will help Shipt understand where we can improve stability and reliability. There will be a focus on the intersection of systems...
- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical...Remote work
- ...the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and scalability of our systems. You...Temporary workWorldwide
- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
$230k - $310k
...daily users while enabling our engineering teams to ship fast. You'll... ...automation and tooling that improves reliability and partnering with... ...prioritize stability. You'll lead incident response, drive systemic... ...'ll bring ~5+ years in site reliability engineering, DevOps...Full timeWork at officeWork from home$194k - $267k
...something more than once, automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$195k - $257.5k
...Staff Site Reliability Engineer Circle (NYSE: CRCL) is one of the world's leading internet financial platform companies, building the foundation of a more open, global economy through digital assets, payment applications, and programmable blockchain infrastructure....Flexible hours$150k - $220k
...and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative... ...achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems...Local area$120.6k - $150.9k
...Staff Site Reliability Engineer (SRE) We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our... ...reliability engineering strategy at WEX. You'll architect and lead efforts that improve availability, performance, and...Flexible hours- ...Job: Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary... ...boundaries. Design, implement, and lead large-scale, cross-functional projects to improve the reliability...
$181k - $263k
...line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability... ...design, automation, and performance optimization standards Lead distributed systems architecture reviews and kickoffs across...Full timeWork from homeWorldwideFlexible hoursNight shift$200k - $260k
...infrastructure that makes Sight Machine the leading provider of Manufacturing Data Pipelines... ...Team as a technical leader driving reliability, automation, and scalability across the... ...reliability practices across teams, mentor senior engineers, and be a primary escalation point for...Casual workWork at officeRemote workFlexible hours- ...guarantees and certifications. We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help...Remote workFlexible hours
$61k - $101k
...Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong... ...engineering. We troubleshoot incidents and problems, lead blameless post-mortems, and drive actions to prevent recurrence...Full time- ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,... ...performance engineering Troubleshoots incidents/problems; leads blameless post‑mortems and drives non‑recurrence actions...
$221.2k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link... ...hybrid schedule as per Google policy. Responsibilities Lead a team of engineers to maintain service uptime while managing...Full timeWork at office- ...design of information and operational support systems. Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted. Proven working experience in installing, configuring, and troubleshooting...Full timeWork experience placementRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead algorithm engineer San Francisco, CA
- lead web developer San Francisco, CA
- lead network engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- lead operating engineer San Francisco, CA
- lead engineer San Francisco, CA
- site reliability engineer San Francisco, CA
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site recruiter San Francisco, CA



