Lead Site Reliability Engineer
$146.7k - $234.3kFederal Reserve System
Company Federal Reserve Bank of San FranciscoWhen you join the Federal Reserve—the nation's central bank—you’ll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we’re building a dynamic and diverse team for our future.
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.Responsibilities
System Reliability & Performance
• Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
• Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
• Lead incident response, conduct root cause analysis, and implement preventive measures
• Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
• Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
• Automate deployment pipelines, monitoring, and operational workflows
• Optimize cloud resource utilization and cost management
Engineering & Development
• Build and maintain internal tools and services to improve operational efficiency
• Collaborate with development teams to implement reliability best practices
• Conduct code reviews and provide technical guidance on system design
• Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
• Integrate security practices into CI/CD pipelines (SAST/DAST)
• Implement and maintain security controls across infrastructure and applications
• Ensure compliance with industry standards and regulatory requirements
• Conduct security assessments and vulnerability management
Leadership & Collaboration
• Mentor junior SRE team members and promote SRE culture across the organization
• Partner with software engineering teams to improve system reliability
• Drive technical initiatives and contribute to architectural decisions
• Document processes, runbooks, and technical specifications
Software Engineering:
- Strong proficiency in Java , Python , and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
- Extensive experience with AWS services including:
- Compute: Lambda, ECS, EC2, Fargate
- Storage: S3, EBS, EFS
- Database: RDS, DynamoDB, Aurora
- Networking: VPC, Route53, CloudFront, API Gateway
- Monitoring: CloudWatch, X-Ray
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
- Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
- Proficiency with configuration management tools
Security:
- Hands-on experience with SAST (Static Application Security Testing) tools
- Knowledge of DAST (Dynamic Application Security Testing) methodologies
- Understanding of security best practices, OWASP Top 10, and compliance frameworks
- Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with aws X-Ray
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
- Eligible Locations for Hire: Richmond, VA, San Francisco, CA
- The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA
Base Salary Range: Min: $146,700 Mid: $190,500 Max: $234,300 (Location: San Francisco)
The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate’s qualifications, internal alignment considerations, district assignment, and geographic location.
The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on aiapply.co .
Full Time / Part Time
Full timeRegular / Temporary
RegularJob Exempt (Yes / No)
YesJob Category
Information Technology Family GroupWork Shift
First (United States of America)The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.
Privacy Notice
- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...SuggestedPermanent employmentWork experience placementWork at officeLocal area
$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...and implementing operational automation at scaleExperience leading or participating in Gamedays, disaster recovery exercises,...SuggestedFull timeFor contractors$114.3k - $235.32k
...verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native platform built on AWS, Kubernetes/EKS, and ArgoCD-driven...SuggestedWork at officeLocal areaRelocationRelocation package$113.4k - $162k
...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at...SuggestedTemporary work$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...SuggestedPermanent employmentLocal areaWorldwideFlexible hours$152.5k - $205k
Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open... ...is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries...Flexible hours$147k - $227k
...mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group... ...systems subject to SLA Experience leading incident response and driving operational improvements...Full timeLocal areaWorldwideFlexible hours- ...billion and backed by world-leading investors including T. Rowe Price... ...’s next.About the teamThe Engineering team at Airwallex is a diverse... ...together to build scalable, reliable, and secure products that empower... ....What you’ll doAs a Senior Site Reliability Engineer, you’ll...Temporary workLocal areaWorldwide
$152.5k - $205k
Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open... ...a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate...Flexible hours- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...
$148.5k - $223.9k
...it all.Ready to level-up your career at the company leading workforce transformation in the agentic era? You’re... ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...Full timeWorldwideWeekend work$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....Work at officeLocal areaRemote workWorldwideFlexible hours- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
$167.7k - $245.2k
...assurance insights within Cisco’s leading Networking, Security,... ...effective.We’re looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$220k - $235k
Ironclad is the leading AI contracting platform that transforms agreements into assets... ...of our cloud platform and champion engineering excellence across Ironclad. In this role... ...and strategic direction for the Site Reliability Engineering team and our broader Cloud...Full timeContract workWork at office$150k - $220k
...and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative... ...achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems...Local area$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$195k - $257.5k
Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open... ...is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and...Flexible hours$204k - $306k
...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,... ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week$194k - $267k
...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is... ...technical challenges. You will serve as a key technical lead within the EPG SRE organization, partnering with software engineers...Local areaWorldwideFlexible hours$217k - $303.9k
...Reddit grow its business. The reliability of our Ads systems directly... ...team partners closely with Ads Engineering teams to improve reliability,... ....We're looking for a Staff Site Reliability Engineer who will... ...infrastructure at Reddit.What you’ll do:Lead reliability initiatives...For contractorsWork experience placementRemote workFlexible hours$174k - $239k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for an experienced Staff TDI Site Reliability...Work experience placementLocal areaWorldwideFlexible hours$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco, CA.... ...Remote unavailable. Modality: On-Site only. Must live within... ...this role, you will take the lead on designing, deploying, and... ...scalability, performance, and reliability across environments. What You...Full timeRemote workRelocationRelocation package- ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'...
$167.7k - $245.2k
...Seattle, Austin or New York.Meet the TeamCisco ThousandEyes is a leading Digital Experience Assurance platform that empowers... ...Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure... ...Responsibilities:Cloud Security Design and Implementation: Help lead the design and deployment of security solutions for cloud...Local areaRemote workWorldwideFlexible hours- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ...CD, ArgoCD). Experience writing and improving runbooks, leading incident response, and doing post-mortem analysis....Flexible hours
- The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ...system surfaces to maintain world-class reliability. Lead incident response with rigor: root cause analysis, post-mortems...
$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense.... ...expects. You will own both the production reliability of our user-facing products and the... ...times. Incident Response and On-Call: Lead incident response with speed and clarity...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead product engineer San Francisco, CA
- lead app. developer San Francisco, CA
- lead web developer San Francisco, CA
- lead industrial engineer San Francisco, CA
- lead algorithm engineer San Francisco, CA
- lead infrastructure engineer San Francisco, CA
- lead network engineer San Francisco, CA
- lead engineer San Francisco, CA
- lead operating engineer San Francisco, CA
- lead system engineer San Francisco, CA

