Lead Site Reliability Engineer
Federal Reserve Bank of San Francisco
The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.
Responsibilities
System Reliability & Performance
* Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
* Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
* Lead incident response, conduct root cause analysis, and implement preventive measures
* Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
* Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
* Automate deployment pipelines, monitoring, and operational workflows
* Optimize cloud resource utilization and cost management
Engineering & Development
* Build and maintain internal tools and services to improve operational efficiency
* Collaborate with development teams to implement reliability best practices
* Conduct code reviews and provide technical guidance on system design
* Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
* Integrate security practices into CI/CD pipelines (SAST/DAST)
* Implement and maintain security controls across infrastructure and applications
* Ensure compliance with industry standards and regulatory requirements
* Conduct security assessments and vulnerability management
Leadership & Collaboration
* Mentor junior SRE team members and promote SRE culture across the organization
* Partner with software engineering teams to improve system reliability
* Drive technical initiatives and contribute to architectural decisions
* Document processes, runbooks, and technical specifications
Software Engineering:
- Strong proficiency in Java , Python , and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
- Extensive experience with AWS services including:
- Compute: Lambda, ECS, EC2, Fargate
- Storage: S3, EBS, EFS
- Database: RDS, DynamoDB, Aurora
- Networking: VPC, Route53, CloudFront, API Gateway
- Monitoring: CloudWatch, X-Ray
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
- Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
- Proficiency with configuration management tools
Security:
- Hands-on experience with SAST (Static Application Security Testing) tools
- Knowledge of DAST (Dynamic Application Security Testing) methodologies
- Understanding of security best practices, OWASP Top 10, and compliance frameworks
- Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with aws X-Ray
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
- Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
- The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA
Screening:
Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.
Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.
Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)
The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.
The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io .
Full Time / Part TimeFull time Regular / TemporaryRegular Job Exempt (Yes / No)Yes Job CategoryInformation Technology Family Group Work ShiftFirst (United States of America)The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.
Privacy Notice
$106k - $130k
...ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering... ...throughout the development lifecycle. Participate in or lead incident response, troubleshooting, service restoration, and...SuggestedHourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours- ...We are seeking a Staff Site Reliability Engineer to serve as the foundational Technical Lead for our Platform Engineering SRE organization. In this role, you will be the primary architect and visionary for the core technology foundations. As the technical lead for all...SuggestedFull time
$150k - $200k
...healthcare organization, creating unique engineering challenges around scale, reliability, security, real-time communication,... .... NOCD is looking for a Senior Site Reliability Engineer (SRE) to help... ...SLOs, and operational metrics and lead incident response and root-cause...SuggestedFull timeWork at office$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...SuggestedWork at officeLocal areaRemote workWorldwideFlexible hours$92.52k - $138.79k
...teamwork, our vision to revolutionize industries, and our goal to lead the future in media and technology, we want you to fast-... ...video advertising work.Job DescriptionWe're looking for a Site Reliability Engineer to own cloud infrastructure, system reliability, and...SuggestedFull timeWorldwide- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...
$91.2k - $136.8k
Reliability Engineer - IE08GEWe’re determined to make a difference and are proud to be an insurance... ...This position will play a crucial role to lead infrastructure resilience in ensuring... ...in Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps...Full timeTemporary workWork at office3 days per week$108.08k - $172.5k
Work with development and platform engineering teams to migrate and maintain applications in Google Cloud. Apply Observability concepts... ...dependents.CME Group: Where Futures are MadeCME Group is the world’s leading derivatives marketplace. But who we are goes deeper than that....Full timeRemote workWorldwide$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...Full timeTemporary workWork experience placementFlexible hours- Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.As a Site Reliability Engineer II at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will use technology to solve...
$130k - $180k
...belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...and documentation over process. You’ll engage in and often lead architectural discussions, reduce toil, and deliver scalable...Work at officeLocal areaRemote workWorldwideMonday to FridayFlexible hours$130k - $150k
...industry experts, and academics. At CRA you will be exposed to leading minds who use economic, financial, and business analysis to... ...is essential for this role. Position OverviewThe Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable...Work at officeWork from home3 days per week$158.5k - $172k
...velocity energy of a powerhouse startup.As a leading U.S. ordering and delivery marketplace,... ....About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will... ...high-impact position driving continuous reliability, deep system optimization, and automation...Full timeTemporary workWork at officeFlexible hours3 days per week- ...StatesIndustry: Trading FirmPosted: 2026-08-17Contact: Ethan HudsonEmail: ****@*****.***: (***) ***-****Job Title: Site Reliability Engineer (Infrastructure & Systems)Location: Chicago, IL (Greater Metro Area)About the OpportunityJoin a premier financial...Local area
$130k - $225k
...expectations, integrity, innovation and a willingness to challenge consensus.The Algorithmic Trading Team is looking for a Site Reliability Engineer for our Chicago office. The SRE team is critical to the success of our trading - ensuring that our production trading...Temporary workWork at officeFlexible hours- Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through... ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform... ...vendor resources Willingness to work on-site at stated location in the job openingDepartment...Contract workFor contractorsWork experience placement
$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in... ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational...Full timeH1b- ...and companies, alikeKlover’s engineering team powers one of the fastest... ...systems that prioritize reliability, security, and performance, and... ...candidateAbout the RoleAs a Senior/Staff Site Reliability Engineer, you... ...with engineering leads to instrument and monitor critical...Work at officeImmediate startRemote work
$160k - $210k
...NinjaTrader! As an industry-leading trading platform and futures... ...you'll do:Join our Platform Engineering team, where you'll ensure the... ...and mentoring engineers across reliability initiativesAnalyze, troubleshoot... ...of experience in DevOps, Site Reliability Engineering, or Platform...Work at officeWorldwideMonday to FridayFlexible hours$194k - $267k
...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk... ...servicesIncident Response: Participate in on-call rotations and lead post-incident reviews to drive systemic improvements and "...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$127k - $249k
...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This...Local areaRemote workWorldwideFlexible hours$204k - $306k
...mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco,... ...Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and... ...capabilities, and robust self-healing patterns.Lead, mentor, and grow a high-performing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours2 days per week- Chicago, IllinoisHybridFull Time$194k - $237kPrincipal Site Reliability Engineer An established fintech institution is seeking a Principal Site... ...observable systems throughout their lifecycle. The Principal SRE leads technical initiatives across the enterprise by leveraging...Hourly payFull timeRemote workFlexible hours
$132.1k - $220.1k
Staff Site Reliability Engineer (SRE) - Platform EngineeringNote: This position follows a hybrid work model, requiring 2 days per week on-site at... ...Engineer to serve as the foundational Technical Lead for our Platform Engineering SRE organization. In this role...Full timeWork at officeLocal areaWorldwide2 days per week$130k - $190k
...quantitative disciplines to deliver high-impact results for our clients. About the Role: Summary Responsible for the operational reliability, observability, and stability of the Strategic Full Revaluation Capability (SFRC) batch platform. This role acts as the first...Full timeTemporary workRemote workWorldwide$130k - $170k
...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to... ...leadership in suitability and risk management with industry‑leading education and the latest technology, Supernova enables advisors...Full timeFlexible hoursShift work- ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and...Full timeRemote work
- ...on this job and more exclusive features. Direct message the job poster from Algo Capital Group Senior Site Reliability Engineer - Observability and Automation A leading high-frequency trading firm is seeking a mid to senior-level Site Reliability Engineer with deep...Full timeWork at officeFlexible hours
$114k - $155k
...applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking and Payment Technology Team, you will solve complex and...Local area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Chicago, IL
- lead operating engineer Chicago, IL
- lead network engineer Chicago, IL
- lead infrastructure engineer Chicago, IL
- lead web developer Chicago, IL
- site reliability engineer sre Chicago, IL
- site reliability engineer Chicago, IL
- site reliability engineer remote Chicago, IL
- official site Chicago, IL
- site services specialist Chicago, IL



