Lead Site Reliability Engineer
Federal Reserve Bank of San Francisco
The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.
Responsibilities
System Reliability & Performance
* Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
* Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
* Lead incident response, conduct root cause analysis, and implement preventive measures
* Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
* Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
* Automate deployment pipelines, monitoring, and operational workflows
* Optimize cloud resource utilization and cost management
Engineering & Development
* Build and maintain internal tools and services to improve operational efficiency
* Collaborate with development teams to implement reliability best practices
* Conduct code reviews and provide technical guidance on system design
* Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
* Integrate security practices into CI/CD pipelines (SAST/DAST)
* Implement and maintain security controls across infrastructure and applications
* Ensure compliance with industry standards and regulatory requirements
* Conduct security assessments and vulnerability management
Leadership & Collaboration
* Mentor junior SRE team members and promote SRE culture across the organization
* Partner with software engineering teams to improve system reliability
* Drive technical initiatives and contribute to architectural decisions
* Document processes, runbooks, and technical specifications
Software Engineering:
- Strong proficiency in Java , Python , and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
- Extensive experience with AWS services including:
- Compute: Lambda, ECS, EC2, Fargate
- Storage: S3, EBS, EFS
- Database: RDS, DynamoDB, Aurora
- Networking: VPC, Route53, CloudFront, API Gateway
- Monitoring: CloudWatch, X-Ray
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
- Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
- Proficiency with configuration management tools
Security:
- Hands-on experience with SAST (Static Application Security Testing) tools
- Knowledge of DAST (Dynamic Application Security Testing) methodologies
- Understanding of security best practices, OWASP Top 10, and compliance frameworks
- Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with aws X-Ray
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
- Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
- The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA
Screening:
Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.
Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.
Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)
The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.
The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io .
Full Time / Part TimeFull time Regular / TemporaryRegular Job Exempt (Yes / No)Yes Job CategoryInformation Technology Family Group Work ShiftFirst (United States of America)The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.
Privacy Notice
- ...solving and decision-making abilities and the highest degree of professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The infrastructure cloud team is responsible for internal services that provide...Suggested
- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...SuggestedContract work
- ...in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.Site Reliability EngineerOnsite: Atlanta, GAJob SummaryAt NCR Voyix, we're looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our...SuggestedFull timeWorldwideFlexible hours
- ...Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing... ...capital and derivative markets. With a leading-edge approach to developing technology... ...people to join our team.We are seeking a Site Reliability Engineer to bring 3+ years of hands-on...Suggested
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementFlexible hours- ...of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence... ...business and technology teams.Responsibilities include leading major incident responses, driving problem management, and...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
- ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion... ...with teams to create SLI/SLO’s Actively monitor and lead troubleshooting of degraded performance and hard to define...Contract workWorldwide
$123.4k - $222.53k
...Responsibilities Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime... ...) ~ Acceptable areas of study include Computer Science, Engineering or related field (Required) ~4-7 years Working in operations...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours- ...Job Title :- Site Reliability Engineer (SRE) Employment Type :- W2 Duration :- Long Term Visa Type :- All Visa applicable which are ready for W2 Location :- Atlanta, GA (Onsite) Job Description We are seeking a highly skilled Site Reliability Engineer (SRE...
$178.13k - $205.4k
...customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re... ...~Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)...Work at officeRemote workFlexible hours$130k - $150k
...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will...Full timeRemote work$141.8k - $195k
.... We're one of the fastest-growing private companies and a leading player in a massive, fast-moving market. With a global workforce... ....Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all...Remote work- ...Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high... ...infrastructure metrics Incident Management Lead technical response for high-severity incidents Drive blameless...Worldwide
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours- ...operational efficiency, accelerate time-to-value, and deliver better customer experiences.About The RoleWe're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'...Work at officeLocal areaRemote workWork from homeWorldwideHome officeFlexible hours
- ...Overview: About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12... ...thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer...Full timeLive inWork at office
$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term... ...SRE) activities, Monitoring & Alerting Our client is a leading Airlines organization and we are currently interviewing to...Contract workLocal areaImmediate start$120k - $175k
...company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range... ...We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting...Full timeRemote workWork visaFlexible hours- ...Site Reliability Engineer At Acuity, you will join an Agile team focused on building and supporting advanced platforms and applications that... ...cross-geo team providing operational & escalation coverage, leading incident response and recovery for critical services....
- ...of this journey! We're looking for a proactive, hands-on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in... ...system performance, reliability, and scalability Leading incident response efforts, conducting postmortems, and driving...Work experience placementFlexible hours
$95k - $171k
...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:... ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to...
$121.4k - $218.6k
...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages... ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company...Work experience placementWork at office- ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE.... ...with teams to create SLI/SLO's . • Actively monitor and lead troubleshooting of degraded performance and hard to define...Work experience placement
$61.09k - $104.36k
...Site Reliability Engineer Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like... ...to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more...Permanent employmentFull timeContract workLocal area- ...rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months... ...) Job Description - Key Responsibilities: Lead and mentor a team of SREs, fostering a culture of collaboration...Contract workLocal areaImmediate start
- ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and...
- ...expertise. We deliver faster, smarter, more reliable insights to insurance carriers and... ...the right place. The Role As a Site Reliability Engineer, you'll be responsible for the... ...our AWS-hosted infrastructure. You'll lead incident response, build the automation...Flexible hours
$81.75k - $138.98k
...Job Schedule Full time Job Description As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible... ...Wesco, we build, connect, power and protect the world. As a leading provider of business‑to‑business distribution, logistics...Full timeWork at officeImmediate startWorldwideShift work- ...Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in... ...leadership skills through a variety of activities, including leading or mentoring technical staff. • Strong verbal/written communication...Immediate start
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Atlanta, GA
- lead operating engineer Atlanta, GA
- lead infrastructure engineer Atlanta, GA
- lead web developer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- site reliability engineer Atlanta, GA
- site reliability engineer remote Atlanta, GA
- official site Atlanta, GA
- site services specialist Atlanta, GA
- construction site safety Atlanta, GA



