Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Federal Reserve Bank of San Francisco

CompanyFederal Reserve Bank of San Francisco When you join the Federal Reserve-the nation's central bank-you'll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we're building a dynamic and diverse team for our future.


The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.


We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities

System Reliability & Performance

* Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure

* Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance

* Lead incident response, conduct root cause analysis, and implement preventive measures

* Develop and maintain disaster recovery and business continuity plans

Infrastructure & Automation

* Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)

* Automate deployment pipelines, monitoring, and operational workflows

* Optimize cloud resource utilization and cost management

Engineering & Development

* Build and maintain internal tools and services to improve operational efficiency

* Collaborate with development teams to implement reliability best practices

* Conduct code reviews and provide technical guidance on system design

* Develop monitoring solutions, alerting systems, and observability frameworks

Security & Compliance

* Integrate security practices into CI/CD pipelines (SAST/DAST)

* Implement and maintain security controls across infrastructure and applications

* Ensure compliance with industry standards and regulatory requirements

* Conduct security assessments and vulnerability management

Leadership & Collaboration

* Mentor junior SRE team members and promote SRE culture across the organization

* Partner with software engineering teams to improve system reliability

* Drive technical initiatives and contribute to architectural decisions

* Document processes, runbooks, and technical specifications

Software Engineering:

  • Strong proficiency in Java , Python , and Node.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code

Cloud Infrastructure (AWS):

  • Extensive experience with AWS services including:
    • Compute: Lambda, ECS, EC2, Fargate
    • Storage: S3, EBS, EFS
    • Database: RDS, DynamoDB, Aurora
    • Networking: VPC, Route53, CloudFront, API Gateway
    • Monitoring: CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred

DevOps & CI/CD:

  • Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
  • Advanced Terraform skills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools

Security:

  • Hands-on experience with SAST (Static Application Security Testing) tools
  • Knowledge of DAST (Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)

Monitoring & Observability:

  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments

The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.

  • Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Screening:

Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.

Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.

Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)

The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.

The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io .

Full Time / Part TimeFull time Regular / TemporaryRegular Job Exempt (Yes / No)Yes Job CategoryInformation Technology Family Group Work ShiftFirst (United States of America)

The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.

Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.

Privacy Notice

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Atlanta, GA vacancy
  •  ...solving and decision-making abilities and the highest degree of professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The infrastructure cloud team is responsible for internal services that provide... 
    Suggested

    Black Knight Financial Services

    Atlanta, GA
    3 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Suggested
    Contract work

    2T Consulting

    Atlanta, GA
    a month ago
  •  ...in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.Site Reliability EngineerOnsite: Atlanta, GAJob SummaryAt NCR Voyix, we're looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our... 
    Suggested
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    3 days ago
  •  ...Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing...  ...capital and derivative markets. With a leading-edge approach to developing technology...  ...people to join our team.We are seeking a Site Reliability Engineer to bring 3+ years of hands-on... 
    Suggested

    Black Knight Financial Services

    Atlanta, GA
    3 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    2 days ago
  •  ...of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence...  ...business and technology teams.Responsibilities include leading major incident responses, driving problem management, and... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Atlanta, GA
    3 days ago
  •  ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion...  ...with teams to create SLI/SLO’s Actively monitor and lead troubleshooting of degraded performance and hard to define... 
    Contract work
    Worldwide

    Motion Recruitment

    Atlanta, GA
    1 day ago
  • $123.4k - $222.53k

     ...Responsibilities Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime...  ...) ~ Acceptable areas of study include Computer Science, Engineering or related field (Required) ~4-7 years Working in operations... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    5 days ago
  •  ...Job Title :- Site Reliability Engineer (SRE) Employment Type :- W2 Duration :- Long Term Visa Type :- All Visa applicable which are ready for W2 Location :- Atlanta, GA (Onsite) Job Description We are seeking a highly skilled Site Reliability Engineer (SRE... 

    Highbrow

    Atlanta, GA
    4 days ago
  • $178.13k - $205.4k

     ...customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re...  ...~​Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)... 
    Work at office
    Remote work
    Flexible hours

    Workday

    Atlanta, GA
    2 days ago
  • $130k - $150k

     ...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will... 
    Full time
    Remote work

    Prestige Staffing

    Atlanta, GA
    2 days ago
  • $141.8k - $195k

     .... We're one of the fastest-growing private companies and a leading player in a massive, fast-moving market. With a global workforce...  ....Why You'll Love This RoleCribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all... 
    Remote work

    Cribl

    Atlanta, GA
    5 days ago
  •  ...Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high...  ...infrastructure metrics Incident Management Lead technical response for high-severity incidents Drive blameless... 
    Worldwide

    Inspire Brands Inc

    Atlanta, GA
    5 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Atlanta, GA
    3 days ago
  •  ...operational efficiency, accelerate time-to-value, and deliver better customer experiences.About The RoleWe're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    Camunda

    Atlanta, GA
    5 days ago
  •  ...Overview: About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12...  ...thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer... 
    Full time
    Live in
    Work at office

    Incident IQ LLC

    Atlanta, GA
    2 days ago
  • $60 - $68 per hour

     ...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term...  ...SRE) activities, Monitoring & Alerting Our client is a leading Airlines organization and we are currently interviewing to... 
    Contract work
    Local area
    Immediate start

    Pyramid Corporation

    Atlanta, GA
    1 day ago
  • $120k - $175k

     ...company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range...  ...We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting... 
    Full time
    Remote work
    Work visa
    Flexible hours

    AEG Presents

    Atlanta, GA
    1 day ago
  •  ...Site Reliability Engineer At Acuity, you will join an Agile team focused on building and supporting advanced platforms and applications that...  ...cross-geo team providing operational & escalation coverage, leading incident response and recovery for critical services.... 

    Acuity

    Atlanta, GA
    4 days ago
  •  ...of this journey! We're looking for a proactive, hands-on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in...  ...system performance, reliability, and scalability Leading incident response efforts, conducting postmortems, and driving... 
    Work experience placement
    Flexible hours

    Rainforest

    Atlanta, GA
    4 days ago
  • $95k - $171k

     ...infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for:...  ...Akamai powers and protects life online. Leading companies worldwide choose Akamai to... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Atlanta, GA
    5 days ago
  •  ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to... 

    Next Level Business Services, Inc.

    Atlanta, GA
    4 days ago
  • $121.4k - $218.6k

     ...and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages...  ...Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company... 
    Work experience placement
    Work at office

    Akamai

    Atlanta, GA
    5 days ago
  •  ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE....  ...with teams to create SLI/SLO's . • Actively monitor and lead troubleshooting of degraded performance and hard to define... 
    Work experience placement

    Fisec Global

    Atlanta, GA
    4 days ago
  • $61.09k - $104.36k

     ...Site Reliability Engineer Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like...  ...to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more... 
    Permanent employment
    Full time
    Contract work
    Local area

    Capgemini

    Atlanta, GA
    1 day ago
  •  ...rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months...  ...) Job Description - Key Responsibilities: Lead and mentor a team of SREs, fostering a culture of collaboration... 
    Contract work
    Local area
    Immediate start

    Navtech

    Atlanta, GA
    1 day ago
  •  ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and... 

    Purple Drive

    Atlanta, GA
    2 days ago
  •  ...expertise. We deliver faster, smarter, more reliable insights to insurance carriers and...  ...the right place. The Role As a Site Reliability Engineer, you'll be responsible for the...  ...our AWS-hosted infrastructure. You'll lead incident response, build the automation... 
    Flexible hours

    Seek Now

    Atlanta, GA
    5 days ago
  • $81.75k - $138.98k

     ...Job Schedule Full time Job Description As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible...  ...Wesco, we build, connect, power and protect the world. As a leading provider of business‑to‑business distribution, logistics... 
    Full time
    Work at office
    Immediate start
    Worldwide
    Shift work

    Anixter

    Atlanta, GA
    2 days ago
  •  ...Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in...  ...leadership skills through a variety of activities, including leading or mentoring technical staff. • Strong verbal/written communication... 
    Immediate start

    Navtech

    Atlanta, GA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!