Lead Site Reliability Engineer
Federal Reserve Bank of San Francisco
The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.
We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.
Responsibilities
System Reliability & Performance
* Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure
* Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance
* Lead incident response, conduct root cause analysis, and implement preventive measures
* Develop and maintain disaster recovery and business continuity plans
Infrastructure & Automation
* Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)
* Automate deployment pipelines, monitoring, and operational workflows
* Optimize cloud resource utilization and cost management
Engineering & Development
* Build and maintain internal tools and services to improve operational efficiency
* Collaborate with development teams to implement reliability best practices
* Conduct code reviews and provide technical guidance on system design
* Develop monitoring solutions, alerting systems, and observability frameworks
Security & Compliance
* Integrate security practices into CI/CD pipelines (SAST/DAST)
* Implement and maintain security controls across infrastructure and applications
* Ensure compliance with industry standards and regulatory requirements
* Conduct security assessments and vulnerability management
Leadership & Collaboration
* Mentor junior SRE team members and promote SRE culture across the organization
* Partner with software engineering teams to improve system reliability
* Drive technical initiatives and contribute to architectural decisions
* Document processes, runbooks, and technical specifications
Software Engineering:
- Strong proficiency in Java , Python , and Node.js
- Experience with microservices architecture and distributed systems
- Solid understanding of data structures, algorithms, and design patterns
- Proficiency in writing clean, maintainable, and testable code
Cloud Infrastructure (AWS):
- Extensive experience with AWS services including:
- Compute: Lambda, ECS, EC2, Fargate
- Storage: S3, EBS, EFS
- Database: RDS, DynamoDB, Aurora
- Networking: VPC, Route53, CloudFront, API Gateway
- Monitoring: CloudWatch, X-Ray
- AWS certifications (Solutions Architect, DevOps Engineer) preferred
DevOps & CI/CD:
- Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
- Advanced Terraform skills for infrastructure provisioning and management
- Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
- Proficiency with configuration management tools
Security:
- Hands-on experience with SAST (Static Application Security Testing) tools
- Knowledge of DAST (Dynamic Application Security Testing) methodologies
- Understanding of security best practices, OWASP Top 10, and compliance frameworks
- Experience with secrets management and identity access management (IAM)
Monitoring & Observability:
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
- Log aggregation and analysis (CloudWatch Logs, Splunk)
- Distributed tracing with aws X-Ray
Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
- 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
- 3+ years in a lead or senior technical position
- Proven track record of managing large-scale production systems
- Experience with on-call rotations and incident management
- GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
- Experience with serverless architectures and event-driven systems
- Familiarity with chaos engineering principles and practices
- Background in Agile/Scrum methodologies
- Experience with multi-cloud or hybrid cloud environments
The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.
- Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
- The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA
Screening:
Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.
Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.
Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)
The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.
The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io .
Full Time / Part TimeFull time Regular / TemporaryRegular Job Exempt (Yes / No)Yes Job CategoryInformation Technology Family Group Work ShiftFirst (United States of America)The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.
Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.
Privacy Notice
- Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through... ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform... ...vendor resources Willingness to work on-site at stated location in the job openingDepartment...SuggestedContract workFor contractorsWork experience placement
$255.7k - $300k
...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system... ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation...SuggestedFull time- ...cloud-native platforms to advanced release engineering practices, our teams are redefining how... ...in coding, testing, and automation. Reliability Engineering: Establish service level... ...Scrum teams with demonstrated success leading improvements (getting better/faster/happier...SuggestedH1bWork at officeRemote workVisa sponsorshipFlexible hours2 days per week3 days per week
- ...part of our global expansion, we're looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our... ...reliability and resiliency are built in from day one. Lead incident response: Drive on-call processes, conduct root-...SuggestedRemote work
$140k - $150k
...learn more. Base pay range $140,000.00/yr - $150,000.00/yr Site Reliability Engineer II | 6-month Contract to Hire | Hybrid (Irving, TX) | 2x onsite per week Optomi, in partnership with a leading financial services company, is seeking a highly skilled, hands‑on...SuggestedFull timeContract work$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...Full timeFlexible hours$55k - $151.47k
...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in... ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational...Full timeH1b$192.4k - $275.8k
...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...facing, and senior leadership audiences 4+ years experience leading post-mortems and root cause analysis for high-severity...Full timeTemporary workLocal areaFlexible hours$172k - $300k
Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property... ...Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more...Full timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$72.1k - $158.62k
...person, one family and one community at a time. Position Summary We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence. The...Hourly payFull timeTemporary workLocal area- ...applications, databases, etc. # Set up SLOs and SLIs using industry-leading tools. # Play the role of an individual contributor and lead... ...self-healing solutions. # Experience in Implementing Chaos Engineering/testing. Seniority level Mid-Senior level...Full time
- Mandatory Skills: AWS/Azure/GCP (GCP is not used very much ). Kubernetes /Helm,Docker,Gitlab,Grafana,Cyberark/Hashicorp Vault, Terraform etc. Experience utilizing Java, Perl, Python, Go and scripting experience in Shell and Perl to automate reports and monitor enterprise...
- ...healthcare fintech innovator, we’re transforming the patient journey and redefining what’s possible in dental care. This role: Site Reliability Engineer (SRE) with deep expertise in monitoring, debugging, and optimizing Azure App Services. This position is critical to...Full timeWork at office3 days per week
- ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market... ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications....
- ...confidential information into the tool.Interactions with the tool are reviewed in order to improve results.If you do not agree with any part of this notice, please close the tool and use the career site. For questions or feedback regarding the tool, submit an HR Connect ticket.
$119k - $170k
...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler...Full timeWork at officeLocal areaRemote workShift work3 days per week- ...Senior Site Reliability Engineer (Permanent Role) Cleveland, OH, Pittsburgh, PA, or Dallas, TX Your future duties and responsibilities... .... Facilitating analysis meetings to discuss incidents. Lead . Identify the automation opportunities for automation specialists...Permanent employmentTemporary workLocal areaFlexible hoursShift workWeekend work
- Company DescriptionAmerica Networks is a leading sensor and networking solutions partner for companies in any Industrial, Manufacturing... ...asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing,...
$160k - $225k
...experience solutions. Our partnerships with leading cloud, design and business intelligence... ...needed on critical paths, establishing engineering guardrails, and leading design reviews.... ...and guide tradeoffs across reliability, performance, and delivery speedHands on...Permanent employmentFull timeTemporary workRemote work- ...stores and communities every day. If you're ready to grow, lead and make a difference, come join our team and help shape the future of convenience.The SRE RunOps Engineer 2 is responsible for ensuring the reliability, availability, and performance of the 7NOW delivery...Hourly payWork experience placement
- ...fulfill travel worldwide.SRE Software Systems Engineer IV - Data Intelligence and AI... ...Systems Engineer, you will drive platform reliability, auto-scaling and cloud cost efficiency... ...running smoothly. This role requires strong Site Reliability Engineering discipline, problem...Full timeWorldwideFlexible hoursWeekend work
- ...874863Reference Number: 25-00760Title: AWS Python ML Developer - Lead LevelPosted Date: 2025-07-10Company: HAN StaffingRole: AWS Python... ...interviewRound 2: Behavioral or combined technical/behavioralA final on-site interview may be required for top candidates to validate skills...Full timeRelocation3 days per week
- Smart Tech Contracting LLC is seeking a Lead System Integrator - Critical Infrastructure to drive engagement in large-scale BAS/EPMS... ...commissioning. You will lead teams, coordinate with design engineers, contractors, vendors and clients to ensure successful delivery...Remote jobFor contractors
- Smart Tech Contracting, LLC is seeking a Lead System Integrator - Critical Infrastructure to guide large-scale BAS/EPMS projects through... ...supervise a team of system integrators, coordinate with design engineers and vendors, and ensure the system meets contract requirements....Contract work
- ...ContractPay Rate: $40/Hr. W2Experience: 3-5 YearsOverviewWe are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps-driven...Remote work
$40 per hour
A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work- OverviewThe Infosys Financial Services unit is a global leader in driving digital transformation for financial institutions. We specialize in leveraging advanced technologies such as AI, cloud, and data-led innovation to help our clients accelerate growth and unlock business...Full timeTemporary workRelocation
- Infosys is seeking a Lead Sterling Integrator Consultant to join our Richardson, TX team. The role focuses on API-led connectivity, integration architecture, and cloud/on-premise server infrastructure, delivering high-quality solutions across global teams. You will work...
- Infosys is seeking a Lead Sterling Integrator Consultant. As a Lead Sterling Integrator Administration Consultant, who understands On Prem and Cloud server infrastructure landscape. You will be working with cross‑functional and global teams and requires strong technical...Immediate start
$115.08k - $218.52k
...part of our flagship in Frisco, Texas, our engineering teams bridge technical rigor with real-... ...solve enterprise challenges.As a Senior Lead, Full-Stack Forward Deployed Engineer,... ...ensuring fast database query execution, reliable data flows, and highly responsive rendering...Minimum wageFull timeTemporary workPart timeWork experience placementLocal areaImmediate startRelocation3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead engineer Dallas, TX
- lead operating engineer Dallas, TX
- lead network engineer Dallas, TX
- lead infrastructure engineer Dallas, TX
- site reliability engineer Dallas, TX
- official site Dallas, TX
- site services specialist Dallas, TX
- construction site safety Dallas, TX
- IT site lead Dallas, TX
- site recruiter Dallas, TX



