Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Federal Reserve Bank of San Francisco

CompanyFederal Reserve Bank of San Francisco When you join the Federal Reserve-the nation's central bank-you'll play a key role, collaborating with leading tech professionals to strengthen and protect our economic, financial and payments systems. We invest in contemporary and emerging technology each year to support the Federal Reserve and our economy, and we're building a dynamic and diverse team for our future.


The Federal Reserve Financial Services portfolio provides technology capabilities that power the U.S. payment systems infrastructure. This portfolio enables critical payment services that are foundational to the nation's financial system, focusing on delivering speed, resilience, and choice to meet evolving marketplace needs.


We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability, and performance of our critical systems. This role combines deep technical expertise in software engineering, cloud infrastructure, and DevOps practices to ensure our services meet the highest standards of availability and operational excellence.

Responsibilities

System Reliability & Performance

* Design, implement, and maintain highly available, scalable, and resilient systems across cloud infrastructure

* Establish and monitor SLIs, SLOs, and SLAs to ensure optimal system performance

* Lead incident response, conduct root cause analysis, and implement preventive measures

* Develop and maintain disaster recovery and business continuity plans

Infrastructure & Automation

* Architect and manage cloud infrastructure on AWS using Infrastructure as Code (Terraform)

* Automate deployment pipelines, monitoring, and operational workflows

* Optimize cloud resource utilization and cost management

Engineering & Development

* Build and maintain internal tools and services to improve operational efficiency

* Collaborate with development teams to implement reliability best practices

* Conduct code reviews and provide technical guidance on system design

* Develop monitoring solutions, alerting systems, and observability frameworks

Security & Compliance

* Integrate security practices into CI/CD pipelines (SAST/DAST)

* Implement and maintain security controls across infrastructure and applications

* Ensure compliance with industry standards and regulatory requirements

* Conduct security assessments and vulnerability management

Leadership & Collaboration

* Mentor junior SRE team members and promote SRE culture across the organization

* Partner with software engineering teams to improve system reliability

* Drive technical initiatives and contribute to architectural decisions

* Document processes, runbooks, and technical specifications

Software Engineering:

  • Strong proficiency in Java , Python , and Node.js
  • Experience with microservices architecture and distributed systems
  • Solid understanding of data structures, algorithms, and design patterns
  • Proficiency in writing clean, maintainable, and testable code

Cloud Infrastructure (AWS):

  • Extensive experience with AWS services including:
    • Compute: Lambda, ECS, EC2, Fargate
    • Storage: S3, EBS, EFS
    • Database: RDS, DynamoDB, Aurora
    • Networking: VPC, Route53, CloudFront, API Gateway
    • Monitoring: CloudWatch, X-Ray
  • AWS certifications (Solutions Architect, DevOps Engineer) preferred

DevOps & CI/CD:

  • Expert-level knowledge of GitLab (CI/CD pipelines, runners, GitOps)
  • Advanced Terraform skills for infrastructure provisioning and management
  • Experience with containerization (Docker) and orchestration (Kubernetes/ECS)
  • Proficiency with configuration management tools

Security:

  • Hands-on experience with SAST (Static Application Security Testing) tools
  • Knowledge of DAST (Dynamic Application Security Testing) methodologies
  • Understanding of security best practices, OWASP Top 10, and compliance frameworks
  • Experience with secrets management and identity access management (IAM)

Monitoring & Observability:

  • Experience with monitoring tools (Grafana, Datadog, New Relic, or similar)
  • Log aggregation and analysis (CloudWatch Logs, Splunk)
  • Distributed tracing with aws X-Ray

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience
  • 7+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years in a lead or senior technical position
  • Proven track record of managing large-scale production systems
  • Experience with on-call rotations and incident management
  • GenAI based Applications: Working knowledge of LLMs and agentic applications a plus
  • Experience with serverless architectures and event-driven systems
  • Familiarity with chaos engineering principles and practices
  • Background in Agile/Scrum methodologies
  • Experience with multi-cloud or hybrid cloud environments

The selected candidate will reside within a reasonable commuting distance, as defined by the employing Reserve Bank, and will work full-time onsite.

  • Eligible Locations for Hire: Boston, MA- New York, NY- Philadelphia, PA- Cleveland, OH- Richmond, VA- Atlanta, GA- Chicago, IL- St. Louis, MO- Minneapolis, MN- Kansas City, MO- Dallas, TX- San Francisco, CA
  • The following Reserve Bank locations are preferred due to the concentration of System IT team members in these locations: San Francisco, and Richmond, VA

Screening:

Due to the nature of access to sensitive information all final offers are subject to the clearance of an enhanced background check. This enhanced screening will require the following items: academic and employment verifications, FBI fingerprint check (criminal and civil cases), credit check, family history, residential records and foreign travel for the previous 7 years, citizenship verification, reference checks, and personal interview with an investigator and can take between 21 - 60 days to clear.

Sponsorship:
Individuals who need immigration sponsorship now or in the future are not eligible for this position.
Must be a U.S Citizen or a Green card holder with intent to become a U.S Citizen.

Base Salary Range: Min: $146,700Mid: $190,500Max: $234,300 (Location: San Francisco)

The listed salary is applicable to 12th District/San Francisco. Final offers are determined by factors including the candidate's qualifications, internal alignment considerations, district assignment, and geographic location.

The Bank is committed to providing reasonable accommodations to individuals with disabilities to participate in the job application or interview process, perform essential job functions and receive other benefits and privileges of employment. The SF Fed is an Equal Opportunity Employer. If you need any assistance or accommodations due to a disability, please let us know at View email address on click.appcast.io .

Full Time / Part TimeFull time Regular / TemporaryRegular Job Exempt (Yes / No)Yes Job CategoryInformation Technology Family Group Work ShiftFirst (United States of America)

The Federal Reserve Banks are committed to equal employment opportunity for employees and job applicants in compliance with applicable law and to an environment where employees are valued for their differences.

Always verify and apply to jobs on Federal Reserve System Careers ( or through verified Federal Reserve Bank social media channels.

Privacy Notice

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Dallas, TX vacancy
  • Qualifications: 8+ years of Software Engineering experience, or equivalent demonstrated through...  ...implement and maintain scalable and reliable infrastructure on Google Cloud Platform...  ...vendor resources Willingness to work on-site at stated location in the job openingDepartment... 
    Suggested
    Contract work
    For contractors
    Work experience placement

    Cedent Consulting

    Dallas, TX
    23 hours ago
  • $255.7k - $300k

     ...designs from peers, providing feedback to ensure best practices in reliability, security, and efficiency.Triage and resolve complex system...  ...execution of software development initiatives.Mentor other engineers and contribute to the engineering community through documentation... 
    Suggested
    Full time

    Google

    Sunnyvale, TX
    2 days ago
  •  ...cloud-native platforms to advanced release engineering practices, our teams are redefining how...  ...in coding, testing, and automation. Reliability Engineering: Establish service level...  ...Scrum teams with demonstrated success leading improvements (getting better/faster/happier... 
    Suggested
    H1b
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    2 days per week
    3 days per week

    GM Financial

    Irving, TX
    23 hours ago
  •  ...part of our global expansion, we're looking for a hands-on Site Reliability Engineer (SRE) to design, scale, and safeguard the reliability of our...  ...reliability and resiliency are built in from day one. Lead incident response: Drive on-call processes, conduct root-... 
    Suggested
    Remote work

    Longbridge Singapore

    Dallas, TX
    23 hours ago
  • $140k - $150k

     ...learn more. Base pay range $140,000.00/yr - $150,000.00/yr Site Reliability Engineer II | 6-month Contract to Hire | Hybrid (Irving, TX) | 2x onsite per week Optomi, in partnership with a leading financial services company, is seeking a highly skilled, hands‑on... 
    Suggested
    Full time
    Contract work

    Optomi

    Irving, TX
    23 hours ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Full time
    Flexible hours

    Proofpoint

    Dallas, TX
    1 day ago
  • $55k - $151.47k

     ...LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in...  ...architecture to support data integrity and accessibility- Leading incident management and resolution efforts to maintain operational... 
    Full time
    H1b

    PwC

    Dallas, TX
    15 hours ago
  • $192.4k - $275.8k

     ...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines...  ...facing, and senior leadership audiences 4+ years experience leading post-mortems and root cause analysis for high-severity... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Richardson, TX
    1 day ago
  • $172k - $300k

    Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property...  ...Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    3 days ago
  • $72.1k - $158.62k

     ...person, one family and one community at a time. Position Summary We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence. The... 
    Hourly pay
    Full time
    Temporary work
    Local area

    CVS Health

    Richardson, TX
    3 days ago
  •  ...applications, databases, etc. # Set up SLOs and SLIs using industry-leading tools. # Play the role of an individual contributor and lead...  ...self-healing solutions. # Experience in Implementing Chaos Engineering/testing. Seniority level Mid-Senior level... 
    Full time

    Infosys

    Richardson, TX
    23 hours ago
  • Mandatory Skills: AWS/Azure/GCP (GCP is not used very much ). Kubernetes /Helm,Docker,Gitlab,Grafana,Cyberark/Hashicorp Vault, Terraform etc. Experience utilizing Java, Perl, Python, Go and scripting experience in Shell and Perl to automate reports and monitor enterprise...

    Omni Inclusive

    Dallas, TX
    2 days ago
  •  ...healthcare fintech innovator, we’re transforming the patient journey and redefining what’s possible in dental care. This role: Site Reliability Engineer (SRE) with deep expertise in monitoring, debugging, and optimizing Azure App Services. This position is critical to... 
    Full time
    Work at office
    3 days per week

    Wellfit Technologies

    Irving, TX
    23 hours ago
  •  ...improving platform infrastructure and applications with high reliability, resiliency, performance & quality, and faster time-to-market...  ...documentation, including runbooks/playbooks; and, Using Chaos Engineering to test the robustness of the systems and applications.... 

    Software Technology Inc

    Dallas, TX
    4 days ago
  •  ...confidential information into the tool.Interactions with the tool are reviewed in order to improve results.If you do not agree with any part of this notice, please close the tool and use the career site. For questions or feedback regarding the tool, submit an HR Connect ticket.

    Texas Instruments

    Dallas, TX
    1 day ago
  • $119k - $170k

     ...impact at the company pioneering security transformation in the AI era? Join us at Zscaler.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (onsite three days a week in San Jose, CA or another Zscaler... 
    Full time
    Work at office
    Local area
    Remote work
    Shift work
    3 days per week

    Zscaler

    Dallas, TX
    4 days ago
  •  ...Senior Site Reliability Engineer (Permanent Role) Cleveland, OH, Pittsburgh, PA, or Dallas, TX Your future duties and responsibilities...  .... Facilitating analysis meetings to discuss incidents. Lead . Identify the automation opportunities for automation specialists... 
    Permanent employment
    Temporary work
    Local area
    Flexible hours
    Shift work
    Weekend work

    System One

    Dallas, TX
    a month ago
  • Company DescriptionAmerica Networks is a leading sensor and networking solutions partner for companies in any Industrial, Manufacturing...  ...asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing,... 

    AmNet Services

    Irving, TX
    23 hours ago
  • $160k - $225k

     ...experience solutions. Our partnerships with leading cloud, design and business intelligence...  ...needed on critical paths, establishing engineering guardrails, and leading design reviews....  ...and guide tradeoffs across reliability, performance, and delivery speedHands on... 
    Permanent employment
    Full time
    Temporary work
    Remote work

    TEKsystems

    Dallas, TX
    3 days ago
  •  ...stores and communities every day. If you're ready to grow, lead and make a difference, come join our team and help shape the future of convenience.The SRE RunOps Engineer 2 is responsible for ensuring the reliability, availability, and performance of the 7NOW delivery... 
    Hourly pay
    Work experience placement

    7 eleven

    Irving, TX
    1 day ago
  •  ...fulfill travel worldwide.SRE Software Systems Engineer IV - Data Intelligence and AI...  ...Systems Engineer, you will drive platform reliability, auto-scaling and cloud cost efficiency...  ...running smoothly. This role requires strong Site Reliability Engineering discipline, problem... 
    Full time
    Worldwide
    Flexible hours
    Weekend work

    Sabre Holdings

    Dallas, TX
    1 day ago
  •  ...874863Reference Number: 25-00760Title: AWS Python ML Developer - Lead LevelPosted Date: 2025-07-10Company: HAN StaffingRole: AWS Python...  ...interviewRound 2: Behavioral or combined technical/behavioralA final on-site interview may be required for top candidates to validate skills... 
    Full time
    Relocation
    3 days per week

    HAN Staffing

    Dallas, TX
    1 day ago
  • Smart Tech Contracting LLC is seeking a Lead System Integrator - Critical Infrastructure to drive engagement in large-scale BAS/EPMS...  ...commissioning. You will lead teams, coordinate with design engineers, contractors, vendors and clients to ensure successful delivery... 
    Remote job
    For contractors

    SmartTech Contracting LLC

    Dallas, TX
    14 hours ago
  • Smart Tech Contracting, LLC is seeking a Lead System Integrator - Critical Infrastructure to guide large-scale BAS/EPMS projects through...  ...supervise a team of system integrators, coordinate with design engineers and vendors, and ensure the system meets contract requirements.... 
    Contract work

    Smart Tech Contracting, LLC

    Dallas, TX
    14 hours ago
  •  ...ContractPay Rate: $40/Hr. W2Experience: 3-5 YearsOverviewWe are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in a DevOps-driven... 
    Remote work

    BayOne Solutions

    Richardson, TX
    2 days ago
  • $40 per hour

    A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants... 
    Long term contract
    Internship
    Remote work

    BayOne Solutions

    Richardson, TX
    2 days ago
  • OverviewThe Infosys Financial Services unit is a global leader in driving digital transformation for financial institutions. We specialize in leveraging advanced technologies such as AI, cloud, and data-led innovation to help our clients accelerate growth and unlock business...
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Dallas, TX
    2 days ago
  • Infosys is seeking a Lead Sterling Integrator Consultant to join our Richardson, TX team. The role focuses on API-led connectivity, integration architecture, and cloud/on-premise server infrastructure, delivering high-quality solutions across global teams. You will work... 

    Infosys

    Richardson, TX
    14 hours ago
  • Infosys is seeking a Lead Sterling Integrator Consultant. As a Lead Sterling Integrator Administration Consultant, who understands On Prem and Cloud server infrastructure landscape. You will be working with cross‑functional and global teams and requires strong technical... 
    Immediate start

    Infosys

    Richardson, TX
    13 hours ago
  • $115.08k - $218.52k

     ...part of our flagship in Frisco, Texas, our engineering teams bridge technical rigor with real-...  ...solve enterprise challenges.As a Senior Lead, Full-Stack Forward Deployed Engineer,...  ...ensuring fast database query execution, reliable data flows, and highly responsive rendering... 
    Minimum wage
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Immediate start
    Relocation
    3 days per week

    Kyndryl

    Dallas, TX
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!