Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Glint Tech Solutions LLC

Job Title: Lead Site Reliability Engineer


Location: Remote within the USA, or onsite in Buffalo, NY / Wilmington, DE (client preference for candidates near these areas). New hires are required to work onsite at the client's office for the first 2–3 weeks (treated as a business trip; travel expenses covered by the company).

Company Overview
Glint Tech Solutions is a women-owned, global IT staffing and recruiting firm serving enterprise clients across the USA and Canada.

Project Description
A leading financial services client is seeking a Lead Site Reliability Engineer responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. This senior individual contributor will design, implement, and improve SRE practices across the software development lifecycle, working closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management, while coaching and influencing others.

Key Responsibilities

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise SRE best practices
  • Define, implement, and monitor SLOs, SLIs, and error budgets for critical business services
  • Develop observability strategies using Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints
  • Lead incident response for high-severity production events and facilitate Root Cause Analysis (RCA)
  • Drive operational excellence through automation of deployments, recovery procedures, and reliability controls
  • Design and execute automated regression testing strategies to validate stability and performance
  • Create and maintain Infrastructure as Code (IaC) solutions using Terraform
  • Support and optimize Microsoft Azure environments, including App Services, scaling, and deployment automation
  • Utilize Azure Monitor, Application Insights, and Log Analytics to improve platform visibility
  • Drive performance testing, resiliency testing, and disaster recovery preparedness
  • Lead capacity planning, performance tuning, and workload optimization
  • Develop operational runbooks, incident playbooks, and standard operating procedures
  • Mentor engineers on observability, cloud engineering, automation, and SRE principles
  • Adhere to Company risk and regulatory standards, policies, and controls

Mandatory Skills

  • Strong hands-on experience with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and centralized logging
  • Proven experience designing and executing automated regression testing frameworks
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform
  • Experience with CI/CD pipelines, deployment automation, and operational tooling
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting
  • Strong understanding of application performance management, distributed systems, and cloud-native architectures
  • Strong experience with Microsoft Azure (App Services, Resource Groups, networking, scaling, deployment/release management)
  • Experience with Azure Monitor, Application Insights, Log Analytics, and Azure dashboards/alerting
  • Experience supporting cloud-native and hybrid infrastructure environments
  • Demonstrated experience implementing SRE practices — SLOs, SLIs, error budgets, incident/problem management, RCA, reliability automation
  • Ability to improve system reliability through performance tuning, capacity planning, and observability-driven insights
  • Experience developing automated recovery mechanisms and self-healing solutions
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures

Nice-to-Have Skills

  • Experience supporting large-scale enterprise applications in regulated environments
  • Experience working in Agile and DevOps operating models
  • Ability to work autonomously and lead complex reliability initiatives
  • Experience partnering with architecture, infrastructure, cybersecurity, and application development teams
  • Scripting/automation experience with PowerShell, Python, or Bash
  • Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering
  • Proven experience leading major incident response and post-incident improvement efforts
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in United States vacancy
  • $90k - $130k

     ...Credence has an immediate opening for a Site Reliability SME who has hands-on experience working as a Cloud Operations Engineer with experience in IT operations to join our...  ...like the Operations Manager/TOPM, technical leads, and environmental engineers to prioritize,... 
    Suggested
    Temporary work
    Work experience placement
    Immediate start
    Worldwide

    Credence

    McLean, VA
    11 days ago
  •  ...build a successful career with opportunities to learn, grow, and make an impact. Join us! Position Summary: The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the... 
    Suggested
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America Corporation

    Charlotte, NC
    a month ago
  • $146.4k - $263.6k

     ...enjoy working with a diverse multi-national team of engineering talents? Join our highly skilled Site Reliability team Our Hardware, Infrastructure, and...  ...datacenters. As a Senior II Site Reliability Engineer Lead, you will be responsible for: Architecting, developing... 
    Suggested
    Full time
    Work experience placement
    Work at office

    Akamai

    Remote
    3 days ago
  • $99k - $225k

    Site Reliability Engineer, LeadThe Opportunity:  As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Chantilly, Loudoun County, VA
    1 day ago
  • $100.1k - $180.2k

     ...'s top brands, offering comprehensive engineering, supply chain, and manufacturing solutions...  ...and a vast network of over 100 sites worldwide, Jabil combines global reach...  ...communities around the globe.Jabil is seeking a Lead Site Reliability Infrastructure and Security Engineer... 
    Suggested
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide

    Jabil Circuit

    Austin, TX
    1 day ago
  • $152k - $195k

     ...and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and...  ...observability — define SLOs, alerts, and dashboards. Lead incident response and postmortems, focusing on root cause and... 
    Remote work

    SecurityScorecard

    United States
    2 days ago
  • $182.8k - $247.3k

     ...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems... 
    Work experience placement

    Socket

    Eastern, KY
    18 hours ago
  •  ...Senior Site Reliability Engineer Remote – Home Based Job Summary We’re partnering with a company in the SaaS space to find a Senior Site Reliability Engineer . In this role, you’ll be part of the IT Operations group responsible for maintaining all environments... 
    Temporary work
    Remote work
    Work from home
    Flexible hours

    SourceDirect Talent

    United States
    3 days ago
  • $125k - $250k

     ...reimagining how developers build reliable, scalable, event-driven...  ...budgets across the platform Lead incident response efforts...  ...possible Partner closely with engineering teams to improve system...  ...5+ years of experience in Site Reliability Engineering, DevOps... 
    Full time
    Immediate start
    Remote work
    Flexible hours

    Orkes

    United States
    1 day ago
  •  ...About the Role We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability...  ...; ensure we meet or exceed targets consistently ~ Lead observability strategy by designing comprehensive... 
    Remote work

    MeridianLink

    United States
    1 day ago
  • $180k - $230k

     ...About Us GridCARE is a leading venture-backed startup solving the most critical constraint...  ...re looking for a Senior SRE to own the reliability, scalability, and observability of our...  ...ll work closely with platform and data engineering to keep high-throughput, data-intensive... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Eastern, KY
    2 days ago
  • $186.82k - $224.18k

     ...Who We Are Babylist is the leading registry, e-commerce, and content platform...  ...tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI...  ...looking for a Senior Software Engineer, Site Reliability to join our Platform team. In this position... 
    Work at office
    Local area
    Immediate start
    Remote work
    Flexible hours
    Shift work

    Babylist

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer Company: Sphera Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Terraform, ARM templates, Kubernetes, Azure, SonarCloud, CheckPoint, Hadoop, Kafka, Presto, NewRelic, CI/CD, Linux, Windows, Redis... 
    Full time
    Remote work

    Sphera

    United States
    1 day ago
  • $140k - $180k

     ...About Us UJET leads the way in AI-powered contact center innovation, delivering a future-proof, cloud platform that redefines...  ...more at Opportunity We’re looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You’ll... 
    Work experience placement
    Local area
    Remote work
    Visa sponsorship
    Work visa

    UJET

    United States
    2 days ago
  •  ...Site Reliability Engineer Company: GitLab Work Type: Remote Employment: Full Time Location: CA, US Seniority: Senior Level Technologies: Terraform, Ansible, Kubernetes, Go, Ruby, Jsonnet, Prometheus, ELK, Grafana Requirements: Senior-level SRE with strong Terraform/IaC... 
    Full time
    Remote work

    GitLab

    United States
    1 day ago
  •  ...customers rely on us in the moments that matter. Engineering delivers on that promise.   The Senior Site Reliability Engineer is responsible for ensuring our SaaS...  ...•    Participate in on-call duties 365/24/7 and lead the triage and RCA of production incidents... 
    Work experience placement
    Remote work
    Flexible hours

    Donnelley Financial Solutions

    United States
    18 hours ago
  •  ...encourage you to apply. The Role  As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry...  ...met. \n What You Will Be Doing Improving production reliability and system resilience within an SRE scoped team Championing... 
    Remote work
    Flexible hours

    Megaport

    United States
    18 hours ago
  •  ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers...  ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    Eastern, KY
    4 days ago
  • $141.8k - $195k

     ...We’re one of the fastest‑growing private companies and a leading player in a massive, fast‑moving market. With a global workforce...  ...You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all... 
    Temporary work
    Remote work

    Cribl

    United States
    1 day ago
  • $7.5k

     ...manager, and we have ambitious goals for the future. As a Site Reliability Engineer (SRE), you will work at the intersection of production...  ...and trading systems Diagnose and fix bugs in code Lead complex deployments Automate manual workflows... 
    Local area
    Remote work

    The Voleon Group

    United States
    1 day ago
  •  ...Senior Site Reliability Engineer Company: CyberArk Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: AWS,...  ...Requirements: Senior SRE with 5+ years AWS infra, 3+ years in senior/lead roles; strong automation with Terraform, Ansible,... 
    Full time
    Remote work

    CyberArk

    United States
    1 day ago
  •  ...Job Title:  Site Reliability Engineer (Azure Government & Infrastructure) Pay Type : SALARIED EXEMPT  Location:  Remote Citizenship Requirement: U.S. Citizen (Required) Summary of Position Role/Responsibilities The Site Reliability Engineer (SRE) for... 
    Full time
    Remote work
    Monday to Friday

    Quzara LLC

    United States
    1 day ago
  •  ...Site Reliability Engineer OXIO is the first NeoTelco. We arebuilding the world’s largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network... 
    Remote work

    OXIO

    United States
    1 day ago
  •  ...Site Reliability Engineer Company: Milestone Systems Work Type: Remote Employment: Full Time Location: US Seniority: Senior Level Technologies: Golang, Python, Linux, Shell scripting, Kubernetes, Docker, Terraform, CI/CD, GitOps, ArgoCD, Spinnaker, Prometheus, Datadog,... 
    Full time
    Remote work

    Milestone Systems Inc

    United States
    1 day ago
  • $135k - $170k

     ...Symmetrio is recruiting a Site Reliability Engineer for its customer, a rapidly growing international healthcare SaaS company aggressively expanding...  ...end. Respond to customer-facing connectivity incidents: lead the call, keep the customer updated, and work the incident... 
    Full time
    Remote work

    Symmetrio

    United States
    2 days ago
  • $147k - $168k

     ...Inc. as one of the most innovative and fastest-growing technology companies in the country. Role Summary As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and operational maturity of the Filevine platform. You’ll... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    2 days per week
    3 days per week

    Filevine

    United States
    1 day ago
  •  ...have come to expect, and help raise the reliability bar as we grow. What you would do:...  ...operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes...  ...technologies. There are many roads leading up to being an SRE. Our team is already... 
    Remote work
    Worldwide
    Flexible hours

    Sanity

    United States
    1 day ago
  •  ...GiveCampus is the world's leading fundraising platform for non-profit educational institutions. Trusted by millions of donors...  ...About the role GiveCampus is looking for a hands-on Site Reliability Engineer to help improve the reliability, performance, and operational... 
    Work at office
    Local area
    Remote work
    Flexible hours

    GiveCampus

    United States
    18 hours ago
  •  ...Site Reliability Engineer Company: Quzara Work Type: Remote Employment: Full Time Location: US Seniority: Mid Level Technologies: Azure, Terraform, Bicep, Ansible, Azure Monitor, Azure Automation, Azure Policy, Azure Site Recovery, TLS/SSL Requirements: 4+ years in SRE... 
    Full time
    Remote work

    Quzara LLC

    United States
    1 day ago
  • $160k - $180k

     ...big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our infrastructural... 
    Local area
    Remote work

    Arkestro

    United States
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!