Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering

SunTrust Investment Services, Inc.

Site Reliability Engineering Lead

The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams. Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime. The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks. Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.

Essential Duties And Responsibilities Following is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.

  • Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.
  • Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.
  • Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
  • Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.
  • Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.
  • Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.
  • Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities.
  • Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience.
  • Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision.

Qualifications Required Qualifications The requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

  • Bachelor's degree in Computer Science, Software Engineering, or related field.
  • Minimum of 7 years of professional experience in software development.
  • Deep knowledge of multiple programming languages, software architecture, and design principles.
  • Deep understanding of software development lifecycle, testing, deployment, and security practices.

Preferred Qualifications

  • Advanced degree in Computer Science or related technical discipline.
  • Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
  • Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps.
  • Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
  • Deep hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
  • Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
  • Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
  • Proven leadership in major incident management and cross-team technical coordination.
  • Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
  • Excellent communication skills, including executive-level situational awareness during critical incidents.
  • Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.

Preferred Qualifications

  • Financial services or regulated industry experience.
  • Experience enabling large-scale SRE transformations or modernization initiatives.
  • Familiarity with chaos engineering, resilience assessments, and service failure modeling.
  • Exposure to hybrid-cloud and multi-cloud operational frameworks.
  • Experience contributing to or leading Center for Enablement functions or Communities of Practice.

Key Responsibilities

  • Incident & Problem Management Leadership
    • Lead major and high-severity incident response efforts, focusing on diagnosing technical root causes therein, and driving multi-team technical resolution.
    • Drive problem management to closure, ensuring systemic fixes replace recurring operational risks.
    • Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks.
  • Reliability Engineering & Automation
    • Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience.
    • Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
    • Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
  • Observability & Operational Excellence
    • Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
    • Define and standardize enterprise observability practices, dashboards, and KPIs.
    • Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
  • Cross-Functional Leadership & Influence
    • Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
    • Act as a change agent to elevate operational maturity and drive transformative improvements across Wholesale.
    • Lead workshops, maturity assessments, and enablement sessions through the SRE C4E and Communities of Practice.
  • Standardization & Documentation
    • Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
    • Contribute to enterprise SRE frameworks, templates, and maturity models.
    • Promote consistent adoption of best practices across domains and lines of business.
  • Mentorship & Technical Development
    • Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
    • Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.

Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering in Atlanta, GA vacancy
  • $35 - $45 per hour

    DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'... 
    Suggested
    Remote work

    KForce

    Atlanta, GA
    3 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    2 days ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Suggested
    Worldwide

    Inspire Brands

    Atlanta, GA
    2 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Atlanta, GA
    2 days ago
  •  ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer... 
    Suggested
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    2 days ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Atlanta, GA
    1 day ago
  •  ...We are currently looking for a Senior Software Engineer to be a part of the Site Reliability Engineering (SRE) team in Atlanta, GA . The SRE team is an innovative team devoted to providing a Docker-based Platform as a Service and assisting a growing number of teams... 
    Contract work
    Work at office
    Local area

    Spartan Technologies

    Atlanta, GA
    2 days ago
  •  ...provisioning, monitoring, and troubleshooting staging and production cloud environments . Experienced in architectural design for reliability, scalability, and performance. Practical application of SRE principles : SLIs, SLOs, error budgets, automation, incident... 

    Purple Drive

    Atlanta, GA
    4 days ago
  •  ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively... 
    Work experience placement

    Fisec Global

    Atlanta, GA
    21 hours ago
  •  ...ideal time and number for communication, and the expected pay rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months Contract Visa : US Citizens/ Green Card Need Local to... 
    Contract work
    Local area
    Immediate start

    Navtech

    Atlanta, GA
    2 days ago
  • $120k - $175k

     ...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues... 
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    Atlanta, GA
    4 days ago
  •  ...Site Reliability Engineer Opportunity Rainforest is an early stage payments-as-a-service startup that has developed a solution that makes monetizing payments for vertically focused software platforms fair and simple. We focus on small-to-mid sized platforms that want... 
    Work experience placement
    Flexible hours

    Rainforest

    Atlanta, GA
    21 hours ago
  • $74.1k - $148.3k

     ...systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining, designing, deploying, and solving key Oracle Cloud services,... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle Defunct

    Atlanta, GA
    4 days ago
  •  ...Senior Site Reliability Engineer Atlanta, Georgia Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered. We're on a mission to empower the healthcare industry to better onboarding, deploy, and manage their workforce. Over... 
    Permanent employment
    Full time
    Work at office
    Remote work
    Work from home
    Work visa

    QGenda

    Atlanta, GA
    3 days ago
  •  ...Site Reliability Engineer (SRE) When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage... 
    Work at office

    Atlanticus

    Atlanta, GA
    21 hours ago
  • $60 - $68 per hour

     ...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested... 
    Contract work
    Local area
    Immediate start

    Pyramid Consulting

    Atlanta, GA
    21 hours ago
  • $104.9k - $174.7k

     ...58 Are you passionate about improving reliability, scalability, and resilience in complex...  ...practices. Own prioritization of reliability engineering tasks within team backlogs. Lead...  ...(IaaS). Background in DevOps, site reliability engineering practices, or related... 
    Full time
    Local area

    RELX

    Atlanta, GA
    2 days ago
  •  ...Responsibilities Kforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot... 
    Hourly pay
    Contract work
    Remote work

    Kforce

    Atlanta, GA
    2 days ago
  • $71.6k - $119.4k

     ...support application teams. Our services provide applications with reliability, security, and better customer experiences. About the Job:...  ...automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You'll gain exposure to a... 
    Full time
    Temporary work
    Internship
    Local area
    Work from home

    RELX

    Atlanta, GA
    2 days ago
  • $121.4k - $218.6k

     ...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with...  ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust... 
    Work experience placement
    Work at office

    Akamai

    Atlanta, GA
    1 day ago
  •  ...resilient software platforms using SRE and AI-native engineering practices. Own production reliability, monitoring, and operational automation while mentoring...  ...), Kubernetes, and production operations. Key Skills Site Reliability Engineering Terraform Python AWS Azure... 
    Temporary work
    Flexible hours

    T-Mobile

    Atlanta, GA
    4 days ago
  • Kforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and optimizing... 
    Temporary work
    Remote work
    Atlanta, GA
    1 day ago
  •  ...English (Required) Work Shift: 1st shift (United States of America) Please review the following job description: The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist Inc

    Atlanta, GA
    2 days ago
  •  ...Role: Site Reliability Engineering (SRE) Architect Location: Atlanta, GA (Hybrid on-site) Contract Role Summary: As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems... 
    Contract work
    Early shift

    AceStack LLC

    Atlanta, GA
    2 days ago
  • Direct message the job poster from STAFFWORXS Delivery Manager @ STAFFWORXS | US IT Recruitment Job Opening: AWS Site Reliability Engineer (SRE) We’re hiring a Site Reliability Engineer (SRE) to join our team in Atlanta, GA. This hybrid role offers the opportunity to work... 
    Contract work

    STAFFWORXS

    Atlanta, GA
    2 days ago
  • $151k - $297k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    2 days ago
  • #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...

    LTM

    Atlanta, GA
    21 hours ago
  •  ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it... 
    Full time
    Live in
    Work at office

    Incident IQ

    Atlanta, GA
    a month ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's...  ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    Atlanta, GA
    26 days ago
  •  ...build, and test of the company's first combined turbojet-ramjet engine and is now being scaled through its first flight vehicle...  ...capabilities to the warfighter. Hermeus is seeking a Senior Software or Site Reliability Engineer to join the Information Team and take charge of... 
    Full time
    Remote work

    Hermeus

    Atlanta, GA
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering. Be the first to apply!