Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineering

Truist Inc

Regular or Temporary: Regular

Language Fluency: English (Required)

Work Shift: 1st shift (United States of America)

Please review the following job description:

The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams.

Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime.

The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks.

Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.

ESSENTIAL DUTIES AND RESPONSIBILITIES

Following is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.

  • Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.
  • Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.
  • Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.
  • Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.
  • Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.
  • Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.
  • Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities.
  • Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience.
  • Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision.
Qualifications
Required Qualifications

The requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

  • Bachelor’s degree in Computer Science, Software Engineering, or related field.
  • Minimum of 7 years of professional experience in software development.
  • Deep knowledge of multiple programming languages, software architecture, and design principles.
  • Deep understanding of software development lifecycle, testing, deployment, and security practices.
Preferred Qualifications
  • Advanced degree in Computer Science or related technical discipline.
  • Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
  • Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps.
  • Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.
  • 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
  • Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
  • Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
  • Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
  • Proven leadership in major incident management and cross-team technical coordination.
  • Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
  • Excellent communication skills, including executive-level situational awareness during critical incidents.
  • Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Preferred Qualifications
  • Financial services or regulated industry experience.
  • Experience enabling large-scale SRE transformations or modernization initiatives.
  • Familiarity with chaos engineering, resilience assessments, and service failure modeling.
  • Exposure to hybrid-cloud and multi-cloud operational frameworks.
  • Experience contributing to or leading Center for Enablement functions or Communities of Practice.
Key Responsibilities
Incident & Problem Management Leadership
  • Lead major and high-severity incident response efforts, focusing on diagnosing technical root causes therein, and driving multi-team technical resolution.
  • Drive problem management to closure, ensuring systemic fixes replace recurring operational risks.
  • Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks.
Reliability Engineering & Automation
  • Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience.
  • Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools.
  • Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization.
Observability & Operational Excellence
  • Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
  • Define and standardize enterprise observability practices, dashboards, and KPIs.
  • Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation.
Cross-Functional Leadership & Influence
  • Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
  • Act as a change agent to elevate operational maturity and drive transformative improvements across Wholesale.
  • Lead workshops, maturity assessments, and enablement sessions through the SRE C4E and Communities of Practice.
Standardization & Documentation
  • Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
  • Contribute to enterprise SRE frameworks, templates, and maturity models.
  • Promote consistent adoption of best practices across domains and lines of business.
Mentorship & Technical Development
  • Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
  • Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.

For this opportunity, Truist will not sponsor an applicant for work visa status or employment authorization, nor will we offer any immigration-related support for this position (including, but not limited to H-1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN-1 or TN-2, E-3, O-1, or future sponsorship for U.S. lawful permanent residence status.)

Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.

General Description of Available Benefits for Eligible Employees of Truist Financial Corporation: All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truist’s defined benefit pension plan, restricted stock units, and/or a deferred compensation plan. As you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work.

Truist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace.

#J-18808-Ljbffr
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineering in Atlanta, GA vacancy
  •  ...in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.Site Reliability EngineerOnsite: Atlanta, GAJob SummaryAt NCR Voyix, we're looking for a Site Reliability Engineer II to help build, support, and scale the cloud platforms that power our... 
    Suggested
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    2 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Suggested
    Contract work

    2T Consulting

    Atlanta, GA
    a month ago
  • $100k - $120k

     ...Overview The Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Remote work
    Flexible hours

    Worky Ltd

    Atlanta, GA
    15 hours ago
  •  ...Site Reliability Engineer Full Description Company Overview: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12... 
    Suggested
    Full time
    Live in
    Work at office

    TrulyHired

    Atlanta, GA
    1 day ago
  •  ...focused on building and supporting advanced platforms and applications that drive our business forward. We are seeking a Site Reliability Engineer (SRE) to help define and raise the reliability bar for our Commerce Platform. As an SRE on the AI Commerce team, you will... 
    Suggested
    Fixed term contract

    Acuity Brands

    Atlanta, GA
    14 hours ago
  •  ...new team members who want to be a part of this journey! Who We’re Looking For We’re looking for a proactive, hands‑on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in fast‑moving startup environments. You’re someone who enjoys... 
    Work experience placement
    Flexible hours

    Rainforest

    Atlanta, GA
    15 hours ago
  • $81.75k - $138.98k

     ...Locations 4125 GA hwy 316, Dacula, GA, 30019, US (Hybrid) Job Schedule Full time Job Description As the Senior Site Reliability Engineer, you will serve as a trusted technical resource responsible for deploying, validating, and operationalizing AI, HPC, Kubernetes... 
    Full time
    Work at office
    Immediate start
    Worldwide
    Shift work

    Anixter

    Atlanta, GA
    14 hours ago
  •  ...Job Purpose At Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect companies...  ...-oriented people to join our team. We are seeking a Site Reliability Engineer to bring 3+ years of hands‑on experience to our SRE... 

    Ice Services

    Atlanta, GA
    15 hours ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's...  ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    Atlanta, GA
    a month ago
  • $178.13k - $205.4k

     ...partial telecommuting. Salary Range: $178,131 - $205,400 About You Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5) years of progressive, post-baccalaureate experience in the job... 
    Work at office
    Remote work
    Flexible hours

    HR Tech Job

    Atlanta, GA
    4 days ago
  •  ...data, and human expertise. We deliver faster, smarter, more reliable insights to insurance carriers and single-family rental...  ...at scale, you’re in the right place. The Role As a Site Reliability Engineer, you'll be responsible for the availability, scalability,... 
    Flexible hours

    Seek Now

    Atlanta, GA
    4 days ago
  • $130k - $145k

     ...Back Site Reliability Engineer Cloud/Infrastructure Atlanta , GA Sep 2, 2026 Site Reliability Engineer Atlanta, GA / Hybrid Blu Omega is seeking a Site Reliability Engineer to support a federal program focused on enterprise cloud modernization. This role operates... 
    Temporary work

    Blu Omega LLC

    Atlanta, GA
    4 days ago
  •  ...Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Worldwide

    Inspire Brands Inc

    Atlanta, GA
    4 days ago
  •  ...Join to apply for the Site Reliability Engineer role at Motion Recruitment Join to apply for the Site Reliability Engineer role at Motion Recruitment Get AI-powered advice on this job and more exclusive features. Every year, nearly 200 million travelers... 
    Contract work
    Worldwide

    Motion Recruitment

    Atlanta, GA
    5 days ago
  • $123.4k - $222.53k

     ...Responsibilities Enhance system reliability and resilience by identifying issues and implementing preventive measures to reduce downtime...  ...) ~ Acceptable areas of study include Computer Science, Engineering or related field (Required) ~4-7 years Working in operations... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    4 days ago
  • $141.8k - $195k

     ...best work, grow fast, and bring their full selves to the herd. Why You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl... 
    Temporary work
    Remote work

    Cribl

    Atlanta, GA
    3 days ago
  • $130k - $150k

     ...recruiter to learn more. Base pay range $130,000.00/yr - $150,000.00/yr Overview: We are seeking a highly skilled Site Reliability Engineer (SRE) to join our team and help build and maintain scalable, reliable, and efficient systems. The ideal candidate will... 
    Full time
    Remote work

    Prestige Staffing

    Atlanta, GA
    1 day ago
  • $149.8k - $241.5k

     ...efficiency, accelerate time-to-value, and deliver better customer experiences. About The Role We're looking for a Senior Site Reliability Engineer who's passionate about building reliable, scalable infrastructure that helps developers ship better software faster. You'... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide
    Home office
    Flexible hours

    Camunda

    Atlanta, GA
    3 days ago
  •  ...Job Title :- Site Reliability Engineer (SRE) Employment Type :- W2 Duration :- Long Term Visa Type :- All Visa applicable which are ready for W2 Location :- Atlanta, GA (Onsite) Job Description We are seeking a highly skilled Site Reliability Engineer (SRE... 

    Highbrow

    Atlanta, GA
    3 days ago
  • $120k - $175k

     ...of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's possible... 
    Full time
    Remote work
    Work visa
    Flexible hours

    AEG Presents

    Atlanta, GA
    5 days ago
  •  ...Site Reliability Engineer We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and... 

    Alembic Technologies

    Atlanta, GA
    1 day ago
  • $113.2k - $188.8k

     ...to work on improving Grid resilience through Software, this is the job for you. We are looking for a Deployment and Site Reliability Engineer to join the GridBeats team. The GridBeats Software Portfolio aggregates all Grid Automation applications which monitor, optimize... 
    Permanent employment
    Contract work
    Remote work
    Relocation

    Openkyber

    Atlanta, GA
    2 days ago
  •  ...English (Required) Work Shift: 1st shift (United States of America) Please review the following job description: The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Fayette Chamber of Commerce

    Atlanta, GA
    4 days ago
  • $61.09k - $104.36k

     ...breadth of their business needs. It delivers end-to-end services and solutions leveraging strengths from strategy and design to engineering, all fueled by its market leading capabilities in AI, generative AI, cloud and data, combined with its deep industry expertise and... 
    Full time
    Local area

    Capgemini

    Atlanta, GA
    3 days ago
  •  ...Site Reliability Engineering (SRE) Architect Location: Atlanta, GA Duration: 12Months+ Extension Hourly Rate: Depending on Experience (DOE) Work Authorization: As an SRE Architect, you will be a pivotal technical leader responsible for designing,... 
    Hourly pay
    Permanent employment
    Contract work
    Local area
    Early shift

    ETHEREUM TECHNOLOGIES LLC

    Atlanta, GA
    4 days ago
  •  ...our company effectively.    The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's...  ...infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline... 
    Remote work
    Flexible hours

    Intellum, Inc.

    Atlanta, GA
    a month ago
  •  ...Site Reliability Engineer We are looking for a Site Reliability Engineer to ensure the reliability, security, and continuous operation of a multi‑cloud application security platform. This role combines platform engineering and security automation, focusing on Kubernetes... 
    Work at office
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    AgileEngine

    Atlanta, GA
    14 hours ago
  • $70 - $75 per hour

     ...Site Reliability Engineering (SRE) Architect Get AI-powered advice on this job and more exclusive features. This range is provided by STAFFWORXS. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range... 
    Contract work

    Staffworxs Inc

    Atlanta, GA
    15 hours ago
  • $178k - $213k

     ...Ventures, and Vista Credit Partners of Vista Equity Partners 2022 Cybersecurity Excellence Award for MDR Manager, Site Reliability Engineering Reports to: VP, Product Engineering Location: While proximity to Tampa is preferred to support hybrid schedule in... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Home office
    Flexible hours

    deepwatch

    Atlanta, GA
    15 hours ago
  •  ...Role: Senior Site Reliability Engineer (SRE) Cloud & Kubernetes Location: Atlanta, GA (Onsite) Contract Role Summary: Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, GCP, and Kubernetes... 
    Contract work

    Noblesoft Technologies

    Atlanta, GA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineering. Be the first to apply!