Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Truist

The position is described below. If you want to apply, click the Apply Now button at the top or bottom of this page. After you click Apply Now and complete your application, you'll be invited to create a profile, which will let you see your application status and any communications. If you already have a profile with us, you can log in to check status.Need Help?If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).Regular or Temporary:RegularLanguage Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams.Responsibilities include leading major incident responses, driving problem management, and implementing automation to reduce service downtime.The role involves standardizing observability practices, mentoring SRE team members, and contributing to enterprise-wide reliability frameworks.Candidates require 7+ years of experience, expertise in distributed systems, Kubernetes, automation scripting, and strong leadership in incident management.ESSENTIAL DUTIES AND RESPONSIBILITIESFollowing is a summary of the essential functions for this job. Other duties may be performed, both major and minor, which are not mentioned below. Specific activities may change from time to time.1. Implements software architecture and engineering approaches for complex initiatives within the job area, contributing to technical plans and working to achieve operational targets with major impact on results.2. Adopts and refines advanced software engineering standards, practices, and governance mechanisms for the job area, influencing how multiple teams improve quality, reliability, and delivery.3. Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches for the domain, providing deep technical insight and proposing solution patterns that inform local roadmaps and priorities.4. Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions for the job area, producing patterns and examples that other technical professionals can follow.5. Independently troubleshoots and resolves complex technical issues in the area of responsibility, designing innovative architectures and performance, reliability, and scalability improvements that advance business objectives.6. Provides ongoing technical guidance, coaching, and training to other engineers, delegating and reviewing work from lower-level technical professionals and raising the technical bar through design reviews and knowledge sharing.7. Evaluates emerging technologies and techniques relevant to the job area, building prototypes and solution concepts that contribute measurable input into new features, products, or capabilities.8. Contributes to the development of long-term technical goals and plans for the area of responsibility through well-reasoned recommendations, design proposals, and implementation experience.9. Leads large or complex initiatives within the job area, coordinating and delegating technical work that may span outside the immediate team, and ensuring cohesive, high-quality outcomes with limited supervision.QualificationsRequired QualificationsThe requirements listed below are representative of the knowledge, skill and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.1. Bachelor’s degree in Computer Science, Software Engineering, or related field.2. Minimum of 7 years of professional experience in software development.3. Deep knowledge of multiple programming languages, software architecture, and design principles.4. Deep understanding of software development lifecycle, testing, deployment, and security practices.Preferred Qualifications1. Advanced degree in Computer Science or related technical discipline.2. Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.3. Deep expertise in cloud-native architectures, microservices, container orchestration, and DevOps.4. Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.5. 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations. 6. Deep hands‑on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling. 7. Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible). 8. Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring. 9. Proven leadership in major incident management and cross-team technical coordination. 10. Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns. 11. Excellent communication skills, including executive-level situational awareness during critical incidents. 12. Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices. Preferred Qualifications Financial services or regulated industry experience. Experience enabling large-scale SRE transformations or modernization initiatives. Familiarity with chaos engineering, resilience assessments, and service failure modeling. Exposure to hybrid-cloud and multi-cloud operational frameworks. Experience contributing to or leading Center for Enablement functions or Communities of Practice. Key Responsibilities Incident & Problem Management Leadership Lead major and high-severity incident response efforts, focusing on diagnosing technical root causes therein, and driving multi-team technical resolution. Drive problem management to closure, ensuring systemic fixes replace recurring operational risks. Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks. Reliability Engineering & Automation Architect and deliver automation solutions that eliminate toil, reduce MTTR, and increase service resilience. Implement intelligent alerting, anomaly detection, and event correlation leveraging AI and AIOps tools. Guide and enforce SLO/SLI adoption across product teams, ensuring metrics inform decision-making and prioritization. Observability & Operational Excellence Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk. Define and standardize enterprise observability practices, dashboards, and KPIs. Ensure operational readiness of applications and platforms through resiliency testing, chaos engineering, and failure-mode validation. Cross-Functional Leadership & Influence Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution. Act as a change agent to elevate operational maturity and drive transformative improvements across Wholesale. Lead workshops, maturity assessments, and enablement sessions through the SRE C4E and Communities of Practice. Standardization & Documentation Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns. Contribute to enterprise SRE frameworks, templates, and maturity models. Promote consistent adoption of best practices across domains and lines of business. Mentorship & Technical Development Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline. Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering. For this opportunity, Truist will not sponsor an applicant for work visa status or employment authorization, nor will we offer any immigration-related support for this position (including, but not limited to H-1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN-1 or TN-2, E-3, O-1, or future sponsorship for U.S. lawful permanent residence status.)Candidate must be willing to work onsite Monday - Friday at either office in Charlotte NC, Raleigh NC, or Atlanta, GA.General Description of Available Benefits for Eligible Employees of Truist Financial Corporation: All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits, though eligibility for specific benefits may be determined by the division of Truist offering the position. Truist offers medical, dental, vision, life insurance, disability, accidental death and dismemberment, tax-preferred savings accounts, and a 401k plan to teammates. Teammates also receive no less than 10 days of vacation (prorated based on date of hire and by full-time or part-time status) during their first year of employment, along with 10 sick days (also prorated), and paid holidays. For more details on Truist’s generous benefit plans, please visit our Benefits site. Depending on the position and division, this job may also be eligible for Truist’s defined benefit pension plan, restricted stock units, and/or a deferred compensation plan. As you advance through the hiring process, you will also learn more about the specific benefits available for any non-temporary position for which you apply, based on full-time or part-time status, position, and division of work.Truist is an Equal Opportunity Employer that does not discriminate on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status, or other classification protected by law. Truist is a Drug Free Workplace.EEO is the LawE-VerifyIER Right to WorkJob SummaryJob number: R0117833Profession: Technology

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Atlanta, GA vacancy
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises... 
    Suggested
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Atlanta, GA
    5 days ago
  • OverviewJob PurposeAt Intercontinental Exchange (NYSE:ICE), we engineer technology, exchanges and clearing houses that connect...  ...motivated, results-oriented people to join our team.We are seeking a Site Reliability Engineer II to bring 3+ years of hands-on experience to our... 
    Suggested

    Black Knight Financial Services

    Atlanta, GA
    2 days ago
  • $98k - $148.5k

     ...Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office,you'll help build and operate the foundational infrastructure that... 
    Suggested
    Work at office
    Local area
    Flexible hours

    PagerDuty

    Atlanta, GA
    5 days ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Atlanta, GA
    1 day ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Suggested
    Worldwide

    Inspire Brands

    Atlanta, GA
    5 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    4 days ago
  • $70 - $85 per hour

     ...redefine what’s possible, give shape to the future—and get there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance models, and engineering guardrails that enable scalable,... 
    Temporary work
    Local area
    Flexible hours
    3 days per week

    Slalom

    Atlanta, GA
    1 day ago
  •  ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview:   We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it... 
    Full time
    Live in
    Work at office

    Incident IQ

    Atlanta, GA
    15 days ago
  •  ...out new infrastructure capabilities to improve platform reliability and scalability. Monitor system health using metrics,...  .... Requirements At least 3 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles. Hands-on experience... 
    Full time
    Work at office
    Local area
    Flexible hours

    PagerDuty

    Atlanta, GA
    16 days ago
  • $167.7k - $245.2k

     ...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Atlanta, GA
    4 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Temporary work

    2T Consulting

    Atlanta, GA
    15 days ago
  • $111.61k - $131.3k

    At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life,...
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Atlanta, GA
    3 days ago
  •  ...our company effectively.    The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's...  ...infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline... 
    Remote work
    Flexible hours

    Intellum, Inc.

    Atlanta, GA
    14 days ago
  • $151k - $297k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    4 days ago
  • #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...

    LTM

    Atlanta, GA
    2 days ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's...  ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    Atlanta, GA
    18 days ago
  •  ...Consultancy and Information Technology Enabled Services.Job DescriptionSCM System EngineerSCM Continuous Integration / Delivery Build Team Engineer with experience in Application Service and Web Application Build, Deployment and Release Management and experience in establishing... 
    Permanent employment
    Full time
    H1b

    Career Guidant

    Atlanta, GA
    2 days ago
  •  ...evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical...  ...highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles... 

    Merican Inc

    Atlanta, GA
    6 days ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Atlanta, GA
    25 days ago
  •  ...technologies to enable scalable, secure, and reliable business operations. Applies strong...  ...infrastructure.3. Manages infrastructure engineering projects and processes aligned with...  ...benefit plans, please visit our Benefits site. Depending on the position and division,... 
    Permanent employment
    Full time
    Part time
    Work experience placement
    H1b
    Remote work
    Work visa
    Shift work
    Weekend work
    Day shift

    Truist

    Atlanta, GA
    3 days ago
  • $101.5k - $169.1k

     ...include an incentive program.Job DescriptionThe Release Train Engineer (RTE) has a primary purpose of supporting an Agile Release Train...  ...organizational AI policies and standards. Monitor AI tool reliability across teams. Create backup plans for system failures. Maintain... 
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    Cox Enterprises

    Atlanta, GA
    5 days ago
  • $105k - $130k

     ...provide the high-speed capabilities our nation and its allies need to maintain a durable, asymmetric advantage. The Mission Systems Engineering (MSE) Team develops the Mission Management System (MMS)—a software platform that integrates mission subsystems, autonomy services... 
    Weekly pay
    Permanent employment
    Full time
    Work at office

    Hermeus

    Atlanta, GA
    1 day ago
  • $141.3k - $237.4k

     ...AT&T, you won’t just imagine the future, you’ll build it.We are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role will drive innovation through automation, enhancement of... 
    Full time
    Temporary work
    Work at office
    Local area
    Relocation

    AT&T

    Atlanta, GA
    3 days ago
  • $68 - $75 per hour

     ...years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles• 2+ years...  ...problem solving abilities• Ability to work on-site in ATL, GA• Bachelor’s degree or...  ...experienceJob DescriptionWe are seeking a Reliability Engineer for our customer at the CDC.In... 
    Contract work
    Temporary work

    TEKsystems

    Atlanta, GA
    2 days ago
  •  ...That’s how we’re UNSTOPPABLE for our employees!Are you ready for the next chapter in your Uncarrier journey? The Sr. System Reliability Engineer (SRE) guides and mentors other SREs and improves and protects the software and systems behind all of T-Mobile's IT services,... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    1 day ago
  • $144k - $191k

     ...customers. We are looking for software engineers, hardware engineers, roboticists, and front...  ...-in-the-loop demonstrations at test sites.Define a strategy to comply with appropriate...  ...(V&V) plans to ensure robust and reliable system performance.Experience writing testable... 
    Full time
    Work experience placement
    Work at office
    Immediate start

    Anduril Industries

    Atlanta, GA
    1 day ago
  • $168.5k - $252.7k

     ...secure.About the RoleAs a Senior Software Engineer, you will play a key role in designing...  ...mentor team members to ensure high-velocity, reliable delivery.What You’ll DoDesign, build, and...  ...not Workday Careers. Please be aware of sites that may ask for you to input your data... 
    Full time
    Contract work
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Atlanta, GA
    4 days ago
  • $51 - $61 per hour

     ...onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities:As a Release Train Engineer, you will be responsible for facilitating Agile Release Train events and processes including communicating with stakeholders... 
    Hourly pay
    Live in
    Work at office
    Local area
    Immediate start
    Flexible hours
    Shift work

    Accenture

    Atlanta, GA
    2 days ago
  •  ...Reference26-00225 Job Title: ( Senior Software Configuration/Release Engineer ) About Kyyba: Founded in 1998 and headquartered in Farmington...  ...structure combined with career development. Job Description On-Site Interviews Only HYBRID - IN THE OFFICE 2 DAYS PER WEEK Minimum... 
    Work at office
    Visa sponsorship
    Work visa
    2 days per week

    Kyyba

    Atlanta, GA
    5 days ago
  • $165k - $190k

    OverviewJob PurposeThe Kubernetes Platform Engineering (KPE) team builds and operates ICE's internal container orchestration platform powered by Red Hat OpenShift. KPE Features, the Solutions Engineering sub-team, serves as the primary interface between the platform and... 
    Full time

    Black Knight Financial Services

    Atlanta, GA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!