Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

Atlanticus

Site Reliability Engineer (SRE)

When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being.

Atlanticus™ technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americans through the use of proprietary analytics. We apply the experience gained and infrastructure built from servicing over 20 million customers and over $40 billion in consumer loans over more than 25 years of operating history to support lenders that originate a range of consumer loan products. These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs.

Office Locations available for this role include:

  • Austin, TX – Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing.
  • Atlanta, GA – Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA.

We foster a collaborative, innovative environment where everyone contributes to building something meaningful. You'll be empowered to lead, grow, and make an impact.

The Role

We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.

The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teams to maintain highly available production services.

The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.

Key Responsibilities

  • Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.
  • Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.
  • Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
  • Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
  • Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
  • Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
  • Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
  • Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
  • Build self-healing and automated operational processes to improve system reliability and reduce manual intervention.
  • Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues.
  • Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks.
  • Partner with Development teams to improve application reliability, resiliency, scalability, and performance.
  • Continuously improve operational processes, monitoring coverage, automation, and deployment practices.
  • Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency.
Qualifications

You're a great fit if you have:

  • 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role.
  • Strong experience supporting Java-based applications in production environments.
  • Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC.
  • Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments.
  • Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis.
  • Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar).
  • Experience supporting MySQL and Oracle databases from an application support perspective.
  • Proficiency in Python, Bash, or other scripting languages for automation.
  • Strong Linux system administration and troubleshooting skills.
  • Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure.
  • Knowledge of networking fundamentals including DNS, TCP/IP, TLS, load balancing, and firewalls.
  • Experience with incident management, problem management, and change management processes.
  • Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency.

Preferred

  • Experience with Helm and GitOps deployment models.
  • Experience with Terraform or Infrastructure as Code.
  • Familiarity with Prometheus, Grafana, or OpenTelemetry.
  • Experience supporting microservices architectures.
  • Knowledge of JVM tuning and Java performance optimization.
  • Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler.

Why You'll Love Working Here

This isn't just a job, it's a place to lead, grow, and thrive. If you believe in your skills and drive, we'll provide the resources and support to help you succeed.

Benefits include:

  • Generous PTO and holiday schedule
  • 401(k) with company match
  • Employee stock purchase plan
  • Ongoing training (lunch & learns, financial and health webinars)
  • Team volunteer outings

Atlanticus is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, religion, gender, sexual orientation, age, veteran status, disability, or other protected status.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in Atlanta, GA vacancy
  •  ...troubleshooting staging and production cloud environments . Experienced in architectural design for reliability, scalability, and performance. Practical application of SRE principles : SLIs, SLOs, error budgets, automation, incident management, and postmortems.... 
    Suggested

    Purple Drive

    Atlanta, GA
    1 day ago
  •  ...Role: Site Reliability Engineering (SRE) Architect Location: Atlanta, GA (Hybrid on-site) Contract Role Summary: As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems... 
    Suggested
    Contract work
    Early shift

    AceStack LLC

    Atlanta, GA
    5 days ago
  •  ...'re in the right place. Job Details: Job Title: SRE Engineer Location: Atlanta GA (Hybrid Duration: 1 year...  ...availability critical application components. 1+ Years in Site Reliability Engineering organization preferred. Overall 4-6 years... 
    Suggested
    Contract work
    Work experience placement
    Work at office
    Remote work

    Staffworxs Inc

    Atlanta, GA
    2 days ago
  • $104.9k - $174.7k

     ...Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    RELX Group

    Atlanta, GA
    5 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability...  ...on our platform.Partners with the larger Cloud Operations, SRE, Engineering teams, and the business-at-large to advance our... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    4 days ago
  • #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...

    LTM

    Atlanta, GA
    1 day ago
  • $138.1k - $198.2k

     ...and more intuitive with technology that simply works.  The SRE Engineering Enablement Team supports our CI Platforms, Developer Environments...  ...Our customers are all engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter of our engineering... 
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Atlanta, GA
    5 days ago
  • $60 - $68 per hour

     ...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential...  ...the automation AWS Pipeline & Infrastructure, DevOps (SRE) activities, Monitoring & Alerting Our client is a... 
    Contract work
    Local area
    Immediate start

    Pyramid Consulting

    Atlanta, GA
    3 days ago
  • $155k - $222.6k

     ...soil. Meet the Team The SRE Fleet team is responsible for...  ...platform. As a team of six engineers distributed across the US, Canada...  ...strong focus on automation, reliability, and operational excellence....  ...~2+ years of experience in Site Reliability Engineering, DevOps... 
    Permanent employment
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    Cisco

    Atlanta, GA
    2 days ago
  •  ...Senior Site Reliability Engineer Atlanta, Georgia Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered...  ...best practices. Actively contribute to fostering an SRE culture within the organization by promoting observability,... 
    Permanent employment
    Full time
    Work at office
    Remote work
    Work from home
    Work visa

    QGenda

    Atlanta, GA
    1 day ago
  • $74.1k - $148.3k

     ..., software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining...  ...Responsibilities Service Ownership –You will be part of the SRE team, whose mission is the shared full stack ownership of a... 
    Temporary work
    Immediate start
    Flexible hours

    Oracle

    Atlanta, GA
    5 days ago
  •  ...Job description Snowflake SRE JD Your Role Accountabilities Primarily responsible for administrating Snowflake environments on AWS Identify, tune, and fix the performance issues on priority. Diagnose and troubleshoot Snowflake related errors and work with team to raise... 

    Rcinfotech

    Atlanta, GA
    1 day ago
  • $95k - $171k

     ...infrastructure? Do you want to build your SRE career on one of the most exciting...  ..., Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Atlanta, GA
    4 days ago
  • $71.6k - $119.4k

     ...Our services provide applications with reliability, security, and better customer experiences...  ...issues, and work closely with senior engineers to learn and apply best practices. You'll...  ...: ~1-3 years of experience in DevOps, SRE, cloud engineering, or related IT roles... 
    Temporary work
    Internship
    Local area

    RELX

    Atlanta, GA
    5 days ago
  •  ...We are currently looking for a Senior Software Engineer to be a part of the Site Reliability Engineering (SRE) team in Atlanta, GA . The SRE team is an innovative team devoted to providing a Docker-based Platform as a Service and assisting a growing number of teams... 
    Contract work
    Work at office
    Local area

    Spartan Technologies

    Atlanta, GA
    5 days ago
  •  ...Site Reliability Engineer Opportunity Rainforest is an early stage payments-as-a-service startup that has developed a solution that makes monetizing...  ...about making a real impact in fintech, and helping shape SRE practices as the company grows, you'll feel right at home at... 
    Work experience placement
    Flexible hours

    Rainforest

    Atlanta, GA
    3 days ago
  • $121.4k - $218.6k

     ...challenges? Join our critical AI Hardware SRE Team! The AI Hardware SRE team is...  ...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure...  ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing... 
    Work experience placement
    Work at office

    Akamai

    Atlanta, GA
    4 days ago
  •  ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively... 
    Work experience placement

    Fisec Global

    Atlanta, GA
    3 days ago
  •  ...rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months...  ...learning, and operational excellence. Drive the adoption of SRE best practices and ensure adherence to reliability and performance... 
    Contract work
    Local area
    Immediate start

    Navtech

    Atlanta, GA
    5 days ago
  •  ..., resilient, and highly available software platforms using Site Reliability Engineering and AI-native engineering practices. The engineer collaborates...  .... This position provides technical leadership for the SRE team, strengthens the India GCC capability, and establishes... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    20 hours ago
  • $160k - $210k

     ...teamwork a cornerstone of our success. We are looking for a Senior Site Reliability engineer to work on expanding our global footprint of datacenters and...  ...of experience in operations, software engineering, or as an SRE. Experience with hybrid cloud/on-prem solutions. Hands on... 
    Work at office
    Local area
    Immediate start
    Remote work

    GrabJobs

    Atlanta, GA
    3 days ago
  • $105.79k - $141.05k

     ...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is... 
    Temporary work
    Remote work

    Lumen Inc

    Atlanta, GA
    1 day ago
  •  ...the best job for you. Role: Platform Engineer - Ansible Automation Platform (AAP / Tower) | DevOps / SRE Location: Atlanta GA Duration: 6 month...  ...maintaining strong controls and approvals SRE & Reliability Engineering Apply SRE principles to improve... 
    Permanent employment
    Contract work
    Remote work

    Tekfortune Inc

    Atlanta, GA
    2 days ago
  • $152.13k - $162.13k

     ...challenge the status-quo.Unum is changing, and we’re excited about what’s next. Join us.General Summary:Unum Group seeks Site Reliability Engineers in Atlanta, GA.Applicants who are interested in this position may apply at (Ref #66753) for consideration.Design, build,... 
    Full time
    Temporary work
    Work at office
    Remote work

    Unum Group

    Atlanta, GA
    3 days ago
  •  ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer... 
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    5 days ago
  • $218.11k - $274.19k

     ...Site Reliability Engineer We are Omnissa! Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from anywhere experiences. We integrate industry-leading solutions—including Unified Endpoint Management, Virtual Apps and Desktops... 
    Work experience placement
    Local area
    Remote work
    Visa sponsorship
    Flexible hours

    Omnissa

    Atlanta, GA
    1 day ago
  •  ...The Home Depot is seeking a Senior Software Reliability Engineer to join the Platform Reliability Engineering team, ensuring the resilience, performance, and security of our enterprise Cloud Platform. You will mentor junior engineers, lead incident triage, root cause... 

    Home Depot

    Atlanta, GA
    1 day ago
  • $116.48k - $174.71k

     ...and society. The ChallengeWe're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this...  ...servicesManage and enforce error budgets to balance system reliability with product feature velocity.Improving alert quality by reducing... 
    Work experience placement
    Work at office
    Local area
    Worldwide
    Flexible hours
    3 days per week
    1 day per week

    OneTrust

    Atlanta, GA
    3 days ago
  • $130k - $180k

     ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Atlanta, GA
    3 days ago
  •  ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-... 
    Full time
    Live in
    Work at office

    Incident IQ

    Atlanta, GA
    14 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!