Site Reliability Engineer (SRE)
Atlanticus
Site Reliability Engineer (SRE)
When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being.
Atlanticus™ technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americans through the use of proprietary analytics. We apply the experience gained and infrastructure built from servicing over 20 million customers and over $40 billion in consumer loans over more than 25 years of operating history to support lenders that originate a range of consumer loan products. These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs.
Office Locations available for this role include:
- Austin, TX – Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing.
- Atlanta, GA – Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA.
We foster a collaborative, innovative environment where everyone contributes to building something meaningful. You'll be empowered to lead, grow, and make an impact.
The Role
We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.
The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teams to maintain highly available production services.
The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.
Key Responsibilities
- Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.
- Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.
- Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
- Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
- Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
- Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
- Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
- Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
- Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
- Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
- Build self-healing and automated operational processes to improve system reliability and reduce manual intervention.
- Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues.
- Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks.
- Partner with Development teams to improve application reliability, resiliency, scalability, and performance.
- Continuously improve operational processes, monitoring coverage, automation, and deployment practices.
- Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency.
Qualifications
You're a great fit if you have:
- 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role.
- Strong experience supporting Java-based applications in production environments.
- Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC.
- Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments.
- Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis.
- Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar).
- Experience supporting MySQL and Oracle databases from an application support perspective.
- Proficiency in Python, Bash, or other scripting languages for automation.
- Strong Linux system administration and troubleshooting skills.
- Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure.
- Knowledge of networking fundamentals including DNS, TCP/IP, TLS, load balancing, and firewalls.
- Experience with incident management, problem management, and change management processes.
- Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency.
Preferred
- Experience with Helm and GitOps deployment models.
- Experience with Terraform or Infrastructure as Code.
- Familiarity with Prometheus, Grafana, or OpenTelemetry.
- Experience supporting microservices architectures.
- Knowledge of JVM tuning and Java performance optimization.
- Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler.
Why You'll Love Working Here
This isn't just a job, it's a place to lead, grow, and thrive. If you believe in your skills and drive, we'll provide the resources and support to help you succeed.
Benefits include:
- Generous PTO and holiday schedule
- 401(k) with company match
- Employee stock purchase plan
- Ongoing training (lunch & learns, financial and health webinars)
- Team volunteer outings
Atlanticus is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, religion, gender, sexual orientation, age, veteran status, disability, or other protected status.
- ...troubleshooting staging and production cloud environments . Experienced in architectural design for reliability, scalability, and performance. Practical application of SRE principles : SLIs, SLOs, error budgets, automation, incident management, and postmortems....Suggested
- ...Role: Site Reliability Engineering (SRE) Architect Location: Atlanta, GA (Hybrid on-site) Contract Role Summary: As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems...SuggestedContract workEarly shift
- ...'re in the right place. Job Details: Job Title: SRE Engineer Location: Atlanta GA (Hybrid Duration: 1 year... ...availability critical application components. 1+ Years in Site Reliability Engineering organization preferred. Overall 4-6 years...SuggestedContract workWork experience placementWork at officeRemote work
$104.9k - $174.7k
...Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role...SuggestedFull timeWork at officeLocal areaRemote workWork from home$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability... ...on our platform.Partners with the larger Cloud Operations, SRE, Engineering teams, and the business-at-large to advance our...SuggestedFull timeTemporary workWork experience placementFlexible hours- #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...
$138.1k - $198.2k
...and more intuitive with technology that simply works. The SRE Engineering Enablement Team supports our CI Platforms, Developer Environments... ...Our customers are all engineers at Cisco. Your Impact As a Site Reliability Engineer, you will be at the epicenter of our engineering...Permanent employmentFull timeTemporary workWork experience placementLocal areaRemote workFlexible hours$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential... ...the automation AWS Pipeline & Infrastructure, DevOps (SRE) activities, Monitoring & Alerting Our client is a...Contract workLocal areaImmediate start$155k - $222.6k
...soil. Meet the Team The SRE Fleet team is responsible for... ...platform. As a team of six engineers distributed across the US, Canada... ...strong focus on automation, reliability, and operational excellence.... ...~2+ years of experience in Site Reliability Engineering, DevOps...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- ...Senior Site Reliability Engineer Atlanta, Georgia Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered... ...best practices. Actively contribute to fostering an SRE culture within the organization by promoting observability,...Permanent employmentFull timeWork at officeRemote workWork from homeWork visa
$74.1k - $148.3k
..., software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining... ...Responsibilities Service Ownership –You will be part of the SRE team, whose mission is the shared full stack ownership of a...Temporary workImmediate startFlexible hours- ...Job description Snowflake SRE JD Your Role Accountabilities Primarily responsible for administrating Snowflake environments on AWS Identify, tune, and fix the performance issues on priority. Diagnose and troubleshoot Snowflake related errors and work with team to raise...
$95k - $171k
...infrastructure? Do you want to build your SRE career on one of the most exciting... ..., Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours$71.6k - $119.4k
...Our services provide applications with reliability, security, and better customer experiences... ...issues, and work closely with senior engineers to learn and apply best practices. You'll... ...: ~1-3 years of experience in DevOps, SRE, cloud engineering, or related IT roles...Temporary workInternshipLocal area- ...We are currently looking for a Senior Software Engineer to be a part of the Site Reliability Engineering (SRE) team in Atlanta, GA . The SRE team is an innovative team devoted to providing a Docker-based Platform as a Service and assisting a growing number of teams...Contract workWork at officeLocal area
- ...Site Reliability Engineer Opportunity Rainforest is an early stage payments-as-a-service startup that has developed a solution that makes monetizing... ...about making a real impact in fintech, and helping shape SRE practices as the company grows, you'll feel right at home at...Work experience placementFlexible hours
$121.4k - $218.6k
...challenges? Join our critical AI Hardware SRE Team! The AI Hardware SRE team is... ...for ensuring best-in-class uptime and reliability of our AI hardware infrastructure... ...when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing...Work experience placementWork at office- ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively...Work experience placement
- ...rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months... ...learning, and operational excellence. Drive the adoption of SRE best practices and ensure adherence to reliability and performance...Contract workLocal areaImmediate start
- ..., resilient, and highly available software platforms using Site Reliability Engineering and AI-native engineering practices. The engineer collaborates... .... This position provides technical leadership for the SRE team, strengthens the India GCC capability, and establishes...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours
$160k - $210k
...teamwork a cornerstone of our success. We are looking for a Senior Site Reliability engineer to work on expanding our global footprint of datacenters and... ...of experience in operations, software engineering, or as an SRE. Experience with hybrid cloud/on-prem solutions. Hands on...Work at officeLocal areaImmediate startRemote work$105.79k - $141.05k
...the future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is...Temporary workRemote work- ...the best job for you. Role: Platform Engineer - Ansible Automation Platform (AAP / Tower) | DevOps / SRE Location: Atlanta GA Duration: 6 month... ...maintaining strong controls and approvals SRE & Reliability Engineering Apply SRE principles to improve...Permanent employmentContract workRemote work
$152.13k - $162.13k
...challenge the status-quo.Unum is changing, and we’re excited about what’s next. Join us.General Summary:Unum Group seeks Site Reliability Engineers in Atlanta, GA.Applicants who are interested in this position may apply at (Ref #66753) for consideration.Design, build,...Full timeTemporary workWork at officeRemote work- ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer...Full timeWorldwideFlexible hours
$218.11k - $274.19k
...Site Reliability Engineer We are Omnissa! Omnissa is the first AI-driven digital work platform, built to support flexible, secure, work-from anywhere experiences. We integrate industry-leading solutions—including Unified Endpoint Management, Virtual Apps and Desktops...Work experience placementLocal areaRemote workVisa sponsorshipFlexible hours- ...The Home Depot is seeking a Senior Software Reliability Engineer to join the Platform Reliability Engineering team, ensuring the resilience, performance, and security of our enterprise Cloud Platform. You will mentor junior engineers, lead incident triage, root cause...
$116.48k - $174.71k
...and society. The ChallengeWe're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this... ...servicesManage and enforce error budgets to balance system reliability with product feature velocity.Improving alert quality by reducing...Work experience placementWork at officeLocal areaWorldwideFlexible hours3 days per week1 day per week$130k - $180k
...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and... ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to...Temporary workWork at officeImmediate startRemote workFlexible hours- ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-...Full timeLive inWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- site reliability engineer remote Atlanta, GA
- site services specialist Atlanta, GA
- construction site safety Atlanta, GA
- site leader Atlanta, GA
- official site Atlanta, GA
- website content developer Atlanta, GA
- on site coordinator Atlanta, GA
- IT site lead Atlanta, GA

