Site Reliability Engineering (SRE) Architect
Cloud Hybrid Technologies LLC
Site Reliability Engineering (SRE) Architect Location: Atlanta,GA Duration: 12Months + Extension Hourly Rate: Depending on Experience (DOE) Work Authorization: As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical services. Moving beyond day-to-day operations, you will focus on the strategic architectural direction of SRE function, defining standards, blueprints, and frameworks that enable development teams and fellow SRE operations team to build and operate highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles to influence technology choices, establish best practices, and foster a proactive culture of reliability across the organization and much beyond observability pillar. Key Responsibilities Reliability Strategy & Design: Architect and design highly available, scalable, secure, and cost-effective infrastructure and application patterns on AWS Define and evangelize SRE best practices, standards, and blueprints for service design, deployment, monitoring, and operational readiness across the engineering organization Review current observability implementation to identify gaps and define steps to reach next level maturity of observability setup to provide deep insights into system health and behaviour With overall maturity lead the definition and implementation strategy for Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets for critical services Design solutions to systematically reduce operational toil through automation and improved system design Evaluate current SRE tools and automation frameworks (e.g., CI/CD pipelines, Infrastructure as Code modules, automated incident remediation, chaos engineering platforms) and suggest enhancement that will help overall enhancement of capability Evaluate, prototype, and recommend new technologies, tools, and methodologies to enhance system reliability, developer productivity, and operational efficiency Technical Leadership & Consultation: Act as a senior technical advisor and subject matter expert on reliability, scalability, and performance for development and platform teams Provide architectural guidance during the design phase of new services and features to ensure reliability principles are embedded early (shift-left) Mentor and coach other SREs and engineers, fostering technical excellence and adherence to SRE principles Lead architectural reviews and production readiness assessments for critical systems Resilience: Lead blameless postmortems for significant incidents, ensuring root causes are identified and systemic architectural improvements are prioritized and implemented Architect and advocate for resilience patterns (e.g., circuit breaking, rate limiting, graceful degradation, chaos engineering) within applications and infrastructure Required Qualifications Proven experience in an architectural role, designing solutions for reliability, scalability, and performance Deep understanding and practical application of SRE principles (SLIs/SLOs, error budgets, toil reduction, automation, incident management, postmortems) Expertise in cloud computing platforms (e.g., AWS) including infrastructure, networking, and security services Strong experience with containerization and orchestration technologies (Kubernetes, Docker, serverless computing) Solid experience designing and implementing observability solutions (e.g., Dynatrace, Prometheus, Grafana, ELK/EFK Stack, Jaeger, OpenTelemetry) Strong programming/scripting skills (e.g., Python, Go, Bash) for automation and tool development Excellent analytical, problem-solving, and strategic thinking skills. Strong communication, collaboration, and leadership skills with the ability to influence technical direction across teams Preferred Qualifications Experience designing and implementing chaos engineering practices and platforms Cloud Hybrid is an equal opportunity employer inclusive of female, minority, disability and veterans, (M/F/D/V). Hiring, promotion, transfer, compensation, benefits, discipline, termination and all other employment decisions are made without regard to race, color, religion, sex, sexual orientation, gender identity, age, disability, national origin, citizenship/immigration status, veteran status or any other protected status. Cloud Hybrid will not make any posting or employment decision that does not comply with applicable laws relating to labor and employment, equal opportunity, employment eligibility requirements or related matters. Nor will Cloud Hybrid require in a posting or otherwise U.S. citizenship or lawful permanent residency in the U.S. as a condition of employment except as necessary to comply with law, regulation, executive order, or federal, state, or local government contract #J-18808-Ljbffr
- ...Role Summary: As an SRE Architect, you will be a pivotal technical leader responsible... ...systems and practices that ensure the reliability, scalability, performance, and efficiency... .... Leverage deep expertise in software engineering, distributed systems, cloud...SuggestedEarly shift
$99.09k - $123.86k
...Site Reliability Engineer (SRE) – AI Systems We are seeking an experienced Site Reliability Engineer who thrives at the intersection of software engineering, infrastructure, and AI systems. The role focuses on ensuring our platforms are scalable, reliable, and secure,...SuggestedLocal areaFlexible hours$130k - $135k
...VARITE is looking for a qualified Senior Site Reliability Engineer (SRE) - 619374 in Atlanta, GA About the client: An American Software company that provides a suite of tools intended to support the development and deployment of large-scale service-oriented software installations...SuggestedFull time- ...Job title: Site Reliability Engineer (SRE) Location: Atlanta, GA 30303 Duration of the project: 12 Months Strong expertise in Ansible with an SRE background Ability to review and test GitLab Duo generated code CICD pipelines (GitLab, GitHub Actions) Infrastructure...Suggested
- ...advances cures by helping the world's most important research sites do their best work. Our solutions are now used by over 30,0... ...What You'll Bring to the Team: We are seeking a Site Reliability Engineer (SRE) to join one of our Scrum teams and help ensure the...SuggestedWork at office
- .... Job Details: Job Title: SRE Engineer Location: Atlanta GA (Hybrid... ...strategies. Work with enterprise security architects to design and implement data security... ...components. 1+ Years in Site Reliability Engineering organization preferred....Contract workWork experience placementWork at officeRemote work
$130k - $160k
...Role At Todyl, our Application Platform Engineering team is dedicated to building... ...work will not only directly impact the reliability and security of our platform but also empower... ...space. Responsibilities As a Platform SRE (Site Reliability Engineer) at Todyl, you will...Temporary workLocal areaFlexible hours- ...Who We’re Looking For We’re looking for a proactive, hands‑on Site Reliability Engineer who thrives in building and scaling cloud infrastructure in... ...about making a real impact in fintech, and helping shape SRE practices as the company grows, you’ll feel right at home at...Work experience placementFlexible hours
- ...qgenda.com or follow us on Instagram or LinkedIn. As a Senior Site Reliability Engineer, you will work with our Infrastructure and Product... ...operations best practices. Actively contribute to fostering an SRE culture within the organization by promoting observability,...Permanent employmentWork at officeRemote workWork from home
$152k - $195k
...Capital, GV and Riverwood Capital. About the Team As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and... ...and cloud infrastructure. Required Qualifications 6+ years in SRE , DevOps , or Infrastructure roles , with significant...$117k - $209.33k
...An exciting new opportunity has opened for a Site Reliability Engineer within the Autodesk PDMS Platform SRE team. The successful candidate will wear multiple... ...hats: first responder, performance analyst, system architect, capacity planner, and monitoring expert. Good technical...Permanent employmentFor contractors- ...I have an opportunity for "SRE " _ (Atlanta, GA - ONSITE)" and I am looking for a candidate who can join Immediately if you are... ...Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in...Immediate start
- ...This position is 60 % SRE and 40% SDE. Also open for candidates... ...metrics so the support engineers can proactively and timely validate... ...with enterprise security architects to design and implement data... ...components. • 1+ Years in Site Reliability Engineering organization...Work experience placement
- ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront... ...for all infrastructure and services within the scope of SRE Preserve operational visibility and response capabilities...
- ...Technical Support Specialist In Site Reliability Engineering (Sre) Mandatory skills: Scripting and programming languages like Python, Java, Ruby. Cloud and infrastructure management – AWS, Google cloud and Azure is a plus- CI/CD Automation, Database Management. The...
$116.48k - $174.71k
...Senior Software Engineer - SRE Atlanta, Georgia Strength in Trust OneTrust's mission is to enable innovation through the responsible... ...Manage and enforce error budgets to balance system reliability with product feature velocity. Improving alert quality by...Work experience placementWork at officeLocal areaWorldwideFlexible hours3 days per week1 day per week$109.5k
...highly motivated, diligent, and skillful Site Reliability Engineer to join the Cyber Security Engineering... ...latest content and functionality. The SRE, coordinating with developers, will... ...grows with us. Responsibilities: Architect new and existing systems to enhance...Temporary workLocal areaRemote work$51.9 per hour
...Company: Allegheny Health Network Job Title: Site Reliability Engineering – Clinical & Facility Services General Overview This role ensures the reliability... ...HIPAA and other regulatory compliance. (15%) Oversee SRE collaboration with Clinical Engineering and Cybersecurity...Local area$178k - $213k
...Partners 2022 Cybersecurity Excellence Award for MDR Manager, Site Reliability Engineering Reports to: VP, Product Engineering Location: While... ...operational resilience—all while mentoring a high‑caliber SRE team. What You’ll Do Lead and grow the SRE team, setting direction...Permanent employmentWork experience placementWork at officeRemote workWork from homeHome officeFlexible hours- ...the best job for you. Role: Platform Engineer - Ansible Automation Platform (AAP / Tower) | DevOps / SRE Location: Atlanta GA Duration: 6 month... ...maintaining strong controls and approvals SRE & Reliability Engineering Apply SRE principles to improve...Permanent employmentContract workRemote work
$126k - $248k
...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and...Local areaRemote workWorldwideFlexible hours- ...an experienced Observability Architect to design, implement, and mature... ...visibility, performance, and reliability. Key Responsibilities... ...infrastructure, application, and SRE teams to ensure high... ...performance. Automation & AI-Driven Engineering • Build automated workflows...
- ...Platform Architect · Evergreen Insight Global Hybrid on site at Insight Global HQ in Dunwoody M-Th... ...in from day one Setting engineering standards for agentic... ...starts Owning platform reliability — SLAs, observability,... ...Apigee, or equivalent) SRE mindset — you've owned...Work from home
$178.13k - $205.4k
...telecommuting. Salary Range: $178,131 - $205,400 About You Basic Qualification Bachelor’s degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5) years of progressive, post‑baccalaureate experience in job offered...Work at officeRemote workFlexible hours$116.48k - $174.71k
...and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In... ...services. Manage and enforce error budgets to balance system reliability with product feature velocity. Improve alert quality by reducing...Work experience placementWork at officeLocal areaWorldwideFlexible hours3 days per week1 day per week- ...Site Reliability Engineer We are looking for a Site Reliability Engineer to ensure the reliability, security, and continuous operation of a multi‑cloud application security platform. This role combines platform engineering and security automation, focusing on Kubernetes...Work at officeRemote workVisa sponsorshipWork visaFlexible hours
- ...looks like. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In this... ...services. Manage and enforce error budgets to balance system reliability with product feature velocity. Improve alert quality by reducing...Local areaFlexible hours
$94.9k - $135.6k
...safely and efficiently. Cardinal Health is seeking a Release Engineer to lead iteration and release management activities supporting mission... ...with Solution Owners, Scrum Masters, Engineering, Testing, SRE, and Operations to align scope, sequencing, dependencies, and...Temporary workLocal areaImmediate startFlexible hours- ...systems. As a Staff Platform Engineer, you will play a critical role... ...role. You will own reliability for major platform domains, design... ...operate their applications Architect, implement, and manage highly... ...years of experience as a Staff SRE with a strong focus on...
- ...customers. As we scale globally, reliability, availability, and... ...features. As a Principal Engineer, you will define and drive the... ...operate their applications • Architect, implement, and manage highly... ...of experience as a Principal SRE with a strong focus on building...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineering (SRE) Architect. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer remote Atlanta, GA
- site reliability engineer sre Atlanta, GA
- IT site lead Atlanta, GA
- website coordinator Atlanta, GA
- on site coordinator Atlanta, GA
- junior website developer Atlanta, GA
- site safety Atlanta, GA
- site services specialist Atlanta, GA
- site recruiter Atlanta, GA

