Site Reliability Engineering (SRE) Architect
Cloud Analytics Technologies LLC
Role Summary:
As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical services. Moving beyond day-to-day operations, you will focus on the strategic architectural direction of SRE function, defining standards, blueprints, and frameworks that enable development teams and fellow SRE operations team to build and operate highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles to influence technology choices, establish best practices, and foster a proactive culture of reliability across the organization and much beyond observability pillar.
Key Responsibilities:
- Reliability Strategy & Design:
- Architect and design highly available, scalable, secure, and cost-effective infrastructure and application patterns on AWS
- Define and evangelize SRE best practices, standards, and blueprints for service design, deployment, monitoring, and operational readiness across the engineering organization
- Review current observability implementation to identify gaps and define steps to reach next level maturity of observability setup to provide deep insights into system health and behaviour
- With overall maturity lead the definition and implementation strategy for Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets for critical services
- Platform Architecture & Automation:
- Design solutions to systematically reduce operational toil through automation and improved system design
- Evaluate current SRE tools and automation frameworks (e.g., CI/CD pipelines, Infrastructure as Code modules, automated incident remediation, chaos engineering platforms) and suggest enhancement that will help overall enhancement of capability
- Evaluate, prototype, and recommend new technologies, tools, and methodologies to enhance system reliability, developer productivity, and operational efficiency
- Technical Leadership & Consultation:
- Act as a senior technical advisor and subject matter expert on reliability, scalability, and performance for development and platform teams
- Provide architectural guidance during the design phase of new services and features to ensure reliability principles are embedded early (shift-left)
- Mentor and coach other SREs and engineers, fostering technical excellence and adherence to SRE principles
- Lead architectural reviews and production readiness assessments for critical systems
- Resilience:
- Lead blameless postmortems for significant incidents, ensuring root causes are identified and systemic architectural improvements are prioritized and implemented
- Architect and advocate for resilience patterns (e.g., circuit breaking, rate limiting, graceful degradation, chaos engineering) within applications and infrastructure
Required Qualifications:
- Proven experience in an architectural role, designing solutions for reliability, scalability, and performance
- Deep understanding and practical application of SRE principles (SLIs/SLOs, error budgets, toil reduction, automation, incident management, postmortems)
- Expertise in cloud computing platforms (e.g., AWS) including infrastructure, networking, and security services
- Strong experience with containerization and orchestration technologies (Kubernetes, Docker, serverless computing)
- Solid experience designing and implementing observability solutions (e.g., Dynatrace, Prometheus, Grafana, ELK/EFK Stack, Jaeger, OpenTelemetry)
- Strong programming/scripting skills (e.g., Python, Go, Bash) for automation and tool development
- Excellent analytical, problem-solving, and strategic thinking skills.
- Strong communication, collaboration, and leadership skills with the ability to influence technical direction across teams
Preferred Qualifications:
- Experience designing and implementing chaos engineering practices and platforms
- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...Suggested
- Direct message the job poster from STAFFWORXS Delivery Manager @ STAFFWORXS | US IT Recruitment Job Opening: AWS Site Reliability Engineer (SRE) We’re hiring a Site Reliability Engineer (SRE) to join our team in Atlanta, GA. This hybrid role offers the opportunity to work...SuggestedContract work
- ...Technology Consultant - Site Reliability Engineer (SRE) Location- - Atlanta, GA (hybrid schedule) Client is currently seeking an experienced Technology Consultant Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability...SuggestedPermanent employment
- ...Job Title: Site Reliability Engineer II (SRE II) Data & Intelligence Location: Atlanta, GA Contract Job Summary The Site Reliability Engineer II (SRE II) is responsible for ensuring the reliability, scalability, performance, security, and operational...SuggestedContract work
- ...Job Description Job Description We are seeking a highly experienced Site Reliability Engineer (SRE) with deep expertise in Dynatrace, observability engineering, and Azure cloud technologies . This role will be exclusively focused on building, enhancing, and managing...SuggestedWork at officeLocal area
$70k - $110k
...architectural position focused on reliability, scalability, and performance... ...practical knowledge of SRE principles, including SLIs/SLOs... ...or running AI-enabled engineering tools in production Expertise... ...preferred Responsibilities: Architect highly available, scalable,...Long term contractFull time- ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability... ...on our platform.Partners with the larger Cloud Operations, SRE, Engineering teams, and the business-at-large to advance our...Full timeTemporary workWork experience placementFlexible hours- Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high... ...has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering...Worldwide
- #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...
- ...review the following job description:The Site Reliability Engineer role focuses on enhancing the... ...standardizing observability practices, mentoring SRE team members, and contributing to... ...Reliability Engineering & Automation Architect and deliver automation solutions that...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$35 - $44 per hour
DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and...Remote work$70 - $85 per hour
...redefine what’s possible, give shape to the future—and get there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance models, and engineering guardrails that enable scalable,...Temporary workLocal areaFlexible hours3 days per week- ...System Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give our customers. All team members receive a competitive base salary and compensation package - this is Total Rewards. Employees...Full timeTemporary workPart timeWork experience placementFlexible hours
- ...scalable, and resilient software platforms using SRE and AI-native engineering practices. Own production reliability, monitoring, and operational automation while mentoring... ..., and production operations. Key Skills Site Reliability Engineering Terraform Python AWS Azure...Temporary workFlexible hours
- ...PwC is seeking an experienced SRE - Enterprise & Cloud Security - AI Driven Security - Manager in Atlanta to lead AI/ML-enabled security... ...firm standards. Leverage data, algorithms, and software engineering to deliver scalable security solutions and drive business growth...
- ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-...Full timeLive inWork at office
$132.23k - $176.31k
...future of AI‑ready connectivity, join us today. The Role We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role...Full timeTemporary workRemote work$113k - $171.6k
...Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office,you'll help build and operate the foundational infrastructure...Work at officeLocal areaFlexible hours- ...an experienced Observability Architect to design, implement, and mature... ...visibility, performance, and reliability. Key Responsibilities... ...infrastructure, application, and SRE teams to ensure high... ...performance. Automation & AI-Driven Engineering • Build automated workflows...
$155.4k - $261.1k
...drive our enterprise forward. Our Systems Reliability and Software Delivery teams are... ...high-quality postmortems and partner with engineering teams to drive corrective actions and long... ...Systems Engineering, ITSM, RM/CMBackground in SRE, Support or QAOne or more of the...Permanent employmentFull timeTemporary workWork at officeLocal areaRelocationShift work$110k - $150k
...Technical Architect Being good neighbors – helping people, investing... ...with product owners, engineering teams, business stakeholders,... ...integration, security, observability, reliability). Drive API-first and... ...practices aligned to DevSecFinOps, SRE, and SLOs/SLIs. ~...H1bWork at officeImmediate startWork from homeRelocation$167.7k - $245.2k
...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze...Full timeTemporary workLocal areaFlexible hours- CBRE Group, Inc. seeks an Operations Consultant to provide technical expertise across complex technology systems and support daily operations within the D&T Support function. You will collaborate with IT teams, optimize configurations, and guide stakeholders on technology...
- ...rollout of new infrastructure capabilities to improve platform reliability and scalability. Monitor system health through metrics... ...Requirements Require 0 to 1+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles. Require hands-...Full timeWork at officeFlexible hours
$142.4k - $213.6k
...wide swath of Workday's technical areas. Our IA team has built several critical tools that support Workday's SaaS properties, its engineers, and its management. Our products and tools typically go from a need based conceptual discussion with our users to production deployment...Full timeWork at officeRemote workHome officeFlexible hoursShift work- The Home Depot is seeking a Systems Engineering Senior Manager to inspire blended teams of systems and software engineers and deliver... ...partner with product teams and vendors, and steer change to improve reliability, developer experience, and #J-18808-Ljbffr The Home Depot
$116.48k - $174.71k
...and society. The Challenge We're looking for a Senior Software Engineer that will report to the Development Manager / R&D Head. In... ...services. Manage and enforce error budgets to balance system reliability with product feature velocity. Improve alert quality by reducing...Work experience placementWork at officeLocal areaWorldwideFlexible hours3 days per week1 day per week- ...our company effectively. The Lead Systems Engineer is a senior individual contributor responsible for the reliability, scalability, and modernization of Intellum's... ...infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline...Remote workFlexible hours
$141.3k - $237.4k
...won’t just imagine the future, you’ll build it.We are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role will drive innovation through automation, enhancement of observability...Full timeTemporary workWork at officeLocal areaRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineering (SRE) Architect. Be the first to apply!
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- website coordinator Atlanta, GA
- on-site clinical research associate (traveling/remote) Atlanta, GA
- site safety Atlanta, GA
- junior website developer Atlanta, GA
- construction site safety Atlanta, GA
- IT site lead Atlanta, GA
- website content developer Atlanta, GA
- site recruiter Atlanta, GA



