Senior Cloud Site Reliability Engineer (SRE)
$104k - $166kPeraton
Responsibilities
Peraton is looking for a Senior Cloud Site Reliability Engineer (SRE) who will be responsible for designing and developing advanced Python-based AWS cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. This person will apply software engineering practices including Infrastructure-as-Code (IaC) with Terraform to build scalable, reusable solutions and utilities that enhance platform reliability across the Federal Reserve System.
Work Location: This is a remote position
What You Will Do
- Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system
- Build automation scripts, APIs, and utilities in Python to reduce toil and improve platform reliability.
- Implement observability and monitoring solutions (Grafana, AWS CloudWatch) leveraging Python for custom metrics and dashboards
- Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles
- Optimize Infrastructure as Code (IaC) with Terraform for AWS resources, integrating Python-based workflows.
- Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
- Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.
- Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development
- Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis
- Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions
- Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
- Produce clear, blameless postmortems with actionable items and documented failure scenarios
Qualifications
Basic Qualifications
- Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level Clearance
- Bachelors Degree and 8 years of experience, or a High School diploma or equivalent and 12 years of experience
- Must have 5+ years of advanced Python development experience, building enterprise-grade, highly available tools, APIs, and utilities for AWS
- 7+ years of extensive experience in software development with focus on reliability and platform engineering
- 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
- 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
- Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
- Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
- Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
- Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
- Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices
Preferred Qualifications
- Experience with GoLang or additional programming languages is a plus
- Bachelors Degree in Computer Science, Information Systems, or similar
Peraton Overview
Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world’s leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can’t be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we’re keeping people around the world safe and secure.
Target Salary Range
$104,000 - $166,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.
EEO
EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.
- ...Job Title: Senior Site Reliability Engineer (SRE) Work Location: Southlake, TX 76092 Contract duration: 12 months Interview Mode- In-person... ...Grafana Python Nice to have skills AI Cloud Detailed Job Description 5 years of experience with...CloudSeniorContract work
$185k - $200k
Back to All JobsSenior Site Reliability Engineer (SRE) Dayton, OH (Remote) full time Top Secret (TS) $185,000 -... ...Position Overview Metronome is seeking a Senior Site Reliability Engineer (SRE) to support AFRL's Google Cloud Platform environment. This engineer will...CloudSeniorFull timeRemote work- ...Opportunity: We are looking for a skilled engineer with disciplines that incorporate... ...approaches to observability and reliability. What you’ll do: • Evangelize SRE mindset and solve problems through... ..., and self-healing workflows for Cloud and Login Platforms. •...CloudSenior
$175k - $229k
...DevOps Engineer Instrumental builds the manufacturing acceleration... ...5 or more years of DevOps or SRE experience deploying and... ...commercial SaaS platforms on public cloud infrastructure, AWS preferred... ...ensure ongoing performance, reliability and efficiency. ~ Network/...CloudSenior$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic... ..., scalability, and long-term reliability of cloud native systems. Our SREs don’t just manage infrastructure...CloudSeniorRemote work- ...Role: Senior Site Reliability Engineer (SRE) Cloud & Kubernetes Location: Atlanta, GA (Onsite) Contract Role Summary: Lead the reliability, scalability, security, and operational excellence of customer-facing platforms across Azure, GCP, and Kubernetes...CloudSeniorContract work
$70k - $150k
...are seeking a highly motivated and technically strong Senior Lead, Site Reliability Engineering (SRE) to provide technical leadership within the DCOE Reliability... .... The role requires deep expertise in SRE principles, cloud and infrastructure operations, AIOps, automation...CloudSeniorFull timeContract workPart time- ...About The Role: We're looking for a Senior Site Reliability Engineer to help us mature and scale the infrastructure behind our multi-cloud SaaS platform. Most of our footprint... ...What You Bring: ~6+ years in SRE, DevOps, or infrastructure engineering roles...CloudSeniorRemote workFlexible hours
$140k - $180k
...center innovation, delivering a future-proof, cloud platform that redefines the customer experience... ...more at Opportunity We’re looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You’ll be a technical leader on a team...CloudSeniorWork experience placementLocal areaRemote workVisa sponsorshipWork visa- ...Job Description Job Description Job Title: Senior AWS Site Reliability Engineer (SRE) Location: Birmingham, Alabama Type: Contract To Hire Work... ...and coordinate resolution across application, cloud, DevOps, security, performance, and production support...CloudSeniorContract workLocal area
$92.7k - $203.94k
...Senior Site Reliability Engineer We're building a world of health around every individual — shaping a more... ..., and operations teams to embed SRE best practices, improve application resiliency... ...& Deployments Champion cloud-native architectures leveraging microservices...CloudSeniorHourly payFull timeTemporary work$63 - $90 per hour
...a leader financial services provider, is seeking a Senior DevOps Engineer / Site Reliability Engineer (SRE) to support enterprise backup, cyber recovery, and platform... ...strategies across virtual, physical, database, cloud-native, and enterprise application environments...CloudSeniorPermanent employmentContract workLocal area- ...The Role Nium is looking for a Senior Manager, Site Reliability Engineering to lead the teams responsible for the... ...reliability budget: tooling investments, cloud cost/performance trade-offs, and... ...leadership. Represent SRE in executive reviews, translating technical...CloudSeniorFull timeWork at officeLocal areaWorldwideFlexible hours3 days per week
$175k - $215k
...exciting experiences.Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and their managers,... ...resilience and innovation. Influences senior internal and external... ...fault-tolerant systems across cloud (AWS, GCP, Azure) and on-prem...CloudSenior$165k - $225k
...bare-metal performance with cloud-native operational simplicity... ...workloads with enterprise-grade reliability and compliance. Your Role:... ...closely with our systems engineers, network engineers, and platform... ...Experience: 5+ years in SRE, DevOps, or infrastructure engineering...CloudSeniorRemote workFlexible hours$98.9k - $228.7k
What you can expectWe are hiring a Senior DevOps Engineer to ensure reliability, scalability, and operational excellence... ...media servers, CDN providers, cloud-native services, and edge networking... ...across teams.Acting as the primary SRE partner for multiple engineering...CloudSeniorFull timeWork at officeRemote workFlexible hours$152k - $241.5k
NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked hands... ....Experience managing production reliability through on-call duties, incident response...CloudSeniorPermanent employmentFull time- ...Sr Application Performance and Observability Engineer At Sequoia Connect, we are a Talent-First Technology Ecosystem that redefines... ...working at the absolute forefront of AI-driven automation and cloud solutions. This is your chance to thrive in a "Customer Success...CloudSeniorRemote workWorldwide
- ...forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of... ...: 6+ years as an SRE, DevOps Engineer, or similar role Cloud: Strong experience with AWS (EKS, EC2, S3, Route53, IAM)...CloudSeniorPermanent employmentWork experience placementLocal area
$175k - $215k
...Sr. Manager, Site Reliability Engineer At Disney, we're storytellers. We make... ...strategic leadership for multiple SRE teams, fostering a culture... ...-tolerant systems across cloud (AWS, GCP, Azure) and on-prem... ...with experience influencing senior stakeholders and driving cross...CloudSeniorLocal area$300k
...mode startup building out their AI and cloud platform, powered by thousands of H10... ..., or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability,... ...Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering...CloudSeniorPermanent employment- ...Detail Description: The AWS Site Reliability Engineer (SRE) is responsible for the operational health, availability, and performance of... ...SOPs). Skills: ~ CloudWatch, performance tuning in cloud environments, IaC tools, Databricks management and performance...Cloud
- ...Summary We are seeking a Senior SRE / DevSecOps Engineer with strong experience in... ...role will focus on platform reliability, incident management, SLO/... ..., and automation across cloud-native environments. Roles... ...SLO/SLI governance and site reliability practices. ~...CloudSeniorContract work
$500 per month
...Spain or Portugal. Key engineering and product teams are based... ...The Role We’re hiring a Senior Site Reliability Engineer to join our... ...sits at the intersection of cloud infrastructure, reliability,... ...make an impact. As an SRE at Maze, you will:...CloudSeniorRemote jobFull timeFlexible hours- ...Site Reliability Engineer (SRE) Immediate need for a talented Site Reliability Engineer (SRE). This is a 12+ months contract opportunity with long... ...and Technology Experience: ~ Must have skills: Azure Cloud Infrastructure. ~ SRE Tools / Technologies: Experience...CloudContract workLocal areaImmediate start
- ...match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-cloud infrastructure. You'll be the person who... ...workflows Partner closely with engineering on reliability reviews and architecture...CloudSenior
- ...Senior Site Reliability Engineer Company: Regrello Work Type: Remote Employment: Full Time Location: US Seniority:... ...Prometheus, Grafana, Splunk, Temporal, Celery, Cloud Run, BigQuery, GKE, Memorystore, CloudSQL, Google Cloud, Git, SRE tooling Requirements: Senior SRE with 4-8...CloudSeniorFull timeRemote work
- ...operational automation for SRE and DevOps environments... .... *Design and manage cloud infrastructure across... ...*Apply AI/ML, prompt engineering, and agentic AI concepts... ...to improve system reliability and operational processes... ...years of SRE, DevOps, or Site Reliability Engineering...CloudSeniorContract work
- ...selected candidate for this role to work on site in the specified location(s).We are seeking a Kafka Site Reliability Engineer to help build, operate, and continuously improve... ...Kafka environments across on-premises and cloud platforms while helping engineering teams deliver...CloudSeniorFull timeWork at office
- ...requirements. We're looking for a Senior SRE to take ownership of reliability on the marketplace side - joining an... ...platform. Mentoring a mid-level engineer and helping shape what the... ...grows back out. Working across cloud infrastructure, container orchestration...CloudSeniorContract workWork at officeRemote workShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Cloud Site Reliability Engineer (SRE). Be the first to apply!
- senior cloud data engineer United States
- cloud network architect United States
- cloud engineer United States
- cloud engineer remote United States
- java cloud engineer United States
- principal cloud computing engineer United States
- big data cloud engineer United States
- google cloud architect United States
- junior cloud engineer United States
- salesforce marketing cloud developer United States




