Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Site Reliability Engineer

PowerPlan

Overview

PowerPlan is seeking a Principal Site Reliability Engineer to sit at the heart of our cloud platform's reliability, scalability, and operational maturity. You'll work hands-on across AWS and Azure environments, solving complex production problems while systematically eliminating the manual toil that creates them.

This role offers significant autonomy, deep technical impact, and the opportunity to shape how reliability engineering is practiced across the organization.

About PowerPlan

PowerPlan helps the companies that power the world unlock greater value from their infrastructure investments. We combine deep industry expertise, trusted technology, and AI-driven innovation to help capital-intensive organizations manage the financial complexity of their assets with confidence.

We're in the middle of one of the most significant transformations in our history. Decades of proven functionality are being reimagined on a modern SaaS foundation, while AI becomes central to how we build, deliver, and evolve our products. The next generation of the PowerPlan platform is being designed right now. Join us, and you'll have the opportunity to influence the technology, engineering practices, and capabilities that will shape it for years to come.

Responsibilities
Your Impact
  • Resolve escalated infrastructure cases across AWS and Azure and deliver targeted automations that reduce manual resolution time.
  • Eliminate or significantly reduce manual intervention for the highest-frequency operational issues through automation and tooling.
  • Establish a consistent, high-quality incident response and post-incident review process for critical production incidents.
  • Deliver a mature, SLO-aligned observability platform with dashboards, tuned alerts, and clear reporting.
  • Coach teams on effective incident communication and decision-making.
What Success Looks Like

Within 90 days, you'll have shipped automations that measurably cut manual resolution time on recurring issues. By month six, you'll have eliminated toil on the highest-frequency operational problems. By month twelve, on-call and engineering teams will be running on an observability platform you built — one with dashboards, tuned alerts, and SLIs/SLOs that make reliability a data-driven practice instead of a guessing game.

Qualifications
What You'll Bring
  • Deep hands-on experience operating production systems in AWS and Azure environments.
  • Strong automation skills using Python and PowerShell in operational contexts.
  • Proven ability to identify repetitive operational work and eliminate it through automation.
  • Experience leading incident response and blameless post-incident reviews.
  • Strong observability expertise, particularly with Grafana and SLI/SLO-driven monitoring.
  • Ability to influence engineering practices without formal authority.
  • Clear written and verbal communication skills across technical and non-technical audiences.
Education & Experience
  • Extensive experience in cloud operations, site reliability engineering, or infrastructure engineering roles, or equivalent professional experience.

PowerPlan is an EOE

Applicant and Candidate Privacy Notice

Please note that this is a hybrid role that involves a combination of onsite work from our corporate office as well as work from home. While we strive to accommodate flexible working arrangements when sensible, there will be times when onsite work is required. This could include scheduled office days, team meetings, client meetings, or special events.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal Site Reliability Engineer in Atlanta, GA vacancy
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud... 
    Suggested
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Atlanta, GA
    3 days ago
  • $35 - $44 per hour

    DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and... 
    Suggested
    Remote work

    KForce

    Atlanta, GA
    2 days ago
  • $113k - $171.6k

     ...Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office,you'll help build and operate the foundational infrastructure... 
    Suggested
    Work at office
    Local area
    Flexible hours

    PagerDuty

    Atlanta, GA
    3 days ago
  • $100k - $120k

    OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Origami Risk

    Atlanta, GA
    2 days ago
  •  ...ideal time and number for communication, and the expected pay rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months Contract Visa : US Citizens/ Green Card Need Local to... 
    Suggested
    Contract work
    Local area
    Immediate start

    Navtech

    Atlanta, GA
    2 days ago
  • $120k - $175k

     ...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues... 
    Full time
    Remote work
    Work visa
    Flexible hours

    PrizePicks

    Atlanta, GA
    4 days ago
  • $70 - $85 per hour

     ...redefine what’s possible, give shape to the future—and get there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance models, and engineering guardrails that enable scalable,... 
    Temporary work
    Local area
    Flexible hours
    3 days per week

    Slalom

    Atlanta, GA
    3 days ago
  •  ...industry leading Digital Platform which will power its 7 existing brands and enable smooth integration of future brands. The Principal Software Engineer will oversee the solution architecture and development of the backend API components of the platform that powers our... 
    Principal

    Focus Brands

    Atlanta, GA
    3 days ago
  • $127k - $249k

     ...Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the Atlas...  ...workloads. Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background. This... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    5 days ago
  •  ...Purple Drive Site Reliability Engineer (SRE) Contractual Atlanta, GA Key Highlights: Proven expertise in Google Cloud Platform (GCP) services, including BigQuery, Cloud Logging, IAM, and Service Accounts. Strong background in provisioning, monitoring, and... 

    Purple Drive

    Atlanta, GA
    4 days ago
  •  ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively... 
    Work experience placement

    Fisec Global

    Atlanta, GA
    5 days ago
  • $55 - $60 per hour

     ...Senior Reliability Engineer Required Skills & Experience• 5+ years of experience in SRE, DevOps, Cloud Operations, or Infrastructure Engineering. Strong expertise in Azure, GCP, Kubernetes, Cloudflare, Splunk, AppDynamics, SQL Server, RabbitMQ, and Zscaler. Solid networking... 

    Insight Global

    Atlanta, GA
    2 days ago
  • $167.7k - $245.2k

     ...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Atlanta, GA
    2 days ago
  • $111.61k - $131.3k

    At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life,...
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Atlanta, GA
    1 day ago
  • The Home Depot is seeking a Senior Principal Software Engineer in Reliability Engineering to design resilience into our store systems, payments, and COM platform. You will lead multiple engineering teams, drive cost-efficient cloud usage, and advocate best practices across... 
    Principal

    The Home Depot

    Atlanta, GA
    5 days ago
  • Position Purpose We're seeking a Software Engineer Senior Principal in Reliability Engineering to be the strategic force behind our store systems, payments and COM platform's stability. This isn't about patching things; it's about fundamentally designing resilience into... 
    Principal
    Remote job
    Work experience placement
    Local area
    Night shift

    The Home Depot

    Atlanta, GA
    5 days ago
  •  ...Role: Site Reliability Engineering (SRE) Architect Location: Atlanta, GA (Hybrid on-site) Contract Role Summary: As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems... 
    Contract work
    Early shift

    AceStack LLC

    Atlanta, GA
    2 days ago
  •  ...speed capabilities our nation and its allies need to maintain a durable, asymmetric advantage.About The Role:The Mission Systems Engineering (MSE) Team develops the Mission Management System (MMS)—a software platform that integrates mission subsystems, autonomy services... 
    Principal
    Weekly pay
    Permanent employment
    Full time
    Work at office

    Hermeus

    Atlanta, GA
    3 days ago
  •  ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview:   We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it... 
    Full time
    Live in
    Work at office

    Incident IQ

    Atlanta, GA
    8 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 
    Temporary work

    2T Consulting

    Atlanta, GA
    8 days ago
  • $143k

     ...thrive.What You'll DoJoin our Data Platform Engineering portfolio, a global team building and...  ...Areas, and Business Systems.As a Principal Database Platform Engineer, you will take...  ...role focused on designing for scale and reliability, driving automation, and setting the platform... 
    Principal
    Work at office
    Local area

    The Boston Consulting Group

    Atlanta, GA
    2 days ago
  • $228k - $342k

     ...product leaders, machine learning engineers, and full-stack builders to...  ....About the RoleAs a Senior/Principal Machine Learning Engineer in...  ...turn them into systems that are reliable, explainable, and built to...  ...Careers. Please be aware of sites that may ask for you to input... 
    Principal
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Atlanta, GA
    4 days ago
  •  ...systems, and policy-as-code enforcement engines that integrate self-service developer experiences...  ...tolerance patterns that ensure platform reliability at enterprise scale. Applies modern...  ...plans, please visit our Benefits site. Depending on the position and division,... 
    Principal
    Full time
    Part time
    Shift work
    Day shift

    Truist

    Atlanta, GA
    5 days ago
  • Job DescriptionPrincipal Consultant - AI Platform & Governance Engineer - Infosys ConsultingInfosys Consulting's Tech Transformation Advisory Practice is seeking a Principal Consultant specializing in AI Platform & Governance EngineerPosition Overview:As a Principal Consultant... 
    Principal
    Full time
    Temporary work
    Local area

    Infosys Technologies

    Atlanta, GA
    2 days ago
  • $148.5k - $313.7k

     ...By applying to the Senior / Lead Software Engineer - Full Stack posting, recruiters and...  ...emerging AI capabilities into practical, reliable product experiences. That means choosing...  ...quickly.Benefits & PerksCheck out our benefits site which explains our various benefits,... 
    Principal
    Full time
    Remote work

    Salesforce

    Atlanta, GA
    1 day ago
  • $196.5k - $294.75k

     ...society.   The Challenge We are looking for a Principal-level, US-based, customer-facing engineer who is deeply hands-on with large-scale,...  ...customers through best practices for scalability, reliability, cost efficiency, and observability when using our... 
    Principal
    Full time
    Work experience placement
    Work at office
    Local area
    Worldwide
    Flexible hours
    3 days per week
    1 day per week

    Onetrust

    Atlanta, GA
    10 hours ago
  • $117.2k - $313.7k

     ...Salesforce.The ExperienceNote: By applying to the Senior / Lead / Principal Software Engineer - Foundations Team posting, recruiters and hiring...  ...talent-density team that cares deeply about performance, reliability, and engineering excellence — driving architecture... 
    Principal
    Full time
    Work experience placement
    Remote work

    Salesforce

    Atlanta, GA
    1 day ago
  • $250.6k - $362.6k

     ...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable capabilities...  ...networking, with a strong emphasis on reliability, interoperability, and long-term...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Principal
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Atlanta, GA
    5 days ago
  • $107.5k - $204.5k

     ...Glenlake Pkwy ~ GLENLAKE, Ste 650 (External Site)Country: United States of AmericaTime...  ...100 years of experience and renowned engineering expertise to meet the needs of today’s mission...  ...Test Capability (SE&TC) Center seeks a Principal Electronic Warfare (EW) Lead System... 
    Principal
    Temporary work
    Work experience placement
    Internship
    Work at office
    Remote work
    Relocation package
    Flexible hours

    Raytheon

    Atlanta, GA
    2 days ago
  •  ...Principal Developer They’ll be working in Salesforce so I’m looking for full stack experience on Salesforce specifically (it sounds like Salesforce has their own customized coding language) – no need to have React/Node/cloud. If they did have broader experience I would... 
    Principal
    Immediate start
    Remote work

    Kaav Inc.

    Atlanta, GA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!