Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer I

$98k - $148.5k

PagerDuty Inc.

PagerDuty (NYSE:PD) is a leader in Digital Operations Management. In an always-on world, organizations of all sizes trust PagerDuty to help them deliver a perfect digital experience to their customers, every time. Teams use PagerDuty to identify issues and opportunities in real time and bring together the right people to fix problems faster and prevent them in the future. Over 13,000 organizations (including 60 of Fortune 100) rely on PagerDuty to succeed with Digital Transformation, Cloud Migration, and DevOps Modernization. Notable customers include GE, Cisco, Genentech, Electronic Arts, Cox Automotive, Netflix, Shopify, Zoom, DoorDash, Lululemon and more. We are expanding rapidly as a platform for Digital Operations Management using AI/ML and Automation and growing our adoption by Development, IT, Customer Service, Security, and other teams across the organization.As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office,you'll help build and operate the foundational infrastructure that powers PagerDuty'sreal-time digital operations platform. Our systems support millions of events and alertsdaily, enabling customers to detect, respond to, and resolve incidents quickly andreliably.You'll work at the intersection of platform evolution and operational excellence, buildingand evolving foundational network, compute, and ingress infrastructure while scalingand hardening existing systems. Your work will directly impact the reliability, scalability,and security of the services our customers rely on to keep their businesses running asPagerDuty continues to grow across products, regions, and customer use cases.Key Responsibilities Support and improve foundational infrastructure, including networking, computeplatforms, Kubernetes clusters, and ingress/traffic management systems. Contribute to the reliability and scalability of PagerDuty's core platform byhardening existing systems and supporting the rollout of new infrastructurecapabilities. Participate in agile rituals (standups, planning, retros) and communicateprogress/risks early You stay current on technical trends to suggest innovative tools and approachesto interesting problems Monitor system health using metrics, logs, and alerts, and participate in 24/7on-call rotations to help detect, respond to, and resolve incidents.Basic Qualifications 0 to 1+ years of experience in Site Reliability Engineering, DevOps, or PlatformEngineering roles Hands-on experience operating Linux-based systems in productionenvironments Working knowledge of networking fundamentals, such as load balancing, DNS,TLS, and ingress traffic flow Experience with container orchestration (e.g., EKS, Kubernetes) Experience working on cloud-native infrastructure (e.g., AWS, GCP, Azure),including networking and compute concepts Proficiency in at least one programming language (e.g., Python, Ruby, Go, etc.) Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)Preferred Qualifications Experience with AWS cloud networking concepts such as VPCs, subnets,routing, security groups, and load balancers Experience operating or contributing to production Kubernetes platforms (e.g.,EKS), including cluster upgrades, networking, or ingress configuration Experience with monitoring, observability, and logging platforms (e.g., DataDog,New Relic, SumoLogic, Splunk, Prometheus, Grafana) Familiarity with service meshes, ingress controllers, or API gateways (e.g.,Envoy, Istio, NGINX)Salary Range: $98,000 to $148,500Hesitant to apply?We encourage you to submit your resume even if you don't meet every requirement. We value potential and consider each candidate's full professional story. Whether you're exploring a career change or taking your next step, we look forward to reviewing your application. If this just isn’t the right role or time - sign up for job alerts!Where we workPagerDuty operates a hybrid work model with offices in 8 major cities: Atlanta, Lisbon, London, San Francisco, Santiago, Sydney, Tokyo, and Toronto. While we offer flexibility within our established locations, we cannot employ candidates residing in:Location restrictions: Australia: Northern Territory, Queensland, South Australia, Tasmania, Western AustraliaCanada: Alberta, Manitoba, Newfoundland, Northwest Territories, Nunavut, PEI, Quebec, Saskatchewan, YukonUnited States: Alaska, Hawaii, Iowa, Louisiana, Mississippi, Nebraska, New Mexico, Oklahoma, Rhode Island, South Dakota, West Virginia, WyomingCandidates must reside in an eligible location, which vary by role.How we workOur values guide how we support customers, collaborate with colleagues, develop products, and foster a culture of belonging. They define not just our actions, but what it means to be Dutonian.People Leaders at PagerDuty are responsible for creating high performance environments that drive accountability. PagerDuty has four key dimensions that define our Leadership Impact: Lead Self, Lead the Team, Lead the Business, and Lead the Future. Each dimension has three associated competencies to give leaders a shared language for guiding their development, career, promotion, and succession planning discussions. Our Manager Expectations serve as a practical guide for managers to understand their responsibilities, prioritize their efforts, and drive engagement and performance.What we offerAs a global organization, our total rewards approach is competitive with industry standards and aligned with local laws and regulations. Learn more, including country-specific offerings, on our benefits site.Your package may include:Competitive salaryComprehensive benefits package Flexible work arrangementsCompany equity*ESPP (Employee Stock Purchase Program)*Retirement or pension plan*Generous paid vacation timePaid holidays and sick leaveDutonian Wellness Days & HibernationDuty - companywide paid days off in addition to PTOPaid parental leave: 22 weeks for pregnant parent, 12 weeks for non-pregnant parent (some countries have longer leave standards and we comply with local laws)*Paid volunteer time off: 20 hours per yearCompany-wide hack weeksMental wellness programs*Eligibility may vary by role, region, and tenureAbout PagerDutyPagerDuty, Inc. (NYSE:PD) is a global leader in digital operations management. The PagerDuty Operations Cloud is an AI-powered platform that empowers business resilience and drives operational efficiency for enterprises. With a generative AI assistant at its core, PagerDuty empowers teams to detect and resolve issues in real time, orchestrate complex workflows, and drive continuous improvement across their digital operations. Trusted by nearly half of both the Fortune 500 and the Forbes AI 50, as well as approximately two-thirds of the Fortune 100, PagerDuty is essential for delivering always-on digital experiences to modern businessesPagerDuty is Great Place to Work-certified, a Fortune Best Workplace for Millennials, a Fortune Best Medium Workplace, a Fortune Best Workplace in Technology, and a top rated product on TrustRadius and G2. Go behind-the-scenes on our careers site and @pagerduty on Instagram.Additional InformationPagerDuty is an equal opportunity employer. PagerDuty does not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, parental status, veteran status, or disability status. Your privacy is important to us. By submitting an application, you confirm that you have read and understand PagerDuty's Privacy Policy.PagerDuty is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application process. Should you require accommodation, please email View email address on click.appcast.io and we will work with you to meet your accessibility needs.PagerDuty uses the E-Verify employment verification program.

Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer I in Atlanta, GA vacancy
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Suggested
    Worldwide

    Inspire Brands

    Atlanta, GA
    4 hours ago
  • $35 - $44 per hour

    DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and... 
    Suggested
    Remote work

    KForce

    Atlanta, GA
    4 days ago
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineer role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises... 
    Suggested
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Atlanta, GA
    4 hours ago
  •  ...System Reliability Engineer (SRE)At T-Mobile, we invest in YOU! Our Total Rewards Package ensures that employees get the same big love we give our customers. All team members receive a competitive base salary and compensation package - this is Total Rewards. Employees... 
    Suggested
    Full time
    Temporary work
    Part time
    Work experience placement
    Flexible hours

    T Mobile US

    Atlanta, GA
    1 day ago
  •  ...resilient software platforms using SRE and AI-native engineering practices. Own production reliability, monitoring, and operational automation while mentoring...  ...), Kubernetes, and production operations. Key Skills Site Reliability Engineering Terraform Python AWS Azure... 
    Suggested
    Temporary work
    Flexible hours

    T-Mobile

    Atlanta, GA
    1 day ago
  • CBRE Group, Inc. seeks an Operations Consultant to provide technical expertise across complex technology systems and support daily operations within the D&T Support function. You will collaborate with IT teams, optimize configurations, and guide stakeholders on technology...

    CBRE Group, Inc.

    Atlanta, GA
    4 days ago
  • $70 - $85 per hour

     ...redefine what’s possible, give shape to the future—and get there.What You’ll Do* Define and establish enterprise reliability standards, Site Reliability Engineering (SRE) practices, SLIs/SLOs, operational governance models, and engineering guardrails that enable scalable,... 
    Temporary work
    Local area
    Flexible hours
    3 days per week

    Slalom

    Atlanta, GA
    7 hours ago
  • $167.7k - $245.2k

     ...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Atlanta, GA
    4 days ago
  •  ...Technology Consultant - Site Reliability Engineer (SRE) Location- - Atlanta, GA (hybrid schedule) Client is currently seeking an experienced Technology Consultant Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability... 
    Permanent employment

    Stellent IT LLC

    Atlanta, GA
    2 days ago
  •  ...evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical...  ...highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles... 
    Early shift

    Cloud Analytics Technologies LLC

    Atlanta, GA
    2 days ago
  • $151k - $297k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    4 days ago
  • #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...

    LTM

    Atlanta, GA
    2 days ago
  •  ...out new infrastructure capabilities to improve platform reliability and scalability. Monitor system health using metrics,...  ...approaches. Requirements At least 3 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles. Hands-on... 
    Full time
    Work at office
    Local area
    Flexible hours

    PagerDuty

    Atlanta, GA
    13 hours ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's...  ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    Atlanta, GA
    3 days ago
  •  ...can create the conditions for educators to teach, students to thrive, and districts to shape the future of education. Site Reliability Engineer (SRE) Overview:   We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it... 
    Full time
    Live in
    Work at office

    Incident IQ

    Atlanta, GA
    8 days ago
  •  ...Consultancy and Information Technology Enabled Services.Job DescriptionSCM System EngineerSCM Continuous Integration / Delivery Build Team Engineer with experience in Application Service and Web Application Build, Deployment and Release Management and experience in establishing... 
    Permanent employment
    Full time
    H1b

    Career Guidant

    Atlanta, GA
    2 days ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Atlanta, GA
    8 days ago
  • $141.3k - $237.4k

     ...AT&T, you won’t just imagine the future, you’ll build it.We are seeking a highly skilled and hands-on Lead Software Engineer to join Software Reliability Engineering (SRE) Onboarding and automation team. This role will drive innovation through automation, enhancement of... 
    Full time
    Temporary work
    Work at office
    Local area
    Relocation

    AT&T

    Atlanta, GA
    3 days ago
  • $101.5k - $169.1k

     ...include an incentive program.Job DescriptionThe Release Train Engineer (RTE) has a primary purpose of supporting an Agile Release Train...  ...organizational AI policies and standards. Monitor AI tool reliability across teams. Create backup plans for system failures. Maintain... 
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    Cox Enterprises

    Atlanta, GA
    7 hours ago
  • $105k - $130k

     ...provide the high-speed capabilities our nation and its allies need to maintain a durable, asymmetric advantage. The Mission Systems Engineering (MSE) Team develops the Mission Management System (MMS)—a software platform that integrates mission subsystems, autonomy services... 
    Weekly pay
    Permanent employment
    Full time
    Work at office

    Hermeus

    Atlanta, GA
    1 day ago
  •  ...an impact, and work with people who care, we'd love to meet you! ABOUT THE ROLE We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering... 
    Local area
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    AgileEngine

    Atlanta, GA
    3 days ago
  •  ...YesApply: JobGeorgia-Pacific's Corrugated Division is seeking a Reliability Engineer to support our growing Mailers network, including facilities...  ...programs, and driving standardization across multiple sites. Success in this position requires strong technical capability... 
    For contractors
    Remote work
    Worldwide
    Visa sponsorship
    Flexible hours

    Georgia-Pacific

    Atlanta, GA
    4 days ago
  • $126k - $167k

     ...autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.ABOUT THE TEAMThe Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Atlanta, GA
    1 day ago
  •  ...coaches. That’s how we’re UNSTOPPABLE for our employees!Are you ready for the next chapter in your Uncarrier journey? The System Reliability Engineer (SRE) improves and protects the software and systems behind all of T-Mobile's IT services, including management of... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    1 day ago
  • $168.5k - $252.7k

     ...secure.About the RoleAs a Senior Software Engineer, you will play a key role in designing...  ...mentor team members to ensure high-velocity, reliable delivery.What You’ll DoDesign, build, and...  ...not Workday Careers. Please be aware of sites that may ask for you to input your data... 
    Full time
    Contract work
    Work at office
    Remote work
    Home office
    Flexible hours

    Workday

    Atlanta, GA
    4 days ago
  • $144k - $191k

     ...customers. We are looking for software engineers, hardware engineers, roboticists, and front...  ...-in-the-loop demonstrations at test sites.Define a strategy to comply with appropriate...  ...(V&V) plans to ensure robust and reliable system performance.Experience writing testable... 
    Full time
    Work experience placement
    Work at office
    Immediate start

    Anduril Industries

    Atlanta, GA
    7 hours ago
  • OverviewJob PurposeIntercontinental Exchange (ICE) presents a unique opportunity for a full-time Engineer to join a team responsible for ICE OpenShift container platform. The Engineer will drive enablement and adoption to migrate high performant and critical applications... 
    Full time

    Black Knight Financial Services

    Atlanta, GA
    7 hours ago
  • OverviewJob PurposeIntercontinental Exchange presents a unique opportunity for a full-time Engineer to join a team responsible for ICE OpenShift container platform. The Engineer will drive enablement and adoption to migrate high performant and critical applications into... 
    Full time

    Black Knight Financial Services

    Atlanta, GA
    2 days ago
  •  ...Sr Release Train Engineer Cox Automotive is seeking a talented Senior Release Train Engineer to join its team. The Senior Release Train Engineer (RTE) is an outcomes-driven problem solver whose primary purpose is to lead large, complex Agile Release Trains (ARTs) to... 
    Live in
    Work at office
    Visa sponsorship
    Flexible hours

    Cox Enterprises

    Atlanta, GA
    3 days ago
  •  ...teamsPropose improvement ideas for the department roadmapProvide on-site support at the client location as neededModel Exotec's values...  ...A minimum of 8+ years of experience in a similar software engineering or software integration role, ideally in the warehousing, fulfillment... 
    Full time
    Work at office
    Local area

    Exotec

    Atlanta, GA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer I. Be the first to apply!