Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Incident IQ LLC

Site Reliability Engineer

Incident IQ North (Alpharetta)

About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.

Purpose: Incident IQ is committed to creating a future where every K-12 district operates with seamless efficiency. When operations are unified on a single platform, districts gain the clarity and control needed to build a stronger foundation for student success. We're focused on delivering the tools, support, and partnerships that help make that vision a reality.

Mission: Incident IQ is on a mission to eliminate the friction of disconnected systems and clunky workflows that slow schools down. We're reimagining the critical work that happens behind the scenes, bringing visibility, efficiency, and impact to the processes that keep classrooms running. By streamlining the complex, automating the routine, and surfacing the insights that matter most, we can create the conditions for educators to teach, students to thrive, and districts to shape the future of education.

Site Reliability Engineer (SRE) Overview:

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.

Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.

We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

Site Reliability Engineer (SRE) Responsibilities:

  • This role is hands-on from day one. Your initial focus will be:
    • SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
    • Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
    • Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
    • Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
    • Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
    • Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.
  • For example: in your first few days, you might stand up an SLO and a burn-rate alert in Grafana for our highest-traffic service. Within a couple of weeks, PagerDuty on-call is configured and the rotation is trained on incident command. That's the pace we operate at here: hours and days, not weeks.

Site Reliability Engineer (SRE) Requirements:

  • The tools below are what we run today. What matters more is the systems literacy and genuine curiosity about reliability that let you reason from first principles when something breaks in a way none of these tools have seen before:
    • Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required. How you acquired it (degree or a rigorous equivalent) is not, since it's what lets you diagnose a novel failure, not just operate a dashboard.
    • AI-Accelerated Execution (core requirement): You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it. This is not a bonus skill here; it's how we expect this role to operate.
    • Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about what you've actually shipped than the number itself.
    • SLI/SLO Methodology: Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
    • Observability Tooling: Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open-source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
    • Incident Management: Proven track record designing on-call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
    • Performance & Chaos Engineering: Hands-on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
    • Automation, Infrastructure & Cloud: Proficient in Python, Go, or Bash; hands-on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
    • Communication: Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something "can't" be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
    • Independence & Pace: You don't need hand-holding or check-ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.

What Success Looks Like:

By the end of your first quarter, core services have defined SLIs and SLOs, live in Grafana dashboards, and are backed by burn-rate-based alerting. A documented incident management process is operating end-to-end, run day to day by trained on-call engineers rather than by you personally, from detection through blameless postmortem. Over time, success looks like measurably reduced alert noise, faster Mean Time to Recovery (MTTR), and an engineering organization that trusts its reliability signals enough to make release and investment decisions based on them.

Bonus Points:

  • Experience standing up an SRE practice from zero to one ("founding SRE").
  • Experience with GitOps and just-in-time production access models.
  • Familiarity with eBPF-based auto-instrumentation (eBPF stands for extended Berkeley Packet Filter), such as Grafana Beyla or OpenTelemetry eB
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Alpharetta, GA vacancy
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Alpharetta, GA
    5 days ago
  •  ...unwavering security to responsibly propel the global lottery industry ever forward.Position SummaryWe are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Suggested
    Permanent employment
    Full time
    Work experience placement
    Local area

    Scientific Games Corporation

    Alpharetta, GA
    2 days ago
  • $128k - $216k

     ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...make a difference at Fiserv.Job TitleSr. Site Reliability EngineerAbout CloverClover is...  ...does a successful Senior Site Reliability Engineer do at Fiserv?As a Senior Site Reliability... 
    Suggested
    Full time
    Worldwide

    Fiserv

    Alpharetta, GA
    2 days ago
  •  ...communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible...  ...across the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of industry... 
    Suggested
    Flexible hours
    Weekend work

    Morgan Stanley

    Alpharetta, GA
    5 days ago
  • $86.6k - $144.4k

     ...platforms and using automation to solve complex security and reliability challenges? Do you enjoy shaping the future of...  ...more about LexisNexis Risk at About our Team Our Site Reliability Engineering (SRE) team plays a critical role in ensuring the reliability... 
    Suggested
    Local area

    RELX Group plc

    Alpharetta, GA
    3 days ago
  •  ...initiatives. Collaborating with product, architecture, and engineering groups to build a platform that streamlines application...  ...Experience: 5+ Years advanced level experience in DevOps, Site Reliability Engineering with expertise in Enterprise Cloud infrastructure... 

    Software Technology Inc

    Alpharetta, GA
    5 days ago
  •  ...be highly motivated, creative, self-directed, and thrive in small project teams. Responsibilities: • Lead a team of software engineers and QA. • Architect, design, develop micro services. • Develop unit and integration tests aligned with the automation platforms... 
    Contract work

    My3Tech Inc

    Alpharetta, GA
    5 days ago
  •  ...I have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite). and I am looking for a candidate who can join Immediately if you are interested, reply to me with your updated resume or if you could refer someone I would really appreciate it. Role... 
    Immediate start
    Relocation

    Navtech

    Alpharetta, GA
    5 days ago
  •  ...Job Title: Site Reliability Engineer (SRE) Experience Required: 8+ Years Industry: Banking / Financial Services Job Description We are seeking an experienced Site Reliability Engineer (SRE) to join our banking client’s technology team. The ideal candidate... 

    Tech M USAAvance Consulting

    Alpharetta, GA
    2 days ago
  • $125k - $175k

     ...our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job family responsible for overseeing the production environment, ensuring... 
    Temporary work

    Morgan Stanley

    Alpharetta, GA
    3 days ago
  •  ...and shape the future of our communities.This is a Software Engineering position at Director level, which is part of the job family...  ...our businesses. This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support,... 

    Morgan Stanley

    Alpharetta, GA
    3 days ago
  • $129k - $161k

     ...Job Description Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote  Grade: 20 About Priority Commerce:  Priority Commerce is a leading financial... 
    Remote work

    Priority Technology Holdings, LLC

    Alpharetta, GA
    21 days ago
  • $125k - $175k

     ...overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize...  ...in quantitative discipline (Computer Science, Computer Engineering). - 5+ years’ experience in leading a small to medium team of... 
    Full time
    Temporary work

    Morgan Stanley

    Alpharetta, GA
    a month ago
  •  ...of your work. You are visible, your talents are valued, and you are empowered to shape the future of payments.As a Senior Site Reliability Engineer (SRE) - Azure & GitOps (CI/CD) in Norcross, GA or Omaha, NE, you will join a diverse, passionate team, dedicated to powering... 
    Full time
    Local area
    Worldwide

    ACI Worldwide

    Norcross, GA
    2 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability... 

    2T Consulting

    Lilburn, GA
    5 days ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 
    Work experience placement

    AmNet Services

    Alpharetta, GA
    1 day ago
  •  ...Job Description Job Description We are hiring an SRE Platform Engineer (IBM BPM/ODM) with our partner in Alpharetta, GA for an onsite role. Job Details: Job Title: SRE Platform Engineer (IBM BPM/ODM) Location- Alpharetta, GA (Onsite) Interview- F2F (Locals... 
    Local area

    Appex Innovation

    Alpharetta, GA
    8 days ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 

    AmNet Services

    Alpharetta, GA
    1 day ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 

    AmNet Services

    Alpharetta, GA
    1 day ago
  •  ...passionate and love what we do! We are at the forefront of future engineering technologies, with solutions that ensure the success of our...  ...shaping the future of the world we live in.Job Description: Reliability EngineerJob Location: Hartford City, INTravel: Approximately... 
    Live in
    Work at office

    Andritz

    Alpharetta, GA
    1 day ago
  • $125.7k - $203.1k

     ...from around the world, working together in Engineering, IT, Supply Chain, Customer Experience,...  ...interactions and ensure product reliability. Apply systems programming expertise to...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Full time
    Temporary work
    Apprenticeship
    Work experience placement
    Local area
    Flexible hours

    CISCO Systems

    Alpharetta, GA
    1 day ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid – large wireless... 

    Amnet

    Alpharetta, GA
    5 days ago
  • $101.6k - $152.4k

    We are seeking a talented Senior Engineer I, Digital Solutions to join our team and take charge of designing, developing, and deploying...  ...alarms, and reports.Travel: Willingness to travel to customer sites as required. Travel is roughly expected to be around 25% but is... 
    Full time
    Temporary work
    Immediate start
    Remote work
    Work from home
    Flexible hours

    Schneider Electric

    Alpharetta, GA
    5 days ago
  • $166.84k - $211.9k

     ...external components, systems, and platforms.REQUIREMENTS: Requires a Bachelor’s degree, or foreign equivalent degree in Computer Engineering, Computer Science, Applied Science, Electrical Engineering, or Math and 5 years of progressive, postbaccalaureate experience in the... 
    Temporary work
    Local area

    AT&T

    Alpharetta, GA
    2 days ago
  •  ...storage tanks, water metering, energy metering, gas monitoring, and asset management. Our founders are hardcore telecommunications engineers with combined 200 + years of experience in designing, optimizing and performance engineering; for several mid - large wireless... 

    AmNet Services

    Alpharetta, GA
    1 day ago
  • $162.4k - $211.9k

     ...JOB TITLE: Lead System Engineering JOB LOCATION: 500 North Point Pkwy., Alpharetta, GA 30022 DUTIES: Responsible for translating...  ...software, cloud, etc.) through functional, performance, and reliability analysis using engineering models and techniques, primarily through... 
    Temporary work
    Local area

    AT&T

    Alpharetta, GA
    2 days ago
  •  ...from supervisor Develop and design web applications and web sites. Responsible for directing web site content creation, enhancement...  ...Experience ~ Bachelor’s degree in computer science, software engineering or relevant field required ~6+ years’ experience required... 
    Full time

    Collabera

    Alpharetta, GA
    1 day ago
  •  ...Choose from a   wide range of exciting opportunities   from our diverse Fortune 500 client base. Job Description • Measure LNRS site usability/effectiveness and present actionable insights and recommendations based on data results and best practices on a regular... 
    Full time

    Collabera

    Alpharetta, GA
    1 day ago
  •  ...is designed & implemented successfully, to enable application reliability/ availability, application quality and performance requirements...  ...internal SQL server databases, mobile apps and other Share Point sites Roles & Responsibilities: • Proficient in designing,... 
    Full time
    Local area
    Flexible hours

    Crowdstaffing

    Alpharetta, GA
    1 day ago
  • Senior Software Engineer - Strong knowledge of Java and Python. Alpharetta GA \ \ Mandatory Areas Must Have Skills – Skill...  ...data processing. Ensure high performance, scalability, and reliability of backend systems. Integrate and manage backend services... 
    Full time

    Gov Services Hub

    Alpharetta, GA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!