Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Incident IQ LLC

Site Reliability Engineer

Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.

Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.

We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

This role is hands-on from day one. Your initial focus will be:

  • SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
  • Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
  • Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
  • Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
  • Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
  • Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.

The tools below are what we run today. What matters more is the systems literacy and genuine curiosity about reliability that let you reason from first principles when something breaks in a way none of these tools have seen before:

  • Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required. How you acquired it (degree or a rigorous equivalent) is not, since it's what lets you diagnose a novel failure, not just operate a dashboard.
  • AI-Accelerated Execution (core requirement): You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it. This is not a bonus skill here; it's how we expect this role to operate.
  • Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about what you've actually shipped than the number itself.
  • SLI/SLO Methodology: Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
  • Observability Tooling: Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open-source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
  • Incident Management: Proven track record designing on-call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
  • Performance & Chaos Engineering: Hands-on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
  • Automation, Infrastructure & Cloud: Proficient in Python, Go, or Bash; hands-on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
  • Communication: Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something "can't" be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
  • Independence & Pace: You don't need hand-holding or check-ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.

By the end of your first quarter, core services have defined SLIs and SLOs, live in Grafana dashboards, and are backed by burn-rate-based alerting. A documented incident management process is operating end-to-end, run day to day by trained on-call engineers rather than by you personally, from detection through blameless postmortem. Over time, success looks like measurably reduced alert noise, faster Mean Time to Recovery (MTTR), and an engineering organization that trusts its reliability signals enough to make release and investment decisions based on them.

Bonus Points:

  • Experience standing up an SRE practice from zero to one ("founding SRE").
  • Experience with GitOps and just-in-time production access models.
  • Familiarity with eBPF-based auto-instrumentation (eBPF stands for extended Berkeley Packet Filter), such as Grafana Beyla or OpenTelemetry eBPF Instrumentation, for legacy or hard-to-modify codebases. It's a newer approach, nice to have rather than expected.
  • .NET experience is a plus, given our engineering stack.
  • Certifications aren't required and aren't a strong signal for us; what you've built matters more than what's on your cert wall. If you happen to have one, Certified Kubernetes Administrator (CKA) or Google Cloud Professional DevOps Engineer are the most relevant.

What makes Incident IQ different:

  • We facilitate whole-person growth where employees can develop personally as well as professionally.
  • We offer an energetic and collaborative environment; everyone's opinion matters!
  • We produce software that empowers K-12 schools to run efficiently, allowing for a better classroom experience for students to THRIVE!
  • We provide excellent work/life balance. Two amazing offices - a Downtown Atlanta office location and one at Halcyon in Alpharetta!

Incident IQ offers a competitive salary based on experience with a benefits package for full-time employees that includes medical, dental, vision, life insurance, 401k match, and paid-time off (PTO).

Incident IQ is an Equal Opportunity Employer

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Alpharetta, GA vacancy
  •  ...I have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite). and I am looking for a candidate who can join Immediately if you are interested, reply to me with your updated resume or if you could refer someone I would really appreciate it. Role... 
    Suggested
    Immediate start
    Relocation

    Navtech

    Alpharetta, GA
    2 days ago
  •  ...Job Title: Site Reliability Engineer (SRE) Experience Required: 8+ Years Industry: Banking / Financial Services Job Description We are seeking an experienced Site Reliability Engineer (SRE) to join our banking client’s technology team. The ideal candidate... 
    Suggested

    Tech M USAAvance Consulting

    Alpharetta, GA
    4 days ago
  •  ...initiatives. Collaborating with product, architecture, and engineering groups to build a platform that streamlines application...  ...Experience: 5+ Years advanced level experience in DevOps, Site Reliability Engineering with expertise in Enterprise Cloud infrastructure... 
    Suggested

    Software Technology Inc

    Alpharetta, GA
    2 days ago
  • $128k - $216k

     ...consumers to one another millions of times a day - quickly, reliably, and securely. Any time you swipe your credit card, pay...  ...scale, come make a difference at Fiserv. Job Title Sr. Site Reliability Engineer About Clover Clover is a pioneer in the fintech... 
    Suggested
    Worldwide

    Fiserv

    Alpharetta, GA
    1 day ago
  •  ...communities.This is a Lead Software Production Management & Reliability Engineering position at Director level which is part of the job family responsible...  ...across the business.Job SummaryWe are looking for a Site Reliability Engineer with a minimum of 5 years of industry... 
    Suggested
    Part time
    Flexible hours
    Weekend work

    Morgan Stanley

    Alpharetta, GA
    3 hours ago
  •  ...and shape the future of our communities.This is a Software Engineering position at Director level, which is part of the job family...  ...our businesses. This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support,... 
    Part time

    Morgan Stanley

    Alpharetta, GA
    3 hours ago
  • $165k - $241.4k

     ...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze... 
    Permanent employment
    Full time
    Temporary work
    Part time
    Local area
    Flexible hours

    CISCO Systems

    Alpharetta, GA
    3 hours ago
  • $104.9k - $174.7k

     ...Management. You can learn more about LexisNexis Risk at the link below, About the Role: We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Work at office
    Local area
    Remote work
    Flexible hours

    RELX

    Buford, GA
    2 days ago
  •  ...security to responsibly propel the global lottery industry ever forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with... 
    Permanent employment
    Work experience placement
    Local area

    SCIENTIFIC GAMES

    Alpharetta, GA
    more than 2 months ago
  • $129k - $161k

     ...Job Description Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20 About Priority: Priority Technology Holdings, Inc. is a leading... 
    Remote work

    Priority Technology Holdings, LLC

    Alpharetta, GA
    18 days ago
  •  ...We are looking for a SRE/DevOps Engineer A highly technical, hands-on engineer to join a critical engineering organization supporting...  ...Python. Improve software delivery through DevOps and Site Reliability Engineering best practices. Support Infrastructure as... 

    ATech Placement

    Alpharetta, GA
    1 day ago
  • $63.22k - $94.55k

     ...maintain Windows Server environments with a focus on performance, reliability, and scalability Develop and maintain PowerShell scripts...  ...excellence and process optimization Mentor junior engineers and contribute to knowledge sharing across the team Required... 
    Remote work

    NTT DATA

    Alpharetta, GA
    13 hours ago
  • Syms Strategic Group (SSG)  is seeking a talented Software Developer Location: Remote Department: Veterans Affairs Type:  Full Time Min. Experience:  Experienced Security Clearance Level:  Public Trust (MBI)  Military Veterans are highly encouraged...
    Full time
    Remote work

    Ssg

    Alpharetta, GA
    23 hours ago
  • $128k - $216k

     ...Senior AI Solutions Engineer Calling all innovators - find your future at Fiserv. We're...  ...another millions of times a day - quickly, reliably, and securely. Any time you swipe your...  ...shaping the future of fintech, this role is on-site Monday through Friday This role... 
    Work at office
    Visa sponsorship
    Monday to Friday

    BentoBox

    Alpharetta, GA
    1 day ago
  •  ...develop models of possible future configurations Create daily test metrics and reporting Occasionally perform other IT systems engineering activities such as requirements, design, installation, operation, sustainment, and support Apply technical principles,... 
    Full time
    Work experience placement
    Remote work

    Ssg

    Alpharetta, GA
    23 hours ago
  •  ...We are looking to add a Software Engineer (Level II) to our Development team to help build out our Bright Suite solutions. Ideal candidates will have the opportunity to work in a fast paced, exciting environment where their work will be noticed and appreciated. As an... 
    Full time
    Work at office
    Relocation

    Deposco

    Alpharetta, GA
    23 hours ago
  • $120k - $160k

     ...We reach higher. We do the right thing—today and for generations to come. Job Purpose and Impact ~ The AI Security Engineering Manager will help solidify foundation for the company's modern business applications. In this role, you will apply your knowledge... 

    Cargill

    Alpharetta, GA
    23 hours ago
  •  ...value. Job Summary Greenstone is seeking a mid-level Software Engineer to join our integrations team. This role is ideal for someone...  ...– 30% Monitor, troubleshoot, and optimize the performance and reliability of integration flows. Ensure integration meets security, auditability... 
    Remote work

    Cultura Technologies

    Alpharetta, GA
    12 hours ago
  • $85.9k - $137k

     ...Senior Software Engineer, WaveLogic Modem Software Development As the global leader in high-speed connectivity, Ciena is committed to a people‑first approach. We prioritize a flexible work environment that empowers individual growth, well‑being, and belonging. Ciena’s... 
    Local area
    Remote work
    Flexible hours

    Ciena

    Alpharetta, GA
    1 day ago
  •  ...We are seeking a Senior Software Engineer to support a large-scale modernization initiative focused on building a cloud-native platform on Azure using .NET Core and Angular. This is a hands-on engineering role–we are looking for strong developers who write high-quality... 

    TalentBridge

    Alpharetta, GA
    3 days ago
  • Position Summary The Senior Software Engineer will provide senior engineering development and leadership on Shared Services products development, enhancement, implementation projects, and maintenance/support needs. The role leads a diverse technical team to ensure the... 

    Scientific Games

    Alpharetta, GA
    1 day ago
  •  ...-oriented, creative problem solver, and collaborative Solutions Engineer to join our Integrated Solutions team. The Solutions Engineer plays...  ..., conduct hospital walk‑throughs, co‑lead in person working on site sessions. Specific vision abilities required by this job include... 
    Work experience placement

    Care Logistics

    Alpharetta, GA
    4 days ago
  • $95.3k - $158.8k

    ## Senior Software Engineer IIApplylocations: Georgia: Florida: Alpharetta, GA: Ohio: Pennsylvaniatime type: Full timeposted on: Posted Todayjob requisition id: R116325Are you passionate about combining software engineering excellence with research-driven problem solving... 
    Local area

    LexisNexis Risk Solutions

    Alpharetta, GA
    11 hours ago
  •  ...positioned for continued growth and is recognized among the nation’s best‑managed firms. Doeren Mayhew is seeking a Senior Software Engineer. The role may be based in Troy, Michigan; Atlanta, Georgia; or Dallas, Texas. Responsibilities Act as a client‑facing full stack... 

    Doeren hew

    Alpharetta, GA
    4 days ago
  • Equifax is seeking creative, high-energy and driven software engineers with hands-on development skills to work on a variety of...  ...to drive engineering excellence in areas of security, quality, reliability, performance, cost optimization, efficiency Generative AI:... 
    Shift work

    Equifax

    Alpharetta, GA
    1 day ago
  •  ...Job Description Job Description The Lead Operating Engineer is responsible for the HVAC system and all mechanical equipment within the building. The position works very closely with the Chief Engineer to ensure that the building systems are functioning properly.... 
    For contractors
    Work at office

    Lincoln Property Company

    Alpharetta, GA
    7 days ago
  •  ...Solution Engineer Intern Location: Alpharetta, GA (hybrid) Reports to: US Lead Solution Engineer Travel: None About Stonebranch Stonebranch builds IT orchestration and automation solutions that transform business IT environments from simple IT task automation... 
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Monday to Friday
    2 days per week
    3 days per week

    Stonebranch

    Alpharetta, GA
    4 days ago
  • $131k - $218.3k

     ...a highly skilled and experienced Senior AI Full Stack Software Engineer to join our innovative team at CMM. This role involves designing...  ...with candidates. McKesson job postings are posted on our career site: careers.mckesson.com. McKesson is an Equal Opportunity Employer... 
    Full time

    McKesson

    Alpharetta, GA
    16 hours ago
  •  ...We are looking to add our first Senior AI Engineer to our Platform and Innovation R&D team to help build out our Supply Chain Intelligence solutions. Ideal candidates will have the opportunity to work in a fast paced, exciting environment where their work will be... 
    Full time
    Work at office
    Relocation

    Deposco

    Alpharetta, GA
    23 hours ago
  • Overview Care Logistics is seeking a Solutions Engineer to join our Integrated Solutions team. The Solutions Engineer will help design and...  ..., and in‑person sessions such as hospital walk‑throughs and on‑site workshops. EEO STATEMENT We are an Equal Opportunity Employer and... 
    Work experience placement

    CareLogistics, LLC

    Alpharetta, GA
    23 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!