Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Incident IQ

Job Description

Job Description

Company Overview:

About Us:
Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.

Purpose:
Incident IQ is committed to creating a future where every K-12 district operates with seamless efficiency. When operations are unified on a single platform, districts gain the clarity and control needed to build a stronger foundation for student success. We're focused on delivering the tools, support, and partnerships that help make that vision a reality.

Mission:
Incident IQ is on a mission to eliminate the friction of disconnected systems and clunky workflows that slow schools down. We're reimagining the critical work that happens behind the scenes, bringing visibility, efficiency, and impact to the processes that keep classrooms running. By streamlining the complex, automating the routine, and surfacing the insights that matter most, we can create the conditions for educators to teach, students to thrive, and districts to shape the future of education.

Site Reliability Engineer (SRE) Overview:

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.

Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.

We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

Site Reliability Engineer (SRE) Responsibilities:

  • This role is hands-on from day one. Your initial focus will be:
    • SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
    • Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
    • Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
    • Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
    • Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
    • Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.
  • For example: in your first few days, you might stand up an SLO and a burn-rate alert in Grafana for our highest-traffic service. Within a couple of weeks, PagerDuty on-call is configured and the rotation is trained on incident command. That's the pace we operate at here: hours and days, not weeks.

Site Reliability Engineer (SRE) Requirements:

  • The tools below are what we run today. What matters more is the systems literacy and genuine curiosity about reliability that let you reason from first principles when something breaks in a way none of these tools have seen before:
    • Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required. How you acquired it (degree or a rigorous equivalent) is not, since it's what lets you diagnose a novel failure, not just operate a dashboard.
    • AI-Accelerated Execution (core requirement): You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it. This is not a bonus skill here; it's how we expect this role to operate.
    • Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about what you've actually shipped than the number itself.
    • SLI/SLO Methodology: Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
    • Observability Tooling: Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open-source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
    • Incident Management: Proven track record designing on-call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
    • Performance & Chaos Engineering: Hands-on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
    • Automation, Infrastructure & Cloud: Proficient in Python, Go, or Bash; hands-on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
    • Communication: Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something "can't" be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
    • Independence & Pace: You don't need hand-holding or check-ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.

What Success Looks Like:

By the end of your first quarter, core services have defined SLIs and SLOs, live in Grafana dashboards, and are backed by burn-rate-based alerting. A documented incident management process is operating end-to-end, run day to day by trained on-call engineers rather than by you personally, from detection through blameless postmortem. Over time, success looks like measurably reduced alert noise, faster Mean Time to Recovery (MTTR), and an engineering organization that trusts its reliability signals enough to make release and investment decisions based on them.

Bonus Points:

  • Experience standing up an SRE practice from zero to one ("founding SRE").
  • Experience with GitOps and just-in-time production access models.
  • Familiarity with eBPF-based auto-instrumentation (eBPF stands for extended Berkeley Packet Filter), such as Grafana Beyla or OpenTelemetry eBPF Instrumentation, for legacy or hard-to-modify codebases. It's a newer approach, nice to have rather than expected.
  • .NET experience is a plus, given our engineering stack.
  • Certifications aren't required and aren't a strong signal for us; what you've built matters more than what's on your cert wall. If you happen to have one, Certified Kubernetes Administrator (CKA) or Google Cloud Professional DevOps Engineer are the most relevant.

What makes Incident IQ different:

  • We facilitate whole-person growth where employees can develop personally as well as professionally.
  • We offer an energetic and collaborative environment; everyone's opinion matters!
  • We produce software that empowers K-12 schools to run efficiently, allowing for a better classroom experience for students to THRIVE!
  • We provide excellent work/life balance. Two amazing offices - a Downtown Atlanta office location and one at Halcyon in Alpharetta!

Incident IQ offers a competitive salary based on experience with a benefits package for full-time employees that includes medical, dental, vision, life insurance, 401k match, and paid-time off (PTO).

Incident IQ is an Equal Opportunity Employer

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Atlanta, GA vacancy
  •  ...Technical Support Specialist In Site Reliability Engineering (Sre) Mandatory skills: Scripting and programming languages like Python, Java, Ruby. Cloud and infrastructure management – AWS, Google cloud and Azure is a plus- CI/CD Automation, Database Management. The... 
    Suggested

    Omni Inclusive

    Atlanta, GA
    3 days ago
  • $75.7k - $136.3k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Suggested
    Work experience placement
    Work at office

    Akamai

    Atlanta, GA
    2 days ago
  •  ...advances cures by helping the world's most important research sites do their best work. Our solutions are now used by over 30,00...  ...What You'll Bring to the Team: We are seeking a Site Reliability Engineer (SRE) to join one of our Scrum teams and help ensure the... 
    Suggested
    Work at office

    Florence

    Atlanta, GA
    1 day ago
  • $60 - $68 per hour

     ...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested... 
    Suggested
    Contract work
    Local area
    Immediate start

    Pyramid Corporation

    Atlanta, GA
    3 days ago
  •  ...configure the monitoring and alerting metrics so the support engineers can proactively and timely validate, troubleshoot and...  ...availability critical application components. • 1+ Years in Site Reliability Engineering organization preferred • Overall 4-6years of experience... 
    Suggested
    Work experience placement

    3B Staffing LLC

    Atlanta, GA
    3 days ago
  • $152.13k - $162.13k

     ...challenge the status-quo. Unum is changing, and we’re excited about what’s next. Join us. General Summary: Unum Group seeks Site Reliability Engineers in Atlanta, GA. Applicants who are interested in this position may apply at (Ref #66753) for consideration. Design,... 
    Temporary work
    Work at office
    Remote work

    Unum Group

    Atlanta, GA
    2 days ago
  •  ...Lead Engineer, Site Reliability Engineering Team As a lead engineer with Retail, Site Reliability Engineering team, you will be at the forefront of Cloud and Big Data technology. In this role you will establish yourself as a technical leader by exposing yourself to... 

    Next Level Business Services, Inc.

    Atlanta, GA
    11 hours ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Atlanta, GA
    3 days ago
  •  ...availability. • Automation Experience with Build/deployment, Software Configuration/Continuous Integration/Continuous Delivery/Release Engineering related tasks in JavaEE/C++ Environments. • Experience in automating manual processes using Python, Ruby, Unix Shell (bash,... 
    Immediate start

    Navtech

    Atlanta, GA
    3 days ago
  • $109.5k

     ...and YouTube. ( Job Description AbbVie Information Security is looking for a highly motivated, diligent, and skillful Site Reliability Engineer to join the Cyber Security Engineering (CSE) Team. The CSE Team, working within the Cyber Security Operations (CSO) function... 
    Temporary work
    Local area
    Remote work

    AbbVie

    Atlanta, GA
    5 days ago
  • $178.13k - $205.4k

     ...Bachelor's degree or foreign degree equivalent in Computer Engineering, Computer Science, Engineering, or related field plus five (5)...  ...websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a job... 
    Work at office
    Remote work
    Flexible hours

    Workday

    Atlanta, GA
    2 days ago
  • $165k - $241.4k

     ...Cisco Meraki, we are responsible for building and growing the cloud that supports these customers and their networks. As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze... 
    Permanent employment
    Full time
    Temporary work
    Part time
    Local area
    Flexible hours

    CISCO Systems

    Atlanta, GA
    4 hours ago
  • $126k - $248k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Atlanta, GA
    3 days ago
  •  ...in the right place. Job Details: Job Title: SRE Engineer Location: Atlanta GA (Hybrid Duration: 1 year Contract...  ...critical application components. 1+ Years in Site Reliability Engineering organization preferred. Overall 4-6 years of... 
    Contract work
    Work experience placement
    Work at office
    Remote work

    Staffworxs Inc

    Atlanta, GA
    5 days ago
  •  ...enterprise initiatives such as public cloud, data science, AI, engineering innovation, and IoT. Our customers include the world's...  ...is founder-led, profitable, and growing. We are hiring a Site Reliability Engineer Our goal is to perfect enterprise infrastructure DevOps... 
    Work at office
    Local area
    Remote work
    Work from home
    Worldwide

    Canonical

    Atlanta, GA
    27 days ago
  •  ...evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical...  ...highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles... 
    Early shift

    Cloud Analytics Technologies LLC

    Atlanta, GA
    18 days ago
  • $105k - $130k

     ...provide the high-speed capabilities our nation and its allies need to maintain a durable, asymmetric advantage. The Mission Systems Engineering (MSE) Team develops the Mission Management System (MMS)—a software platform that integrates mission subsystems, autonomy services... 
    Weekly pay
    Permanent employment
    Full time
    Work at office

    Hermeus

    Atlanta, GA
    1 day ago
  •  ...We have an immediate need for a Senior Release Train Engineer for a contract assignment located in Carmel, Indiana . The Release Train Engineer (RTE) has a primary purpose of supporting an Agile Release Train (ART) by steering it to success and navigating the complexity... 
    Contract work
    Work at office
    Immediate start

    Spartan Technologies

    Atlanta, GA
    3 days ago
  •  ...Job Description - Agile -Release Train Engineer (RTE) Duration: FULL TIME Location: Atlanta, GA Role and Responsibilities: Serve as the key facilitator for the Agile Release Train (ART) in Mobile/Web/Services IT projects. Collaborate... 
    Full time

    ACI Infotech

    Atlanta, GA
    5 days ago
  • $101.5k - $169.1k

     ...Company Cox Automotive - USA Job Family Group Engineering / Product Development Job Profile Sr Release Train Engineer...  ...organizational AI policies and standards. Monitor AI tool reliability across teams. Create backup plans for system failures.... 
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours
    Shift work

    Cox Communications

    Atlanta, GA
    3 days ago
  • $140k

     ...their lives. About the Role: As a Stable Kernel Senior Software Engineer, you play an essential role in setting our portfolio of world-...  ...as you learn new technologies. Your knowledgeable practice, reliability, and consultative nature make you an engineer that... 
    Full time
    Contract work
    Temporary work
    Visa sponsorship
    Work visa
    Flexible hours

    Stable Kernel

    Atlanta, GA
    17 hours ago
  • $117.8k - $212.5k

     ...year-round money coaches. That’s how we’re UNSTOPPABLE for our employees! Are you ready to join the Un-carrier movement? The Engineer, Forward Deployment is a hands-on technical practitioner who embeds within a T-Mobile business unit to design, build, and iterate... 
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Atlanta, GA
    3 days ago
  • $124.6k - $148.2k

     ...serve them, and shape the consumer experience. Our product and engineering organizations bring together small, empowered teams that move...  ...on connected service tools that improve resolution speed and reliability, data and API platforms that unlock growth and decisioning,... 
    Full time
    Temporary work
    Local area
    Relocation

    The Coca-Cola Company

    Atlanta, GA
    3 days ago
  •  ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems...  ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes... 

    Saviynt

    Atlanta, GA
    11 days ago
  • $142k - $214k

     ...applies practical experience with modern software technologies to redefine robotics using autonomy. We are looking for Full Stack Engineers to join our team. As a Full Stack Engineer at Anduril you will be architecting and building out user interface applications from... 
    Full time
    Work experience placement
    Local area
    Remote work
    Relocation package

    Anduril Industries

    Atlanta, GA
    1 day ago
  • $143k - $191k

     ...computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems survive... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Atlanta, GA
    5 days ago
  •  ...Saviynt’s platform is mission-critical for our customers. As we scale globally, reliability, availability, and performance are not optional—they are core product features. As a Principal  Engineer, you will define and drive the reliability strategy for our SaaS platform.... 

    Saviynt

    Atlanta, GA
    28 days ago
  • $61 - $80 per hour

     ...Description Anduril is seeking a Reliability Engineer to drive product reliability across the full lifecycle of advanced autonomous defense systems-from early concept development through qualification, production, and field deployment. This individual will partner... 
    Contract work
    Temporary work

    Actalent

    Atlanta, GA
    4 days ago
  •  ...Job Description Position Summary The Construction Project Engineer supports the Project Manager and project team in the planning,...  ...monthly pay application reviews and progress verification Field & Site Support Participate in site visits, inspections, and progress... 
    Contract work
    For contractors
    For subcontractor

    Johnson Construction Services

    Atlanta, GA
    18 days ago
  • $106.1k - $176.8k

     ...this position We are seeking a DevOps Engineer (P3) to support and enhance cloud infrastructure...  ...Operations teams to improve platform reliability, scalability, and automation. Support...  ...candidate is expected to work on-site at our Las Colinas office a minimum of two... 
    Full time
    H1b
    Work at office
    Remote work
    2 days per week

    McKesson

    Atlanta, GA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!