Site Reliability Engineer
Incident IQ
Job Description
Job Description
Company Overview:
About Us:
Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.
Purpose:
Incident IQ is committed to creating a future where every K-12 district operates with seamless efficiency. When operations are unified on a single platform, districts gain the clarity and control needed to build a stronger foundation for student success. We're focused on delivering the tools, support, and partnerships that help make that vision a reality.
Mission:
Incident IQ is on a mission to eliminate the friction of disconnected systems and clunky workflows that slow schools down. We're reimagining the critical work that happens behind the scenes, bringing visibility, efficiency, and impact to the processes that keep classrooms running. By streamlining the complex, automating the routine, and surfacing the insights that matter most, we can create the conditions for educators to teach, students to thrive, and districts to shape the future of education.
Site Reliability Engineer (SRE) Overview:
We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.
Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.
We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.
Site Reliability Engineer (SRE) Responsibilities:
- This role is hands-on from day one. Your initial focus will be:
- SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
- Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
- Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
- Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
- Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
- Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.
- For example: in your first few days, you might stand up an SLO and a burn-rate alert in Grafana for our highest-traffic service. Within a couple of weeks, PagerDuty on-call is configured and the rotation is trained on incident command. That's the pace we operate at here: hours and days, not weeks.
Site Reliability Engineer (SRE) Requirements:
- The tools below are what we run today. What matters more is the systems literacy and genuine curiosity about reliability that let you reason from first principles when something breaks in a way none of these tools have seen before:
- Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required. How you acquired it (degree or a rigorous equivalent) is not, since it's what lets you diagnose a novel failure, not just operate a dashboard.
- AI-Accelerated Execution (core requirement): You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn't be possible without it. This is not a bonus skill here; it's how we expect this role to operate.
- Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about what you've actually shipped than the number itself.
- SLI/SLO Methodology: Proven, hands-on track record implementing the SLI/SLO/error-budget model in a prior role, the discipline formalized in Google's SRE Workbook.
- Observability Tooling: Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open-source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
- Incident Management: Proven track record designing on-call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
- Performance & Chaos Engineering: Hands-on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
- Automation, Infrastructure & Cloud: Proficient in Python, Go, or Bash; hands-on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
- Communication: Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something "can't" be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
- Independence & Pace: You don't need hand-holding or check-ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.
What Success Looks Like:
By the end of your first quarter, core services have defined SLIs and SLOs, live in Grafana dashboards, and are backed by burn-rate-based alerting. A documented incident management process is operating end-to-end, run day to day by trained on-call engineers rather than by you personally, from detection through blameless postmortem. Over time, success looks like measurably reduced alert noise, faster Mean Time to Recovery (MTTR), and an engineering organization that trusts its reliability signals enough to make release and investment decisions based on them.
Bonus Points:
- Experience standing up an SRE practice from zero to one ("founding SRE").
- Experience with GitOps and just-in-time production access models.
- Familiarity with eBPF-based auto-instrumentation (eBPF stands for extended Berkeley Packet Filter), such as Grafana Beyla or OpenTelemetry eBPF Instrumentation, for legacy or hard-to-modify codebases. It's a newer approach, nice to have rather than expected.
- .NET experience is a plus, given our engineering stack.
- Certifications aren't required and aren't a strong signal for us; what you've built matters more than what's on your cert wall. If you happen to have one, Certified Kubernetes Administrator (CKA) or Google Cloud Professional DevOps Engineer are the most relevant.
What makes Incident IQ different:
- We facilitate whole-person growth where employees can develop personally as well as professionally.
- We offer an energetic and collaborative environment; everyone's opinion matters!
- We produce software that empowers K-12 schools to run efficiently, allowing for a better classroom experience for students to THRIVE!
- We provide excellent work/life balance. Two amazing offices - a Downtown Atlanta office location and one at Halcyon in Alpharetta!
Incident IQ offers a competitive salary based on experience with a benefits package for full-time employees that includes medical, dental, vision, life insurance, 401k match, and paid-time off (PTO).
Incident IQ is an Equal Opportunity Employer
- Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence...SuggestedWorldwide
$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SuggestedFull timeWork at officeLocal areaRemote workWork from home$100k - $120k
OverviewThe Site Reliability Engineer is a key force behind improving Origami’s time to resolution and advancing overall site reliability and scalability. This person participates in efforts to identify root causes during post-incident investigations, while also identifying...SuggestedFull timeTemporary workWork experience placementFlexible hours$35 - $45 per hour
DescriptionKforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot of Java/API'...SuggestedRemote work$115k - $140k
...Job Title:Senior Site Reliability Engineer (SRE) Job Description We're Concentrix. The intelligent transformation partner. Solution-focused. Tech-powered. Intelligence-fueled. The global technology and services leader that powers the world’s best brands, today and...SuggestedFull timeWork at officeImmediate start- ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer...Full timeWorldwideFlexible hours
- ...OpenShift - Site Reliability Engineer Atlanta , GA / Onsite Qualifications: This position is 60 % SRE and 40% SDE. Required Skillset • Manage and optimize data streaming and API components in OpenShift Onpremise and AWS. • Proactively...Work experience placement
- ...ideal time and number for communication, and the expected pay rate for C2C/1099/W2. Job Description: Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months Contract Visa : US Citizens/ Green Card Need Local to...Contract workLocal areaImmediate start
- ...provisioning, monitoring, and troubleshooting staging and production cloud environments . Experienced in architectural design for reliability, scalability, and performance. Practical application of SRE principles : SLIs, SLOs, error budgets, automation, incident...
- ...We are currently looking for a Senior Software Engineer to be a part of the Site Reliability Engineering (SRE) team in Atlanta, GA . The SRE team is an innovative team devoted to providing a Docker-based Platform as a Service and assisting a growing number of teams...Contract workWork at officeLocal area
$95k - $171k
.... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts...Permanent employmentWork experience placementWork at officeRemote workWork from homeWorldwideFlexible hours- ...Site Reliability Engineer Opportunity Rainforest is an early stage payments-as-a-service startup that has developed a solution that makes monetizing payments for vertically focused software platforms fair and simple. We focus on small-to-mid sized platforms that want...Work experience placementFlexible hours
$120k - $175k
...Senior Site Reliability Engineer (SRE) Atlanta, GA preferred, Remote At PrizePicks, we are the fastest-growing sports company in North America, as recognized by Inc. 5000. As the leading platform for Daily Fantasy Sports, we cover a diverse range of sports leagues...Remote workWork visaFlexible hours$71.6k - $119.4k
...support application teams. Our services provide applications with reliability, security, and better customer experiences. About the Job:... ...automation, troubleshoot issues, and work closely with senior engineers to learn and apply best practices. You'll gain exposure to a...Full timeTemporary workInternshipLocal areaWork from home$121.4k - $218.6k
...will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings. Partner with... ...and defend them when they are breached. As a Senior Site Reliability Engineer, you will be responsible for: Developing and scaling robust...Work experience placementWork at office- ...Responsibilities Kforce has a client that is seeking a remote Site Reliability Engineer to join their team.Summary:The team consists of systems that can track lead management, job management and sales management. It is built on Salesforce but underpinned by a lot...Hourly payContract workRemote work
- ...Senior Site Reliability Engineer Atlanta, Georgia Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered. We're on a mission to empower the healthcare industry to better onboarding, deploy, and manage their workforce. Over...Permanent employmentFull timeWork at officeRemote workWork from homeWork visa
$60 - $68 per hour
...Site Reliability Engineer Immediate need for a talented Site Reliability Engineer. This is a 12+ months contract opportunity with long-term potential and is located in Atlanta, GA (Onsite). Please review the job description below and contact me ASAP if you are interested...Contract workLocal areaImmediate start- ...Site Reliability Engineer (SRE) When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage...Work at office
$104.9k - $174.7k
...58 Are you passionate about improving reliability, scalability, and resilience in complex... ...practices. Own prioritization of reliability engineering tasks within team backlogs. Lead... ...(IaaS). Background in DevOps, site reliability engineering practices, or related...Full timeLocal area$74.1k - $148.3k
...systems. Facilitate service capacity planning and demand forecasting, software performance analysis, and system tuning. As a Site Reliability Engineer, you will solve interesting technical challenges by defining, designing, deploying, and solving key Oracle Cloud services,...Temporary workImmediate startFlexible hours- ...resilient software platforms using SRE and AI-native engineering practices. Own production reliability, monitoring, and operational automation while mentoring... ...), Kubernetes, and production operations. Key Skills Site Reliability Engineering Terraform Python AWS Azure...Temporary workFlexible hours
- Kforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and legacy platform retirement initiative. This role will focus on maintaining and optimizing...Temporary workRemote work
- ...English (Required) Work Shift: 1st shift (United States of America) Please review the following job description: The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
- A technology company is seeking a skilled Site Reliability Engineer (SRE) with expertise in AEM to ensure application reliability, performance, and scalability. Responsibilities include implementing monitoring solutions, automating deployments, and optimizing cloud costs...
- ...Role: Site Reliability Engineering (SRE) Architect Location: Atlanta, GA (Hybrid on-site) Contract Role Summary: As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems...Contract workEarly shift
- Direct message the job poster from STAFFWORXS Delivery Manager @ STAFFWORXS | US IT Recruitment Job Opening: AWS Site Reliability Engineer (SRE) We’re hiring a Site Reliability Engineer (SRE) to join our team in Atlanta, GA. This hybrid role offers the opportunity to work...Contract work
$151k - $297k
...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB's cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and...Local areaRemote workWorldwideFlexible hours- #CareersJC 1483593Qualifications· Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.· Hands-on experience with incident management and 24/7 production support models.· Proficiency with monitoring and observability...
- ...build, and test of the company's first combined turbojet-ramjet engine and is now being scaled through its first flight vehicle... ...capabilities to the warfighter. Hermeus is seeking a Senior Software or Site Reliability Engineer to join the Information Team and take charge of...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer remote Atlanta, GA
- site reliability engineer Atlanta, GA
- site reliability engineer sre Atlanta, GA
- junior website developer Atlanta, GA
- website content developer Atlanta, GA
- on site coordinator Atlanta, GA
- website coordinator Atlanta, GA
- site leader Atlanta, GA
- site recruiter Atlanta, GA
- historic site Atlanta, GA


