Site Reliability Engineer
Incident IQ LLC
About Us:
Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.
Company Overview:
About Us:
Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.
Purpose:
Incident IQ is committed to creating a future where every K-12 district operates with seamless efficiency. When operations are unified on a single platform, districts gain the clarity and control needed to build a stronger foundation for student success. We’re focused on delivering the tools, support, and partnerships that help make that vision a reality.
Mission:
Incident IQ is on a mission to eliminate the friction of disconnected systems and clunky workflows that slow schools down. We’re reimagining the critical work that happens behind the scenes, bringing visibility, efficiency, and impact to the processes that keep classrooms running. By streamlining the complex, automating the routine, and surfacing the insights that matter most, we can create the conditions for educators to teach, students to thrive, and districts to shape the future of education.
Site Reliability Engineer (SRE) Overview:
We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what “reliable” means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.
Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.
We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.
Site Reliability Engineer (SRE) Responsibilities:
- This role is hands-on from day one. Your initial focus will be:
- SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
- Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
- Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
- Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
- Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
- Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.
- For example: in your first few days, you might stand up an SLO and a burn-rate alert in Grafana for our highest‑traffic service. Within a couple of weeks, PagerDuty on‑call is configured and the rotation is trained on incident command. That’s the pace we operate at here: hours and days, not weeks.
Site Reliability Engineer (SRE) Requirements:
- The tools below are what we run today. What matters more is the systems literacy and genuine curiosity about reliability that let you reason from first principles when something breaks in a way none of these tools have seen before:
- Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent formal training, with real depth in operating systems, databases, and networking. This fundamental understanding is required. How you acquired it (degree or a rigorous equivalent) is not, since it's what lets you diagnose a novel failure, not just operate a dashboard.
- AI-Accelerated Execution (core requirement): You actively use AI tools daily to multiply your own output, not just experiment with them on the side. We expect you to use AI to write and debug code faster, stand up dashboards and alerts faster, and generally ship at a pace that wouldn’t be possible without it. This is not a bonus skill here; it’s how we expect this role to operate.
- Track Record Over Tenure: A demonstrated history of independently driving big, ambiguous reliability or infrastructure projects to completion, typically reflecting 5+ years in an SRE, DevOps, or production engineering role. We care far more about what you’ve actually shipped than the number itself.
- SLI/SLO Methodology: Proven, hands‑on track record implementing the SLI/SLO/error‑budget model in a prior role, the discipline formalized in Google’s SRE Workbook.
- Observability Tooling: Strong experience with Grafana and PromQL (Prometheus Query Language), Grafana Alloy for Loki logs, and a metrics backend such as Prometheus or Datadog. Experience instrumenting with OpenTelemetry and a tracing/Application Performance Monitoring (APM) backend (open‑source preferred: SigNoz, Uptrace, Tempo; commercial: Datadog, New Relic), plus Real User Monitoring (RUM) and synthetic monitoring (e.g., Grafana Faro, Grafana Synthetic Monitoring / k6).
- Incident Management: Proven track record designing on‑call rotations and incident command practices elsewhere, with tooling such as PagerDuty or equivalent.
- Performance & Chaos Engineering: Hands‑on with a load/performance framework (Locust, k6, or JMeter) and chaos engineering exercises to validate reliability under real conditions.
- Automation, Infrastructure & Cloud: Proficient in Python, Go, or Bash; hands‑on with Infrastructure as Code (Terraform, Ansible, or equivalent), Kubernetes, and at least one major cloud platform (Amazon Web Services (AWS), Google Cloud Platform (GCP), or Azure).
- Communication: Experienced, versatile communicator: able to go deep with developers on root cause, tradeoffs, and implementation detail; comfortable pushing back with a real technical path when a team says something “can’t” be done; precise about the difference between a mitigation and an actual fix when reporting status; and able to translate reliability status, risk, and priorities clearly for business and engineering stakeholders.
- Independence & Pace: You don’t need hand‑holding or check‑ins to make progress. Comfortable resolving ambiguous problems in hours, not weeks, taking full ownership of outcomes, and juggling multiple threads under real time pressure without dropping the ball.
What Success Looks Like:
By the end of your first quarter, core services have defined SLIs and SLOs, live in Grafana dashboards, and are backed by burn‑rate‑based alerting. A documented incident management process is operating end‑to‑end, run day to day by trained on‑call engineers rather than by you personally, from detection through blameless postmortem. Over time, success looks like measurably reduced alert noise, faster Mean Time to Recovery (MTTR), and an engineering organization that trusts its reliability signals enough to make release and investment decisions based on them.
Bonus Points:
- Experience standing up an SRE practice from zero to one (“founding SRE”).
- Experience with GitOps and just‑in‑time production access models.
- Familiarity with eBPF‑based auto‑instrumentation (eBPF stands for extended Berkeley Packet Filter), such as Grafana Beyla or OpenTelemetry eBPF Instrumentation, for legacy or hard‑to‑modify codebases. It’s a newer approach, nice to have rather than expected.
- .NET experience is a plus, given our engineering stack.
- Certifications aren’t required and aren’t a strong signal for us; what you’ve built matters more than what’s on your cert wall. If you happen to have one, Certified Kubernetes Administrator (CKA) or Google Cloud Professional DevOps Engineer are the most relevant.
What makes Incident IQ different:
- We facilitate whole‑person growth where employees can develop personally as well as professionally.
- We offer an energetic and collaborative environment; everyone’s opinion matters!
- We produce software that empowers K‑12 schools to run efficiently, allowing for a better classroom experience for students to THRIVE!
- We provide excellent work/life balance. Two amazing offices - a Downtown Atlanta office location and one at Halcyon in Alpharetta!
Incident IQ offers a competitive salary based on experience with a benefits package for full‑time employees that includes medical, dental, vision, life insurance, 401k match, and paid‑time off (PTO).
Incident IQ is an Equal Opportunity Employer
#J-18808-Ljbffr- ...security to responsibly propel the global lottery industry ever forward. Position Summary We are looking for a skilled Site Reliability Engineer (SRE) to enhance the stability, performance, and reliability of our production systems. The SRE will work closely with...SuggestedPermanent employmentWork experience placementLocal area
- ...Overview About the role You are a Senior Site Reliability Engineer based our Alpharetta, GA office. You lead Cloud Operations for Business Central Online customers in the Americas, executing escalations in-region and environment work that previously required...SuggestedCasual workWork at officeFlexible hours
- ~4+ years of experience in an SRE, DevOps, or cloud infrastructure role. ~ Strong experience with Azure cloud services and infrastructure. ~ Hands-on experience with java and Terraform and Terragrunt for infrastructure-as-code. ~ Proficiency with Kubernetes (preferably...Suggested
$61.07 per hour
...Our client, a technology and cloud operations organization is seeking a Senior Site Reliability Engineer to join their team. As a Senior Site Reliability Engineer, you will be part of the Hosted Operations Team supporting Azure-focused cloud systems and production reliability...SuggestedWeekly payTemporary workWork experience placementFlexible hours$129k - $161k
...Job Description Job Description Job title: Senior Site Reliability Engineer Reports to: Director, Site Reliability Engineering Department: Cloud Platforms Location: Remote Grade: 20 About Priority Commerce: Priority Commerce is a leading financial...SuggestedRemote work$60 - $63 per hour
...more. Base pay range $60.00/hr - $63.00/hr Our client, a global information and analytics company, is looking for a Site Reliability Engineer to join their team in Alpharetta, GA! This is a 12-month initial contract and is hybrid so local candidates are...Full timeContract workTemporary workLocal areaFlexible hours- ...initiatives. Collaborating with product, architecture, and engineering groups to build a platform that streamlines application... ...Experience: 5+ Years advanced level experience in DevOps, Site Reliability Engineering with expertise in Enterprise Cloud infrastructure...
$125k - $175k
...overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize... ...in quantitative discipline (Computer Science, Computer Engineering). - 5+ years’ experience in leading a small to medium team of...Full timeTemporary work- ...shape the future of our communities. This is a Software Engineering position at Director level, which is part of the job family... ...businesses. This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support...
- ...We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability...
- ...your work. You are visible, your talents are valued, and you are empowered to shape the future of payments. As a Senior Site Reliability Engineer (SRE) - Azure & GitOps (CI/CD) in Norcross, GA or Omaha, NE, you will join a diverse, passionate team, dedicated to...Full timeWorldwide
- ...The Site Reliability Engineer (Observability) to join its Platform Engineering organization and help develop, implement, and mature a robust enterprise observability platform. This is a highly technical, hands‑on role focused on improving visibility across complex infrastructure...
- ...Job Description Job Description We are hiring an SRE Platform Engineer (IBM BPM/ODM) with our partner in Alpharetta, GA for an onsite role. Job Details: Job Title: SRE Platform Engineer (IBM BPM/ODM) Location- Alpharetta, GA (Onsite) Interview- F2F (Locals...Local area
- Primary Responsibilities Develop, maintain, and optimize system-level applications on Unix/Linux platforms using C. Design and implement network programming solutions using sockets, TCP/IP, IPC, and related protocols. Develop automation scripts and system utilities...Contract work
$65 - $90 per hour
...Immediate need for a talented Reliability Engineer . This is a 06+months contract opportunity with long-term potential and is located in Alpharetta, GA (Onsite). Job ID:26-29242 Pay Range: $65 - $90/hour. Employee benefits include, but are not limited to,...Contract workLocal areaImmediate start$92k - $114.63k
...infrastructure, optimize operations, and deliver energy that is more reliable, resilient, accessible, safe, and sustainable.Behind every... ...Level of Education: BS, MS in Electrical/Electronic/Chemical Engineering, or equivalentLandis+Gyr is a global leader in energy...Worldwide- ...Select how often (in days) to receive an alert: Job title: Reliability Engineer Job family: Engineering Business area: Pulp & Paper Service Contract location: Alpharetta, GA, US Working location: Hartford City, IN, US Location type: Office Location /...Permanent employmentContract workLive inWork at office
$34.67 - $54.18 per hour
...We are seeking a hands-on Platform Engineer with strong full-stack software development and AI engineering skills to join the CPS PEArch... ..., infrastructure, and production-management teams to improve reliability, accelerate delivery, and strengthen engineering standards....$162.4k - $211.9k
...JOB TITLE: Lead System Engineering JOB LOCATION: 500 North Point Pkwy., Alpharetta, GA 30022 Responsible for translating the core... ...software, cloud, etc.) through functional, performance, and reliability analysis using engineering models and techniques, primarily through...Temporary workLocal area- ...security to responsibly propel the global lottery industry ever forward.Position SummaryAdvanced Software EngineerThis Advanced Software Engineer position is part of the lottery division engineer team at Scientific Games. Your work on this team will help enable the funding of...Full timeLocal area
- ...origin, or disability.Envision yourself at BarracudaThe Cloud-to-Cloud Backup team is looking for an experienced Senior Software Engineer to join our team in Barracuda's Data Protection division. You'll be part of the team building the next generation of our Cloud-to-...WorldwideFlexible hours
- ...media companies deliver exceptional customer experiences through reliable, efficient, and secure operations at scale. We provide... ...efficiencies to the software or business processes by utilizing software engineering tools, various innovative techniques, and reusing existing...Worldwide
- ...and help establish development standards that support scalable, reliable, and maintainable solutions.ResponsibilitiesDesign, develop,... ..., estimate effort, and prioritize enhancements.Apply software engineering best practices, including object-oriented programming, design...Full time
- ...Requisition ID: 7222 Job Title: Senior Software Engineer - Alpharetta, GA Job Country: United States (US) Here at Avanos... ...implementation and firmware upgrade mechanisms, including secure and reliable update processes. ~ Strong knowledge of operating-system...Temporary workFor contractorsImmediate start
$63.22k - $94.55k
...environments with a focus on performance, reliability, and scalability Develop and maintain... ...and process optimization Mentor junior engineers and contribute to knowledge sharing across... ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and...Work at officeRemote workFlexible hours- Software Systems Engineer - IV America Networks is a leading sensor and networking solutions partner for companies in any Industrial, Manufacturing, and Waste management space. We design and manufacture sensors for storage tanks, water metering, energy metering, gas monitoring...
- ...SpringBoot and Microservices architecture Design and implement cloudnative applications with a focus on performance security and reliability Collaborate with crossfunctional teams to deliver highquality software solutions aligned with business requirements Write clean maintainable...For contractorsOverseas
$158k
...shape the future of our communities.What you’ll do in the role:Morgan Stanley Services Group Inc. is seeking an Associate, Software Engineer in Alpharetta, Georgia to develop a framework to streamline web services development, reducing costs and improving team efficiency...Temporary workRemote workWorldwide2 days per week- ...production, ensuring high availability and reliability.Mentor junior developers and contribute... ...’s degree in Computer Science, Engineering, or equivalent work experiencePreferred... ...hire locally to NTT DATA offices or client sites. This ensures we can provide timely and...Full timeTemporary workWork at officeRemote workFlexible hours
$110 per hour
Job Title: Senior Java Microservices Developer - Bill Rate: $110 - Duration: Experience building and deploying cloud-native applications on Azure is a plus (09/14/2026 to 09/12/2027) - Location/Schedule: Oakland/Rancho Cordova/Alpharetta; Hybrid - 2 Days in-office. In-Person...Work at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- site reliability engineer Alpharetta, GA
- site safety Alpharetta, GA
- on-site clinical research associate (traveling/remote) Alpharetta, GA
- site services specialist Alpharetta, GA
- construction site safety Alpharetta, GA
- junior website developer Alpharetta, GA
- historic site Alpharetta, GA
- IT site lead Alpharetta, GA
- site leader Alpharetta, GA
- official site Alpharetta, GA






