Senior Site Reliability Engineer
PlayOn LLC
Senior Site Reliability Engineer
Playon is looking for an experienced Senior Site Reliability Engineer to help us strengthen the reliability, performance, and scalability of our systems. This role sits at the intersection of software engineering and operations — focused on building the tools, automation, and visibility that enable our teams to deliver resilient software at scale. You'll work closely with application engineers, DevOps, and QA teams to evolve our infrastructure, CI/CD pipelines, observability frameworks, and reliability practices. This is a hands-on engineering role with a strong emphasis on automation, performance analysis, and continuous improvement.
The outcomes you'll deliver:
In the first few months, you'll focus on building a clear understanding of our systems and establishing the foundation for stronger observability across our platforms. As you settle in, your scope will grow to include broader reliability and performance initiatives.
- Assess and improve visibility: Work with engineering teams to review our current dashboards, metrics, and logs, identify the biggest gaps, and make targeted improvements that help us better understand system health.
- Tighten monitoring and alerting: Refine alerts and dashboards for the most critical services so we can catch issues earlier and respond faster.
- Build observability into delivery: Add instrumentation and telemetry into existing build and deploy processes to make reliability checks part of our normal release workflow.
- Clarify what "reliable" means: Help define initial SLIs and SLOs for a few core user flows, aligning the team on what good performance and availability look like.
- Streamline incident response: Partner with the Event Commander/on-call rotation to improve how we communicate, coordinate, and follow up during incidents.
- Reduce manual effort: Automate routine checks and monitoring tasks to free up engineers for more impactful work.
Over time, you'll take on a larger role shaping how we measure, monitor, and improve reliability across all services — setting standards, mentoring others, and helping engineering teams make data-driven decisions about performance and stability.
In this role, you can expect to:
- Contribute to system observability i.e implementing, improving metrics, alerting, and dashboards for better insight and faster recovery.
- Develop automation, tooling, and monitoring solutions to support high service availability.
- Partner with application and quality engineering teams to implement best practices in reliability, release automation, and testing.
- Drive operational excellence through proactive incident prevention, blameless postmortems, and capacity planning.
- Participate in on-call rotations to support critical services and ensure rapid response to incidents.
To thrive in this role, you have:
- Solid experience in Python, especially for automation, tooling, and data-driven operational tasks.
- Proficiency in at least one (Java, C++, or Go).
- Strong understanding of Linux systems, cloud infrastructure (AWS, GCP, or Azure), and modern deployment practices (Docker, Kubernetes, Terraform).
- Experience with CI/CD pipelines, version control, and automated testing frameworks.
- Experience with observability tools (e.g., Prometheus, Grafana, ELK, Datadog, etc.) and log/metric analysis for diagnosing issues.
- Proven experience facilitating and documenting Critical User Journeys translating them to actionable SLA/SLO for automation.
- Demonstrated ability to collaborate with cross-functional teams and communicate clearly in high-impact situations.
- A problem-solver who approaches reliability as a shared responsibility across engineering.
- Familiarity with AI-augmented development tools (Claude, Codex) as part of a modern engineering workflow.
Nice to have:
- Experience writing or maintaining end-to-end or integration tests for distributed systems.
- Background in performance testing, capacity planning, or chaos engineering.
- Contributions to internal developer tooling or reliability-focused frameworks.
- Exposure to security, compliance, or change management processes in production environments.
- Relevant certifications.
How you play:
Ownership over Participation: You take responsibility for holistic outcomes, prioritize key objectives, adapt quickly, and follow through against the toughest challenges.
Team over Stars: You are a bridge builder, rallying teams around common goals, finding win-win solutions, and helping others succeed.
Growth over Comfort: You actively seek to expand your comfort zone and skills, embracing new challenges and treating failure as a chance to learn.
Fairness over Popularity: You approach decisions with a scientist's mindset, stay objective, weigh long-term impact, and seek out other perspectives.
PlayOn is where high school sports come to life. Through GoFan, NFHS Network, and MaxPreps, we give every fan a front-row seat to the moments that matter most: the buzzer-beaters, the comeback wins, the senior nights, the rivalries that define a town.
We built our technology for the people who live and breathe high school athletics — the parents who never miss a game, the alumni still cheering from across the country, the communities that show up week after week. From buying tickets to watching a live stream to reliving the highlights, we make it simple to stay close to the sports and the athletes you love most.
Backed by KKR, we build the technology that powers high school athletics from the inside out: Schools trust us to handle ticketing, streaming, fundraising, concessions, merchandise, and more so the people running programs can stay focused on the athletes and fans we all serve together.
We're a growth-stage company on a mission to make high school sports more accessible, more memorable, and more connected than ever before.
When being there means everything, we make sure you never miss a moment.
Why you'll love working at PlayOn
Product, potential, and people. We're a leader in the high school event space, constantly evolving our product to meet the needs of administrators. We focus on solving real challenges, learning quickly, and creating impactful solutions. This is a growth-stage company, meaning your contributions have real impact. You'll have opportunities to grow your skills, tackle meaningful problems, and make a difference in the lives of schools and the students and fans they serve. Our culture is built on accountability, collaboration, growth, and fairness. We don't just show up—we show up for each other. Everyone wears the same jersey, and we play hard, make the extra pass, and cheer one another on. Losses teach us, challenges motivate us, and persistence drives us forward. We value integrity over shortcuts, choosing to do what's right even when it's hard. Together, we strive to be better every day—because we know that's how we win as a team.
The benefits we offer
Multiple medical insurance plans to choose from Dental, vision life and disability insurance Employee Emergency Fund Company equity (stock options) Open PTO policy 401K plan with company match Hybrid/flexible work environment Note: Must be a full-time employee to participate in the company's employee health benefit plan. Part-time employees and interns are not eligible to participate
- ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that... ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to...SeniorFull time
$170k - $220k
Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating...Senior- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8... ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will...SeniorRemote work
$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role...SeniorRemote work- IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion...SeniorWork at officeImmediate start
$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers...SeniorFull time- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with...SeniorFlexible hours
$210k - $230k
GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation...SeniorCurrently hiringRemote work- ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and... ...and networking teams to improve service reliability and deployment workflowsDeploy and... ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering...SeniorWork at officeLocal areaWork from homeFlexible hours
$152.6k - $191.5k
...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include... ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior...SeniorFull timeWork at officeDay shift$168k - $270.25k
...phenomenal people like you to help us accelerate the next wave of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance...SeniorFull time$104.9k - $174.7k
...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory...SeniorFull timeWork at officeLocal areaRemote workWork from home- ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering... ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or...SeniorWork at officeLocal areaWork from homeFlexible hours
- LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is...SeniorFull timeWork at office2 days per week
- ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS...SeniorTemporary work
$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind...SeniorFlexible hours- Job Description:About the Role: We are looking for a Senior SRE to join our Platform Engineering team as the operations owner of our observability platforms. You’ll be responsible for the reliability, scalability, and continued evolution of the tools that give our engineering...SeniorFull time
$119.8k - $234.7k
...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual... ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft... ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability...SeniorOngoing contractLocal area3 days per week- ...the team takes that seriously!The RoleThe Senior SRE at 2K is a hands-on technical leader... ...regions while partnering with network engineers, systems architects, and game studio developers... ...technical direction, influencing reliability from architecture review through production...Senior
- ...professionalism. We are seeking an experienced AWS solution design engineer/architect to join our infrastructure cloud team. The... ...product features efficiently and confidently them into production.As Senior SRE, you will be responsible for providing leadership, design and...Senior
$267k - $356k
...day is currently Tuesday.Lambda's Storage Engineering team is the backbone behind our world-... ...workloads in the industry, which means reliability and performance aren't just goals—they're... ...defined storage across new and existing sites using tools such as Ansible, Jenkins etc...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours- ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation...SeniorFull timeWorldwideFlexible hours
$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage...SeniorWork at officeLocal areaRemote work$101k - $161k
...several prestigious awards, such as Best Engineering Team, Best Company for Diversity,... ...DescriptionWho You'll Work WithWe’re looking for Site Reliability Engineers to join our growing Arista’s... ...: EngineeringExperience level: Mid-Senior LevelIndustry: Computer NetworkingSenior$148k - $235.75k
...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer...SeniorFull time$166k - $258k
...Seattle office a minimum of 4 days/week in order to be considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'll help build the scalable, reliable, and resilient...SeniorFull timeWork at office- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
- Job Description:Note: Fidelity will not provide immigration sponsorship for this positionThe RoleOur Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high...SeniorFull time
$80k - $140k
Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and...SeniorFull timeFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer United States
- lead site reliability engineer United States
- site reliability engineering manager United States
- site reliability engineer remote United States
- site reliability engineer sre United States
- senior maintenance supervisor United States
- senior lead project manager United States
- senior robotics software engineer United States
- senior firewall engineer United States
- senior devops engineer remote United States
