Engineering Manager, Site Reliability Engineering
$250k - $325kReplit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.
About the Role
Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.
This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure . You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries.
This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team.
What You'll Do
Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.
Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.
Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.
Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations.
Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes.
Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.
Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.
What You'll Bring
Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor.
Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.
Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions.
Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.
Nice to Have
Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.
Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.
Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.
Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.
Full-Time Employee Benefits Include:
Competitive Salary & Equity
401(k) Program with a 4% match ( US Only )
⚕️ Health, Dental, Vision and Life Insurance
Short Term and Long Term Disability
Paid Parental, Medical, Caregiver Leave
Flexible Time Off (FTO) + Holidays
Commuter Benefits ( In-Office & US Only )
Monthly Wellness Stipend
Autonomous Work Environment
In Office Set-Up Reimbursement ( In-Office Only )
Quarterly Team Gatherings
☕ In Office Amenities ( In-Office Only )
Want to learn more about what we are up to?
Self-driving Company
Replit Agent at Scale
AI Adoption
Build Open-Source Apps
Interviewing + Culture at Replit
Operating Principles
Reasons not to work at Replit
To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
$130k - $200k
...used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal... ...work hours and being on callStrong communication and time management skillsThe base salary range for this full-time position is...SuggestedFull timeWork at officeImmediate start- ...Site Reliability Engineer As a Site Reliability Engineer, you have a mindset to maximize system availability through both proactive and reactive... ...stable and performant Mentor other team members on managing end-to-end availability and performance of mission...SuggestedLive in
$196.75k - $243.29k
...scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service...SuggestedFull timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$200k - $285k
...control, air quality sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the market: more than 30,0... ...our Command product- it will include both frontend and backend engineers in your team Be able to roll up your sleeves and actively...SuggestedFull timeWork experience placementShift work$220k - $315k
...control, air quality sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the market: more than 30,... ...70+ countries. About the Team Atlas is the founding engineering team responsible for the Maps mandate at Verkada. Our mission...SuggestedFull timeWork visaFlexible hoursShift work$200k - $315k
...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the... ...Our mandate is to guarantee unmatched reliability, endurance, and security of local storage... ...and accessible instantly. As the Engineering Manager for Storage Systems, you will...Full timeWork at officeLocal areaWork visaFlexible hoursShift work- ...Engineering ManagerAs an Engineering Manager, you will oversee the delivery and support of engineering deliverables with high quality and on time, working closely with product management, customer success, field and other key stakeholders. You will be responsible for...
$180k - $230k
...re looking for a Senior SRE to own the reliability, scalability, and observability of our... ...ll work closely with platform and data engineering to keep high-throughput, data-intensive... ...toil — deployment pipelines, capacity management, self-healing systems Partner with engineering...Work at officeLocal areaImmediate startRemote work3 days per week- ...achieving great things together. Role Summary: As an Engineering Manager, you will lead and inspire a team of engineers while driving... ..., testing, and deployment to ensure high levels of system reliability and maintainability. Foster an environment of collaboration...Work at officeRemote work3 days per week
$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible... ...through monitoring, maintenance, and troubleshooting. Manage and support public cloud platforms such as AWS, Azure, and...Full time$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make... ...at this stage are primarily reactive — on-call, incident management, keeping services stable. At Mithril, you'll also be...Work at officeLocal area1 day per week$165k - $280k
...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we're leveraging our experience in... ...infrastructure to support a multi-region environment Manage petabyte scale bare metal compute clusters Closely collaborate...Temporary workWorldwideWeekend work- ...unified, hybrid workforce, with comprehensive management and oversight. And it's already operating at scale... ...for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual...Shift work
$200k - $350k
...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the... ...We are looking for a seasoned engineering manager to scale up an established team... ...builds with an emphasis on resilience and reliability. Expertise in distributed/financial systems...Full timeWork experience placementWork visaFlexible hoursShift work$295.25k - $345.04k
...aspect of immersive experience is performance - less friction means players want to stay longer and come back for more. As the Engineering Manager for Consumer Apps - Performance, you will lead a team focused on making Roblox more performant, responsive, and resilient...Full timeWork experience placementH1bWork at officeLocal areaRemote workVisa sponsorshipMonday to Friday- ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure... ...highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models. We run nearly 30...Shift work
$200k - $315k
...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the... ...experience. We’re looking for an Engineering Manager who enjoys building products, growing... ...discipline. Help build highly reliable, observable, and scalable distributed...Full timeWork visaFlexible hoursShift work$295.25k - $345.04k
...providing the critical infrastructure that connects global brands with our massive community of creators. We are seeking an Engineering Manager to lead the Ads Experience team, responsible for the full spectrum of our monetization experience: Advertiser Experience (...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$200k - $340k
...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the... ...the Role We are looking for an Engineering Manager to lead a new team focused on building... ...that improve engineering velocity, reliability, and developer experience across the...Full timeWork visaFlexible hoursShift work$160k - $225k
...control, air quality sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the market: more than 30,... ...the Role We are looking for a driven Associate Solutions Engineering Program Leader to lead and mentor a team of Associate...Full timeWork at officeWork visaFlexible hoursShift work$125k - $160k
...with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we're leveraging our experience in... ...and GPU platforms. You will develop automation to deploy and manage on-premise compute resources, create highly scalable and...Temporary workImmediate startWeekend work- ...guiding teams to successful project outcomes. Proficient in Microsoft Office Suite (Project, Visio, Excel, PowerPoint) for project management and reporting. Expertise in HP ALM for managing testing cycles, with broad technical knowledge of SaaS and Cloud systems in a...Work at office
$144k - $374k
...opportunity We are seeking a senior, hands-on Distinguished Engineer to lead the delivery of AI-native systems into highly... ...related technical field. Expertise with AI/ML lifecycle management, observability, and validation approaches. Experience working...Full timeTemporary workSummer holidayImmediate startFlexible hours$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development... ...or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure...Full time- ...everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that... ...tooling to reduce developer cognitive overhead and Help manage our AWS footprint by identifying opportunities for better...Full timeContract work
$345.04k - $399.42k
...more civil shared experiences for everyone. As a Senior Engineering Manager, Communications, you'll lead the team responsible for in-game... ...You'll balance the demands of massive scale and rock-solid reliability with a relentless focus on the user experience, shipping...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$144k - $193k
...who is equally comfortable driving a release end to end and engineering the systems that streamline it. You will work closely with our... ...scale automation that removes manual release work, from branch management and merge workflows to release visibility and metricsPartner...Full timeTemporary workRelocation package$250k - $325k
...underneath that experience must make it straightforward to launch services, isolate workloads, and run reliable systems at scale. We're hiring a hands-on Engineering Manager to lead Cloud Infrastructure: the shared infrastructure as code (IaC), networking, storage,...Full timeTemporary workWork at officeWorldwideFlexible hours$219.1k - $350.8k
...Description Visa's Data Trust & Platform Engineering (DTPE) organization delivers the... ...cloud adoption, engineering excellence, reliability, and operational maturity while... ...and platform adoption. Lead incident management governance, operational reviews, risk management...Work experience placementWork at officeLocal area- ...lead, you will establish and mature the reliability practices used across our cloud... ...platform services. You will work with Cloud Engineering and product teams to define reliability... ...SRE and Cloud Engineering Define and manage short-and-long term SRE roadmap, distributing...Full timeTemporary workPart timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Engineering Manager, Site Reliability Engineering. Be the first to apply!
- site safety Foster, CA
- on-site clinical research associate (traveling/remote) Foster, CA
- junior website developer Foster, CA
- construction site safety Foster, CA
- site reliability engineer sre
- site reliability engineer
- site reliability engineering manager
- site reliability engineer remote
- junior site reliability engineer
- lead site reliability engineer

