Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, Site Reliability Engineering

$250k - $325k
Full-time

Replit

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the Role

Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.

This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure . You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries.

This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team.

What You'll Do

  • Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.

  • Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.

  • Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.

  • Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners—not just recommendations.

  • Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools—including Replit—to prototype and automate. Apply rigorous review and verification to AI-generated changes.

  • Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.

  • Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.

What You'll Bring

  • Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team—not only acted as its strongest individual contributor.

  • Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.

  • Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions.

  • Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.

Nice to Have

  • Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.

  • Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.

  • Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.

  • Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.

Full-Time Employee Benefits Include:

Competitive Salary & Equity

401(k) Program with a 4% match ( US Only )

⚕️ Health, Dental, Vision and Life Insurance

Short Term and Long Term Disability

Paid Parental, Medical, Caregiver Leave

Flexible Time Off (FTO) + Holidays

Commuter Benefits ( In-Office & US Only )

Monthly Wellness Stipend

‍ Autonomous Work Environment

In Office Set-Up Reimbursement ( In-Office Only )

Quarterly Team Gatherings

☕ In Office Amenities ( In-Office Only )

Want to learn more about what we are up to?

  • Self-driving Company

  • Replit Agent at Scale

  • AI Adoption

  • Build Open-Source Apps

Interviewing + Culture at Replit

  • Operating Principles

  • Reasons not to work at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Engineering Manager, Site Reliability Engineering in Foster, CA vacancy
  • $130k - $200k

     ...used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal...  ...work hours and being on callStrong communication and time management skillsThe base salary range for this full-time position is... 
    Suggested
    Full time
    Work at office
    Immediate start

    IXL Learning

    San Mateo, CA
    3 days ago
  •  ...Site Reliability Engineer As a Site Reliability Engineer, you have a mindset to maximize system availability through both proactive and reactive...  ...stable and performant Mentor other team members on managing end-to-end availability and performance of mission... 
    Suggested
    Live in

    Omega Solutions Inc

    San Mateo, CA
    3 days ago
  • $196.75k - $243.29k

     ...scale, and helping to create safer, more civil shared experiences for everyone. The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service... 
    Suggested
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $200k - $285k

     ...control, air quality sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the market: more than 30,0...  ...our Command product- it will include both frontend and backend engineers in your team Be able to roll up your sleeves and actively... 
    Suggested
    Full time
    Work experience placement
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  • $220k - $315k

     ...control, air quality sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the market: more than 30,...  ...70+ countries. About the Team Atlas is the founding engineering team responsible for the Maps mandate at Verkada. Our mission... 
    Suggested
    Full time
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  • $200k - $315k

     ...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the...  ...Our mandate is to guarantee unmatched reliability, endurance, and security of local storage...  ...and accessible instantly. As the Engineering Manager for Storage Systems, you will... 
    Full time
    Work at office
    Local area
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  •  ...Engineering ManagerAs an Engineering Manager, you will oversee the delivery and support of engineering deliverables with high quality and on time, working closely with product management, customer success, field and other key stakeholders. You will be responsible for... 

    Alluxio Inc

    San Mateo, CA
    4 days ago
  • $180k - $230k

     ...re looking for a Senior SRE to own the reliability, scalability, and observability of our...  ...ll work closely with platform and data engineering to keep high-throughput, data-intensive...  ...toil — deployment pipelines, capacity management, self-healing systems Partner with engineering... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    2 days ago
  •  ...achieving great things together. Role Summary: As an Engineering Manager, you will lead and inspire a team of engineers while driving...  ..., testing, and deployment to ensure high levels of system reliability and maintainability. Foster an environment of collaboration... 
    Work at office
    Remote work
    3 days per week

    Notable

    San Mateo, CA
    4 days ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible...  ...through monitoring, maintenance, and troubleshooting. Manage and support public cloud platforms such as AWS, Azure, and... 
    Full time

    OPPO

    Palo Alto, CA
    18 hours ago
  • $170k - $230k

     ...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make...  ...at this stage are primarily reactive — on-call, incident management, keeping services stable. At Mithril, you'll also be... 
    Work at office
    Local area
    1 day per week

    Mithril

    Palo Alto, CA
    18 hours ago
  • $165k - $280k

     ...with the ultimate goal of enabling human life on Mars. SR. SITE RELIABILITY ENGINEER (STARLINK) At SpaceX we're leveraging our experience in...  ...infrastructure to support a multi-region environment Manage petabyte scale bare metal compute clusters Closely collaborate... 
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    3 days ago
  •  ...unified, hybrid workforce, with comprehensive management and oversight. And it's already operating at scale...  ...for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual... 
    Shift work

    Wand AI

    Palo Alto, CA
    18 hours ago
  • $200k - $350k

     ...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the...  ...We are looking for a seasoned engineering manager to scale up an established team...  ...builds with an emphasis on resilience and reliability. Expertise in distributed/financial systems... 
    Full time
    Work experience placement
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  • $295.25k - $345.04k

     ...aspect of immersive experience is performance - less friction means players want to stay longer and come back for more. As the Engineering Manager for Consumer Apps - Performance, you will lead a team focused on making Roblox more performant, responsive, and resilient... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Remote work
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  •  ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure...  ...highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models. We run nearly 30... 
    Shift work

    AI Chopping Block

    Menlo Park, CA
    2 days ago
  • $200k - $315k

     ...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the...  ...experience. We’re looking for an Engineering Manager who enjoys building products, growing...  ...discipline. Help build highly reliable, observable, and scalable distributed... 
    Full time
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  • $295.25k - $345.04k

     ...providing the critical infrastructure that connects global brands with our massive community of creators. We are seeking an Engineering Manager to lead the Ads Experience team, responsible for the full spectrum of our monetization experience: Advertiser Experience (... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $200k - $340k

     ...sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the...  ...the Role We are looking for an Engineering Manager to lead a new team focused on building...  ...that improve engineering velocity, reliability, and developer experience across the... 
    Full time
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  • $160k - $225k

     ...control, air quality sensors, alarms, intercoms, and visitor management. We’ve got serious momentum in the market: more than 30,...  ...the Role We are looking for a driven Associate Solutions Engineering Program Leader to lead and mentor a team of Associate... 
    Full time
    Work at office
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    2 days ago
  • $125k - $160k

     ...with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we're leveraging our experience in...  ...and GPU platforms. You will develop automation to deploy and manage on-premise compute resources, create highly scalable and... 
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  •  ...guiding teams to successful project outcomes. Proficient in Microsoft Office Suite (Project, Visio, Excel, PowerPoint) for project management and reporting. Expertise in HP ALM for managing testing cycles, with broad technical knowledge of SaaS and Cloud systems in a... 
    Work at office

    Exaways Corporation

    Foster, CA
    3 days ago
  • $144k - $374k

     ...opportunity We are seeking a senior, hands-on Distinguished Engineer to lead the delivery of AI-native systems into highly...  ...related technical field. Expertise with AI/ML lifecycle management, observability, and validation approaches. Experience working... 
    Full time
    Temporary work
    Summer holiday
    Immediate start
    Flexible hours

    EY

    San Mateo, CA
    2 days ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development...  ...or a similar role, with a strong, objective background in managing large-scale distributed systems. Cloud & Infrastructure... 
    Full time

    Zoox

    Foster, CA
    more than 2 months ago
  •  ...everyone. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that...  ...tooling to reduce developer cognitive overhead and Help manage our AWS footprint by identifying opportunities for better... 
    Full time
    Contract work

    Rivian and Volkswagen Group Technologies

    Palo Alto, CA
    4 days ago
  • $345.04k - $399.42k

     ...more civil shared experiences for everyone. As a Senior Engineering Manager, Communications, you'll lead the team responsible for in-game...  ...You'll balance the demands of massive scale and rock-solid reliability with a relentless focus on the user experience, shipping... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $144k - $193k

     ...who is equally comfortable driving a release end to end and engineering the systems that streamline it. You will work closely with our...  ...scale automation that removes manual release work, from branch management and merge workflows to release visibility and metricsPartner... 
    Full time
    Temporary work
    Relocation package

    Zoox

    San Mateo, CA
    3 days ago
  • $250k - $325k

     ...underneath that experience must make it straightforward to launch services, isolate workloads, and run reliable systems at scale. We're hiring a hands-on Engineering Manager to lead Cloud Infrastructure: the shared infrastructure as code (IaC), networking, storage,... 
    Full time
    Temporary work
    Work at office
    Worldwide
    Flexible hours

    Replit

    Foster, CA
    3 days ago
  • $219.1k - $350.8k

     ...Description Visa's Data Trust & Platform Engineering (DTPE) organization delivers the...  ...cloud adoption, engineering excellence, reliability, and operational maturity while...  ...and platform adoption. Lead incident management governance, operational reviews, risk management... 
    Work experience placement
    Work at office
    Local area

    Visa

    Foster, CA
    2 days ago
  •  ...lead, you will establish and mature the reliability practices used across our cloud...  ...platform services. You will work with Cloud Engineering and product teams to define reliability...  ...SRE and Cloud Engineering  Define and manage short-and-long term SRE roadmap, distributing... 
    Full time
    Temporary work
    Part time
    Worldwide

    Shield AI

    San Mateo, CA
    24 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, Site Reliability Engineering. Be the first to apply!