Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

Replit

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the role:

Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit’s infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability.

We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure’s reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit.

You Will:
  • Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection.

  • Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed.

  • Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution. Conduct thorough, blameless post-mortems and drive the implementation of preventative measures. Develop and refine runbooks and build automation to reduce Mean Time To Recovery (MTTR).

  • Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios.

  • Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments, with a deep focus on Kubernetes, Docker, and GCP. Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions.

  • Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack. Use your findings to design and implement long-term fixes that make our systems and products more robust, operable, and easier to diagnose.

  • Provide Staff-Level Guidance: Review feature and system designs from across the company, acting as a key owner for the reliability, scalability, security, and operational integrity of those designs.

  • Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems, making reliability a core value of the Replit engineering culture.

  • Build and Integrate: Write high-quality, well-tested code in Python or Go to meet the needs of your customers, whether it’s building new internal tools or integrating with third-party vendors.

Required Skills and Experience:
  • 8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).

  • Strong programming skills in languages like Python or Go. You write high-quality, well-tested code.

  • Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture.

  • Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.

  • Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing).

  • Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure.

  • Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.

  • Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices.

  • Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels.

  • A willingness to dive into understanding, debugging, and improving any layer of the stack.

  • You’re passionate about making software creation accessible and empowering the next generation of builders.

Bonus Points:
  • Deep experience with Google Cloud Platform (GCP) services and tools.

  • Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry).

  • Experience designing and building reliable systems capable of handling high throughput and low latency.

  • Significant experience with Go and Terraform.

  • Familiarity with working in rapid-growth, startup environments.

  • Experience writing company-facing blog posts and training materials.

Full-Time Employee Benefits Include:

Competitive Salary & Equity

401(k) Program with a 4% match (US Only)

Health, Dental, Vision and Life Insurance

Short Term and Long Term Disability

Paid Parental, Medical, Caregiver Leave

Flexible Time Off (FTO) + Holidays

Commuter Benefits (In-Office & US Only)

Monthly Wellness Stipend

Autonomous Work Environment

In Office Set-Up Reimbursement (In-Office Only)

Quarterly Team Gatherings

In Office Amenities (In-Office Only)

Want to learn more about what we are up to?
  • Self-driving Company

  • Replit Agent at Scale

  • AI Adoption

  • Build Open-Source Apps

Interviewing + Culture at Replit
  • Operating Principles

  • Reasons not to work at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Foster, CA vacancy
  • $130k - $200k

    ## Senior Site Reliability Engineer### San Mateo, CAIXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance... 
    Suggested
    Full time
    Work at office
    Immediate start

    I Xl

    San Mateo, CA
    4 days ago
  • $240k - $300k

     ...’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most. About The Role We are looking for a hands‑on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers... 
    Suggested
    Full time
    Local area
    Relocation package

    Skydio

    San Mateo, CA
    20 hours ago
  • $100k - $200k

     ...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about... 
    Suggested
    Full time

    OPPO

    Palo Alto, CA
    2 days ago
  • $140k - $165k

     ...'s most complex electronics. We capture digital exhaust and engineering context from assembly lines - images, test logs, BOM data, performance...  ...and the best access to that technology to win. As a Site Reliability Engineer, you'll operate, improve, and scale our AWS-based... 
    Suggested

    instrumental-inc-

    Palo Alto, CA
    20 hours ago
  •  ...Site Reliability Engineer - 100% Remote Site Reliability Engineers (SREs) are responsible for working with different developer teams to keep our systems running smoothly. They are a blend of pragmatic operators and software craftspeople that apply excellent problem-... 
    Suggested
    Remote work
    Shift work

    Talentify.io

    Redwood City, CA
    1 day ago
  • $180k - $230k

     ...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the... 
    Work at office
    Local area
    Immediate start
    Remote work
    3 days per week

    GridCARE, Inc.

    Redwood City, CA
    4 days ago
  •  ...technologies. Our mission is to double America’s compute capacity without building new data centers. We are seeking a skilled Site Reliability Engineer to join our growing team. The ideal candidate will help ensure the reliability, scalability, and performance of our hybrid... 
    Work at office
    Weekend work

    FLUIX

    Palo Alto, CA
    2 days ago
  •  ...Elise AI, IBM and Accern. Position Summary We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE... 
    Shift work

    Wand AI

    Palo Alto, CA
    2 days ago
  • $210.38k - $243.21k

     ...Manager, Site Reliability Engineer (Hybrid in South San Francisco) About the Role We are seeking an experienced and hands‑on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring... 

    Twist Bioscience

    South San Francisco, CA
    3 days ago
  •  ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently... 
    Shift work

    AI Chopping Block

    Menlo Park, CA
    4 days ago
  • $243.29k - $295.25k

     ...everyone wants to play, by making higher-fidelity avatars and environments performant at platform scale.You Will:Design and ship 3D engine systems for avatar and environment geometry, including mesh processing, level of detail, transcoding, streaming, runtime loading,... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $196.75k - $243.29k

     ...civil shared experiences for everyone.Portal, part of Roblox’s Engine Productivity team, builds productivity tools for the hundreds of...  ...productivity and take solutions from early experiments to reliable tools engineers use every day.You will:Own improvements to developer... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    2 days ago
  • $110k - $186k

     ...future of humanoid robotics. We partner closely with leaders across Engineering, Manufacturing, Supply Chain, Operations, and Corporate...  ...can scale without compromising performance or culture. As a Staff Manufacturing Recruiter, you will serve as a senior talent partner... 
    Full time
    Temporary work
    Local area
    Work from home
    Flexible hours

    1x

    San Carlos, CA
    a month ago
  • $160k - $271k

    SummaryJoin Guidewire as a Senior Software Engineer, Application Platform, and be part of the team building our next-generation platform...  .... Your work will drive measurable outcomes—improving platform reliability, security, and customer value—while supporting Guidewire’s... 
    Full time
    Part time
    Worldwide
    Flexible hours

    Guidewire Software, Inc.

    San Mateo, CA
    2 days ago
  •  ...our SRE function.   As the SRE lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services. You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and... 
    Full time
    Temporary work
    Part time
    Worldwide

    Shield AI

    San Mateo, CA
    1 day ago
  • $153.12k - $196.75k

     ...evaluation, and developer tooling while learning how to build reliable systems at scale. You’ll design, code, test, launch, and operate...  ...across product, research, data, infrastructure, safety, creator, engine, discovery, and economy. You will report to an Engineering... 
    Full time
    Work experience placement
    Internship
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    20 hours ago
  • $150k - $265k

     ..."Native Systems Layer"—ensuring that our mobile foundation is reliable, scalable, and easy for other teams to build upon. You will also...  ...how we can best balance native performance with global engineering velocity.What You Bring6+ years of professional Android development... 
    Full time
    Worldwide
    Work visa
    Flexible hours
    Shift work

    Verkada

    San Mateo, CA
    1 day ago
  • $136.09k - $168.11k

     ...products, designing exceptional experiences, building scalable platforms, and delighting customers. We are seeking a Lead Solution Engineer to join our growing GTM team in North America, focusing on our Enterprise business. If you have strong analytical skills and a... 
    Work at office
    Flexible hours

    Freshworks

    San Mateo, CA
    a month ago
  • $126.82k - $164.12k

     ...scientific data to drive decisions throughout drug discovery. The Staff Software Engineer, Research Systems partners directly with scientists to...  ...infrastructure.Monitor application performance, reliability, and adoption.Identify opportunities to improve usability,... 
    Full time
    For contractors
    Local area

    GILEAD Sciences

    Foster, CA
    3 days ago
  • $163.7k - $245.5k

     ...synonymous with entertainment excellence and creativity.Software Engineer II DevEx Location: San Mateo (Hybrid)DevEx team builds...  ...PlayStation engineers develop, test, and troubleshoot software in a reliable, repeatable environment. These tools shorten feedback cycles, support... 

    Sony Interactive Entertainment America

    San Mateo, CA
    1 day ago
  • $295.25k - $345.04k

     ...unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.As a Principal Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday
    3 days per week

    Roblox

    San Mateo, CA
    2 days ago
  • $116k - $150k

    IXL Learning, developer of personalized learning products used by millions of people globally, is seeking Software Engineers who have a passion for technology and education to help us add new features to our extremely successful educational products and build new, innovative... 
    Full time
    Work at office

    IXL Learning

    San Mateo, CA
    2 days ago
  • $123k - $190.9k

     ...Progress starts with you.Job DescriptionSoftware Development Engineers are expert problem-solvers and builders who design, implement,...  ...investing time in training resources to improve product availability, reliability, efficiency, observability, and performance.This is a hybrid... 
    Full time
    Work experience placement
    Work at office
    Local area

    Visa

    San Mateo, CA
    1 day ago
  • $219k - $271k

     ...application teams to move them onto it. We are looking for an engineer who is opinionated about what the right path looks like, who has...  ...to more riders and more cities.In this role, you will:Own the reliability and operational excellence of our deployment platformBuild and... 
    Full time
    Temporary work
    Relocation package

    Zoox

    San Mateo, CA
    3 days ago
  • $243.29k - $295.25k

     ...and underpins every online data workload at Roblox. As a senior engineer on the database team, you will shape the architecture, build...  ...launch critical database capabilities that keep our services fast, reliable and efficient at global scale. You will report to the Technical... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    1 day ago
  • $243.29k - $295.25k

     ...civil shared experiences for everyone.Are you a seasoned engineer with a passion for reliability and scalability? We’re looking for exceptional Software...  ...years of experience with added advantage working in the Site Reliability space in SRE or Software EngineeringPassion... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    3 days ago
  • $114k - $202k

     ...Protocol) technologies?Join our highly collaborative Solution Engineering and Operations team to design and deliver intelligent, next-...  ...and GenAI components with enterprise standards for security, reliability, and performance.Ensure scalability, performance, and security... 
    Full time
    Part time

    Guidewire Software, Inc.

    San Mateo, CA
    20 hours ago
  • $243.29k - $295.25k

     ...Drive the integration of service mesh with Kubernetes, ensuring reliable sidecar injection, mTLS, traffic policies, and observability...  ...network stack.Act as a senior voice on the team, mentoring junior engineers and promoting best practices in testing, deployment, and... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    2 days ago
  • $243.29k - $295.25k

     ...detection, translation, serving and rendering. As a Senior Software Engineer, you will design and build highly scalable systems to help...  ...you’ve designed and led implementation on highly scalable and reliable distributed backend systems. You have hands-on experience in microservices... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Worldwide
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    20 hours ago
  • $153.12k - $196.75k

     ...solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.As a Software Engineer on the Discovery UX team, you'll build frontend features across discovery surfaces like Home and Search, helping users find games... 
    Full time
    Work experience placement
    H1b
    Work at office
    Local area
    Visa sponsorship
    Monday to Friday

    Roblox

    San Mateo, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!