Staff Site Reliability Engineer
Replit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.
About the role:
Join our Site Reliability Engineering (SRE) team and help ensure the reliability, scalability, and performance of Replit’s infrastructure that serves millions of developers worldwide. As a Staff Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability.
We are seeking Staff SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to proactively find and analyze reliability problems across our stack, then design and implement software and systems to create step-function improvements. You will design robust observability solutions, lead incident response, automate operational tasks, and continuously improve our infrastructure’s reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit.
You Will:
-
Architect and Implement Observability: Design, build, and lead the implementation of comprehensive monitoring, logging, and tracing solutions. Create dashboards and metrics that provide real-time visibility into system health and performance, enabling proactive issue detection.
-
Define and Drive Reliability Standards: Work with product and engineering teams to define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to monitor and report on these metrics, holding teams accountable and ensuring we maintain high reliability standards while balancing innovation speed.
-
Lead Incident Management and Response: Act as a senior leader during high-impact incidents, guiding the team to rapid resolution. Conduct thorough, blameless post-mortems and drive the implementation of preventative measures. Develop and refine runbooks and build automation to reduce Mean Time To Recovery (MTTR).
-
Drive Automation and Infrastructure as Code: Architect, build, and improve automation to eliminate toil and operational work. Design and maintain CI/CD pipelines and infrastructure automation using tools like Terraform or Pulumi. Create self-healing systems that can automatically respond to common failure scenarios.
-
Optimize Performance on Kubernetes: Collaborate with core infrastructure and product teams to performance-tune and optimize our large-scale cloud deployments, with a deep focus on Kubernetes, Docker, and GCP. Identify and resolve performance bottlenecks, implement capacity planning strategies, and reduce latency across global regions.
-
Debug and Harden Distributed Systems: Dive deep into debugging extremely difficult technical problems across the stack. Use your findings to design and implement long-term fixes that make our systems and products more robust, operable, and easier to diagnose.
-
Provide Staff-Level Guidance: Review feature and system designs from across the company, acting as a key owner for the reliability, scalability, security, and operational integrity of those designs.
-
Educate and Mentor: Educate, mentor, and hold accountable the broader engineering team to improve the reliability of our systems, making reliability a core value of the Replit engineering culture.
-
Build and Integrate: Write high-quality, well-tested code in Python or Go to meet the needs of your customers, whether it’s building new internal tools or integrating with third-party vendors.
Required Skills and Experience:
-
8-10 years of experience in Site Reliability Engineering or similar roles (e.g., DevOps, Systems Engineering, Infrastructure Engineering).
-
Strong programming skills in languages like Python or Go. You write high-quality, well-tested code.
-
Deep understanding of distributed systems. You’ve designed, built, scaled, and maintained production services and know how to compose a service-oriented architecture.
-
Deep experience with container orchestration platforms, specifically Kubernetes, and cloud-native technologies.
-
Proven track record of designing, implementing, and maintaining sophisticated monitoring and observability solutions (e.g., metrics, logging, tracing).
-
Strong incident management skills with extensive experience leading incident response for complex systems and demonstrated critical thinking under pressure.
-
Experience with infrastructure as code (e.g., Terraform, Pulumi) and configuration management tools.
-
Excellent written and verbal communication skills, with an ability to explain complex technical concepts clearly and simply and a bias toward open, transparent cultural practices.
-
Strong interpersonal skills, with experience working with and mentoring engineers from junior to principal levels.
-
A willingness to dive into understanding, debugging, and improving any layer of the stack.
-
You’re passionate about making software creation accessible and empowering the next generation of builders.
Bonus Points:
-
Deep experience with Google Cloud Platform (GCP) services and tools.
-
Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry).
-
Experience designing and building reliable systems capable of handling high throughput and low latency.
-
Significant experience with Go and Terraform.
-
Familiarity with working in rapid-growth, startup environments.
-
Experience writing company-facing blog posts and training materials.
Full-Time Employee Benefits Include:
Competitive Salary & Equity
401(k) Program with a 4% match (US Only)
Health, Dental, Vision and Life Insurance
Short Term and Long Term Disability
Paid Parental, Medical, Caregiver Leave
Flexible Time Off (FTO) + Holidays
Commuter Benefits (In-Office & US Only)
Monthly Wellness Stipend
Autonomous Work Environment
In Office Set-Up Reimbursement (In-Office Only)
Quarterly Team Gatherings
In Office Amenities (In-Office Only)
Want to learn more about what we are up to?
-
Self-driving Company
-
Replit Agent at Scale
-
AI Adoption
-
Build Open-Source Apps
Interviewing + Culture at Replit
-
Operating Principles
-
Reasons not to work at Replit
To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
#J-18808-Ljbffr$130k - $200k
## Senior Site Reliability Engineer### San Mateo, CAIXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance...SuggestedFull timeWork at officeImmediate start$240k - $300k
...’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most. About The Role We are looking for a hands‑on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers...SuggestedFull timeLocal areaRelocation package$100k - $200k
...OPPO US Research Center is seeking a skilled and proactive Site Reliability Engineer (SRE) to join our team. In this role, you will be responsible for ensuring the stability, scalability, and performance of our application systems. The ideal candidate is passionate about...SuggestedFull time$140k - $165k
...'s most complex electronics. We capture digital exhaust and engineering context from assembly lines - images, test logs, BOM data, performance... ...and the best access to that technology to win. As a Site Reliability Engineer, you'll operate, improve, and scale our AWS-based...Suggested- ...Site Reliability Engineer - 100% Remote Site Reliability Engineers (SREs) are responsible for working with different developer teams to keep our systems running smoothly. They are a blend of pragmatic operators and software craftspeople that apply excellent problem-...SuggestedRemote workShift work
$180k - $230k
...Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...Work at officeLocal areaImmediate startRemote work3 days per week- ...technologies. Our mission is to double America’s compute capacity without building new data centers. We are seeking a skilled Site Reliability Engineer to join our growing team. The ideal candidate will help ensure the reliability, scalability, and performance of our hybrid...Work at officeWeekend work
- ...Elise AI, IBM and Accern. Position Summary We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function. This is a deeply hands-on individual contributor role, to build and operate SRE...Shift work
$210.38k - $243.21k
...Manager, Site Reliability Engineer (Hybrid in South San Francisco) About the Role We are seeking an experienced and hands‑on Site Reliability Engineering (SRE) Manager to lead our Site Operations and infrastructure initiatives. This role is responsible for ensuring...- ...About the Role We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently...Shift work
$243.29k - $295.25k
...everyone wants to play, by making higher-fidelity avatars and environments performant at platform scale.You Will:Design and ship 3D engine systems for avatar and environment geometry, including mesh processing, level of detail, transcoding, streaming, runtime loading,...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$196.75k - $243.29k
...civil shared experiences for everyone.Portal, part of Roblox’s Engine Productivity team, builds productivity tools for the hundreds of... ...productivity and take solutions from early experiments to reliable tools engineers use every day.You will:Own improvements to developer...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$110k - $186k
...future of humanoid robotics. We partner closely with leaders across Engineering, Manufacturing, Supply Chain, Operations, and Corporate... ...can scale without compromising performance or culture. As a Staff Manufacturing Recruiter, you will serve as a senior talent partner...Full timeTemporary workLocal areaWork from homeFlexible hours$160k - $271k
SummaryJoin Guidewire as a Senior Software Engineer, Application Platform, and be part of the team building our next-generation platform... .... Your work will drive measurable outcomes—improving platform reliability, security, and customer value—while supporting Guidewire’s...Full timePart timeWorldwideFlexible hours- ...our SRE function. As the SRE lead, you will establish and mature the reliability practices used across our cloud infrastructure and platform services. You will work with Cloud Engineering and product teams to define reliability targets, improve observability, and...Full timeTemporary workPart timeWorldwide
$153.12k - $196.75k
...evaluation, and developer tooling while learning how to build reliable systems at scale. You’ll design, code, test, launch, and operate... ...across product, research, data, infrastructure, safety, creator, engine, discovery, and economy. You will report to an Engineering...Full timeWork experience placementInternshipH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$150k - $265k
..."Native Systems Layer"—ensuring that our mobile foundation is reliable, scalable, and easy for other teams to build upon. You will also... ...how we can best balance native performance with global engineering velocity.What You Bring6+ years of professional Android development...Full timeWorldwideWork visaFlexible hoursShift work$136.09k - $168.11k
...products, designing exceptional experiences, building scalable platforms, and delighting customers. We are seeking a Lead Solution Engineer to join our growing GTM team in North America, focusing on our Enterprise business. If you have strong analytical skills and a...Work at officeFlexible hours$126.82k - $164.12k
...scientific data to drive decisions throughout drug discovery. The Staff Software Engineer, Research Systems partners directly with scientists to... ...infrastructure.Monitor application performance, reliability, and adoption.Identify opportunities to improve usability,...Full timeFor contractorsLocal area$163.7k - $245.5k
...synonymous with entertainment excellence and creativity.Software Engineer II DevEx Location: San Mateo (Hybrid)DevEx team builds... ...PlayStation engineers develop, test, and troubleshoot software in a reliable, repeatable environment. These tools shorten feedback cycles, support...$295.25k - $345.04k
...unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.As a Principal Software Engineer on the Engine DataModel team, you will own and innovate on the foundational components that form the backbone of the Roblox...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday3 days per week$116k - $150k
IXL Learning, developer of personalized learning products used by millions of people globally, is seeking Software Engineers who have a passion for technology and education to help us add new features to our extremely successful educational products and build new, innovative...Full timeWork at office$123k - $190.9k
...Progress starts with you.Job DescriptionSoftware Development Engineers are expert problem-solvers and builders who design, implement,... ...investing time in training resources to improve product availability, reliability, efficiency, observability, and performance.This is a hybrid...Full timeWork experience placementWork at officeLocal area$219k - $271k
...application teams to move them onto it. We are looking for an engineer who is opinionated about what the right path looks like, who has... ...to more riders and more cities.In this role, you will:Own the reliability and operational excellence of our deployment platformBuild and...Full timeTemporary workRelocation package$243.29k - $295.25k
...and underpins every online data workload at Roblox. As a senior engineer on the database team, you will shape the architecture, build... ...launch critical database capabilities that keep our services fast, reliable and efficient at global scale. You will report to the Technical...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$243.29k - $295.25k
...civil shared experiences for everyone.Are you a seasoned engineer with a passion for reliability and scalability? We’re looking for exceptional Software... ...years of experience with added advantage working in the Site Reliability space in SRE or Software EngineeringPassion...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$114k - $202k
...Protocol) technologies?Join our highly collaborative Solution Engineering and Operations team to design and deliver intelligent, next-... ...and GenAI components with enterprise standards for security, reliability, and performance.Ensure scalability, performance, and security...Full timePart time$243.29k - $295.25k
...Drive the integration of service mesh with Kubernetes, ensuring reliable sidecar injection, mTLS, traffic policies, and observability... ...network stack.Act as a senior voice on the team, mentoring junior engineers and promoting best practices in testing, deployment, and...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday$243.29k - $295.25k
...detection, translation, serving and rendering. As a Senior Software Engineer, you will design and build highly scalable systems to help... ...you’ve designed and led implementation on highly scalable and reliable distributed backend systems. You have hands-on experience in microservices...Full timeWork experience placementH1bWork at officeLocal areaWorldwideVisa sponsorshipMonday to Friday$153.12k - $196.75k
...solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.As a Software Engineer on the Discovery UX team, you'll build frontend features across discovery surfaces like Home and Search, helping users find games...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- technology administrator Foster, CA
- assistant engineer Foster, CA
- staff engineer Foster, CA
- senior staff systems engineer Foster, CA
- engineering aide Foster, CA
- site safety Foster, CA
- on-site clinical research associate (traveling/remote) Foster, CA
- construction site safety Foster, CA
- junior website developer Foster, CA
- staff process engineer




