Senior Site Reliability Engineer
Replit
Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.
About the role:
Join our Site Reliability Engineering team and help ensure the reliability, scalability, and performance of Replit's infrastructure that serves millions of developers worldwide. As a Site Reliability Engineer, you will bridge the gap between development and operations, implementing automation and establishing best practices that enable our platform to scale efficiently while maintaining high availability.
We are seeking SREs who are passionate about building and maintaining resilient systems at scale. Your mission will be to design and implement robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability and performance.
You will:
- Design and Implement Observability Solutions : Develop comprehensive monitoring and alerting systems using modern observability tools. Create dashboards and metrics that provide real-time visibility into system health and performance. Implement logging strategies that enable quick problem identification and resolution.
- Drive Automation and Infrastructure as Code : Architect and implement infrastructure automation solutions using tools like Terraform, Ansible, or Pulumi. Design and maintain CI/CD pipelines that enable reliable and consistent deployments. Create self-healing systems that can automatically respond to common failure scenarios.
- Establish SLOs and SLIs : Work with product and engineering teams to define and implement Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Build systems to track and report on these metrics, ensuring we maintain high reliability standards while balancing innovation speed.
- Incident Management and Response : Lead incident response efforts, conducting thorough post-mortems, and implementing improvements to prevent future occurrences. Develop and maintain runbooks for critical services. Build tools and processes that reduce Mean Time To Recovery (MTTR).
- Performance Optimization : Identify and resolve performance bottlenecks across our infrastructure. Implement capacity planning strategies and optimize resource utilization. Work on reducing latency and improving system efficiency across global regions.
Required skills and experience:
- 4-8 years of experience in Site Reliability Engineering or similar roles (DevOps, Systems Engineering, Infrastructure Engineering)
- Strong programming skills in languages commonly used for automation (Python, Go, or similar)
- Deep understanding of distributed systems
- Experience with container orchestration platforms (Kubernetes) and cloud-native technologies
- Proven track record of implementing and maintaining monitoring/observability solutions
- Strong incident management skills with experience leading incident response
- Experience with infrastructure as code and configuration management tools
Bonus Points:
- Experience with Google Cloud Platform (GCP) services and tools
- Knowledge of modern observability platforms (Prometheus, Grafana, Datadog, etc.)
What we value:
- Problem-solving mindset: Ability to approach complex operational challenges systematically and devise effective solutions
- Self-directed and autonomous: Capable of working independently while collaborating effectively with cross-functional teams
- Strong communication skills: Ability to explain complex technical concepts to both technical and non-technical audiences
- Continuous learning: Passion for staying current with industry best practices and new technologies
- Focus on automation: Strong belief in automating repetitive tasks and building self-healing systems
Full-Time Employee Benefits Include:
- Competitive Salary & Equity
- 401(k) Program with a 4% match (US Only)
- Health, Dental, Vision and Life Insurance
- Short Term and Long Term Disability
- Paid Parental, Medical, Caregiver Leave
- Flexible Time Off (FTO) + Holidays
- Commuter Benefits (In-Office & US Only)
- Monthly Wellness Stipend
- Autonomous Work Environment
- In Office Set-Up Reimbursement (In-Office Only)
- Quarterly Team Gatherings
- In Office Amenities (In-Office Only)
Want to learn more about what we are up to?
- Self-driving Company
- Replit Agent at Scale
- AI Adoption
- Build Open-Source Apps
Interviewing + Culture at Replit
- Operating Principles
- Reasons not to work at Replit
To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.
#J-18808-Ljbffr- ...Discover exciting DevOps job opportunities and connect with 28,396 DevOps professionals. The Senior Site Reliability Engineer role at Jobicy is designed for experienced professionals who are passionate about enhancing system reliability and operational efficiency....SeniorRemote workFlexible hours
$182.8k - $247.3k
...mission to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement- ...developer-tooling company whose product is used by engineering teams at thousands of software companies for... ...commitments and the SRE team is a senior, well-resourced group of nine. As Senior SRE you will lead reliability initiatives across the platform — from defining...Senior
- ...We need a Senior SRE to ensure Xident's verification platform runs at 99.99% uptime.... ...SLOs, incident response, and production reliability for a system that processes millions of... ...structured logging Implement chaos engineering practices to proactively identify failure...SeniorRemote work
$180k - $230k
...Power Acceleration Job Description We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the...SeniorWork at officeLocal areaImmediate startRemote work3 days per week- ...As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture...SeniorWork experience placement
- ...SQS, Cloudflare, GitHub Actions, PostgreSQL, Redis/BullMQ, Node.js/NestJS, Datadog, TypeScript, React, SQL Position: Senior Site Reliability Engineer Engagement period: Ongoing Interview timeline: ASAP Interview process: 1) CV review 2) Interview with our CTO...SeniorContract workImmediate start
- ...Quarterhill is seeking a Senior Site Reliability Engineer (SRE) to join our growing team. This role is an exciting opportunity to contribute to the reliability and performance of smart transportation systems, including a next-generation, cloud-native tolling platform...SeniorLocal area
- ...infra has to match. The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi... ...-assisted development workflows Partner closely with engineering on reliability reviews and architecture decisions...Senior
$150k - $220k
...Senior Site Reliability EngineerJob detailsDepartment / EngineeringRemoteFull-time$150,000 USD - $220,000 USD## About UsMetaRouter is a customer... ...architecture.## About The RoleAs a Senior Site Reliability Engineer, you own significant pieces of our infrastructure and...SeniorFull timeRemote work$150.4k - $277.6k
...Services The Media Platforms SRE team under the Apple Service Engineering division is one of the most exciting examples of Apple’s long... ...field with 4+ years experience At least 6 years in a Reliability Engineering, DevOps or infrastructure focused role Advanced...SeniorRelocationDay shift$148.5k - $223.9k
...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service...SeniorWorldwideWeekend work$180k - $200k
...Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options... ...equities traders rely on every market day. As our first Senior Site Reliability Engineer, you'll define what reliability means at tastytrade,...SeniorWork at office3 days per week- ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise.The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our customers...SeniorWork experience placementFlexible hours
$110k - $145k
...is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role...SeniorFlexible hours$104k - $178k
## Sr. Site Reliability Engineer IApply: Hybrid: NYC Global HQ: Full time: Posted 12 Days Ago: JR00000779# ****Who We Are****DV is the leader in digital performance solutions, helping our advertiser and agency partners Verify the quality of their digital campaigns, Optimise...SeniorFull time$135.2k - $181.2k
...enhance electrical, mechanical, and sensor-based systems to ensure reliability and performance. Configure, calibrate, and validate... ...professional development, including an interest in emerging data engineering tools and methodologies. Preferred Qualifications: ~8+...SeniorWorldwide- ...Onshape is hiring a Principal Software Engineer (SRE) in a hybrid role based in Boston, MA. You will lead reliability initiatives, shape strategy, and act as a technical authority to ensure the platform is fast, resilient, and scalable for customers. You will drive...
- ...services, including Gen Insurance and Engine by Gen, delivering trusted digital and... ...consumers at scale. As a Principal Site Reliability Engineer, you will define and lead the... ...excellence strategy for our platform. This is a senior technical leadership role for an...
$169.3k - $304.7k
...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting...Work experience placementWork at officeRemote work$160k - $180k
...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical...Flexible hours$194k - $237k
## Principal Site Reliability EngineerApplylocations: Scottsdaletime type: Full timeposted on:... ...Purpose**The Principal Site Reliability Engineer partners with development teams by designing... ...techniques.* Be a thought leader: a senior point of expertise on site reliability...Hourly payWork at officeImmediate startVisa sponsorshipWork visaFlexible hours$165k - $185k
...more at tandemdiabetes.com**A DAY IN THE LIFE:**The Principal Site Reliability Engineer (SRE) is responsible for the reliability, availability, and... ...coverage that works across time zones. Participates as a senior escalation tier for high-severity incidents.* Builds...Permanent employmentContract workLocal areaRemote workFlexible hoursShift work- ...A senior Site Reliability Engineer will join an established infrastructure function responsible for highly available, security-conscious cloud systems supporting complex business-critical workloads. You’ll take significant ownership of reliability, scalability, and...Full timeRemote work
$114k - $148k
...Total compensation is based on experience, skills, and location using objective, job-related criteria. Summary As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant, and highly available. If...Work experience placement- ...Engineering, Product, Design, and Marketing Engineering Compensation ~ Zone 1 Base Pay: $214K – $260K Superhuman offers... ...role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them...WorldwideHome officeFlexible hours
$110k - $145k
...operations. You will liaise with product and engineering teams to ensure applications and... ...feedback loop for platform and product reliability. The ideal candidate is a solutions-oriented... ...experience as a platform engineer, site reliability engineer, systems engineer...Work experience placement$107.9k - $195.05k
...The Digital Sector at Leidos currently has an opening for a Site Reliability Engineer (SRE) / Senior Cloud Engineer to work in our Baltimore, Maryland office. This is an exciting opportunity to use your experience helping the Center for Medicare and Medicaid Services...Contract workWork at office- ...personalized care faster. We are building AI agents to support the full arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and maintain the tooling, alerts and incident-response playbooks that keep...
$140k - $195k
...Improve reliability, observability, service health, incident response, and operational readiness. CodeVertex works across data... ..., secure systems, and operational clarity matter. The Site Reliability Engineer role helps turn business needs into reliable execution, whether...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer Eastern, KY
- senior manager customer operations Eastern, KY
- senior product manager mobile Eastern, KY
- senior software engineer ruby on rails Eastern, KY
- sr finance manager Eastern, KY
- sr marketing manager Eastern, KY
- senior customer service Eastern, KY
- senior business manager Eastern, KY
- senior account executive Eastern, KY
- senior account director Eastern, KY

