Staff Site Reliability Engineer
$200k - $270kBluesky
Role Description
We're looking for a Staff Site Reliability Engineer to help design, implement, and operate the infrastructure that powers Bluesky and atproto. This is a hands-on role for someone who has operated high-scale production systems, understands how distributed systems fail, and wants to build the operational foundation for an open social network.
- Work across bare-metal systems, cloud services, data infrastructure, observability, incident response, capacity planning, and reliability engineering for systems serving millions of users.
- Own reliability, availability, and operational excellence for our production systems, including observability, incident response, deployment, and rollback systems.
- Improve production readiness for services, migrations, and infrastructure changes.
- Develop software that pushes the state of the art in performance, automation, observability, and other areas.
- Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities.
- Reduce toil through automation, tooling, and thoughtful engineering practices.
- Partner with engineers across all our teams to help design services with strong operational characteristics.
- Lead incident reviews and turn contributing factors into concrete engineering improvements as we practice continuous improvement.
- Perform capacity planning and cost management across compute, storage, database, and networking workloads.
- Manage various vendor relationships to ensure we can provide high quality services at a reasonable TCO.
- Mentor engineers on reliability, operability, debugging, and distributed systems practices and help define a culture of operational excellence across the org.
Qualifications
- +10 years experience operating high-scale production systems, including bare metal.
- Strong fundamentals in Linux, networking, storage, databases, and distributed systems.
- Built and operated high-scale systems where correctness, latency, throughput, and availability were critical.
- Can write production-quality software in Go.
- Comfortable debugging across application code, operating systems, databases, networks, and hardware.
- Experience with observability systems, alert design, incident response, capacity planning, Kubernetes, and production automation.
- Like working on very small, fast-moving teams at a startup.
- Have read the AT Protocol docs, feel aligned with the mission, and want to contribute!
Requirements
- Fully remote team with an overlap of working hours with PST required.
- Willingness to travel to team meetups once every 3-4 months.
- On-site component may be required for interviews.
Benefits
- Health, dental, and vision insurance.
Additional Notes
The anticipated base salary range for this position is $200,000 - $270,000 USD, excluding equity. Equity will be considered in the total compensation package. Final base salary for this role will be based on the individual's geographic location, as well as experience level, skill set, training, licenses, and certifications.
- ...At Vynca, our mission is to provide comprehensive care for more quality days at home. About the job We're looking for a Site Reliability Engineer (E3) to help build and operate the infrastructure that powers Vynca's healthcare technology platform. In this role, you'll...SuggestedFull timeLocal areaRemote work
$158.5k - $172k
...exceptional value they deserve. About The Opportunity As a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeWork at office3 days per week$160k - $180k
...have a big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our infrastructural...SuggestedFull timeLocal areaRemote work$90k - $180k
...medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division. We...SuggestedFull timeRemote workShift work$175k - $185k
...together. Come join our team as we develop new ways to improve the lives of working Americans. About the role: As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of our...SuggestedDaily paidRemote workHome officeFlexible hours- ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-... ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,...Contract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$100k - $120k
...position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in United States. The Site Reliability Engineer will play a critical role in improving platform reliability,...Full timeTemporary workRemote workFlexible hours$87.4k - $123.4k
...the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.*** The Site Reliability Engineer will help ensure the reliability, scalability, and performance of Empower’s financial services platform. This person will...16 hoursFull timeContract workTemporary workWork experience placementCasual workWork at officeLocal areaRemote workWork from homeWork visaFlexible hours$104.43k - $156.65k
...Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual... ..., Fox, Disney, NBC, Paramount+, and many others. Our Site Reliability Engineering (SRE) team is at the heart of our mission to deliver...Permanent employmentFull timeWork experience placementWork at officeRemote workWorldwideFlexible hours$150k - $190k
...to help shape how we operate. You'll be responsible for the reliability, performance, and availability of Develocity instances serving... ...Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you...Full timeRemote workWork from homeShift work- ...cloud infrastructure. You'll lead initiatives that improve reliability, scalability, and operational excellence. Key... ...management tools. Required Qualifications • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps. • Hands-on experience...Full timeRemote work
$230k - $255k
Join Aya Healthcare, winner of multiple Top Workplace awards! We're looking for a highly experienced Manager, Site Reliability Engineering to lead the team behind one of healthcare's most relied-on workforce platforms. In this leadership role, you'll guide and grow a...Full timeLocal areaRemote work$187k - $243k
...through early diagnosis and longitudinal care management of chronic conditions. We're looking for a Senior Manager of Site Reliability Engineering to join our team. You'll lead a team of ~10 SREs across North America, UK, HK, and New Zealand — owning both the day-to-...Full timeWork experience placementWork at officeRemote workFlexible hoursShift work$175k - $250k
...been customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market fit and... ...compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability...Full timeRelocation package$140k - $230k
...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services—from designing fault...Full time- Role Description We are looking for a Site Reliability Engineer (SRE) who is passionate about infrastructure reliability, automation, and building scalable production systems. ~Own and improve production infrastructure reliability and stability ~Prepare, execute, and...Full timeRemote work
$106.5k - $177.5k
Role Description The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability...Full timeRemote work- ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our...Full timeWork experience placement
- Role Description We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance...Full timeTemporary work
- Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration...Full timeImmediate startRemote work
- Role Description We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions...Full time
- Role Description As a Site Reliability Engineer on the Central AI team, you will help Health Catalyst engineer teams adopt AI responsibly and effectively. You bring deep experience solutioning and implementing AI systems, and you use that expertise to evaluate architectures...Full time
$130k - $170k
...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software solution to automate securities‑based lending from origination through the life of the loan. By combining thought...Full timeFlexible hoursShift work$114k - $148k
Role Description As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant... ...other team members as needed. You will interact with internal staff, managers, and customers to implement and maintain...Full timeTemporary workWork experience placement- Role Description Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at scale for...Full timeWork at office
$110k - $137.49k
Role Description The Sr Site Reliability Engineer, Release will prototype, write, maintain, and test code in multiple stages of the release process and in multiple environments in order to rapidly deliver automated solutions to our application releases. Assesses unusual...Full timeWork experience placementRemote work- ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD...Full timeRemote work
$126k - $248k
...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and...Full timeLocal areaRemote workWorldwideFlexible hours$130k - $180k
...of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference... ...and AI R&D. THE ROLE Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to...Temporary workWork at officeImmediate startRemote work$190k - $240k
Role Description As a Sr. Site Reliability Engineer (SRE) at ICD, you will play a critical role in ensuring the reliability and seamless operation of our global platform and AWS infrastructure to create scalable and highly reliable software systems. Job Responsibilities...Full timeWork at officeImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- staff security engineer Remote
- project engineer assistant project manager Remote
- assistant chief engineer Remote
- staff data engineer Remote
- senior staff engineer Remote
- staff design engineer Remote
- engineering aide Remote
- software engineer staff Remote
- assistant engineer Remote
- assistant engineering manager Remote












