Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

$200k - $270k
Full-time

Bluesky

Role Description

We're looking for a Staff Site Reliability Engineer to help design, implement, and operate the infrastructure that powers Bluesky and atproto. This is a hands-on role for someone who has operated high-scale production systems, understands how distributed systems fail, and wants to build the operational foundation for an open social network.

  • Work across bare-metal systems, cloud services, data infrastructure, observability, incident response, capacity planning, and reliability engineering for systems serving millions of users.
  • Own reliability, availability, and operational excellence for our production systems, including observability, incident response, deployment, and rollback systems.
  • Improve production readiness for services, migrations, and infrastructure changes.
  • Develop software that pushes the state of the art in performance, automation, observability, and other areas.
  • Scale systems running on dense, latest-generation, bare-metal servers in our own colocation facilities.
  • Reduce toil through automation, tooling, and thoughtful engineering practices.
  • Partner with engineers across all our teams to help design services with strong operational characteristics.
  • Lead incident reviews and turn contributing factors into concrete engineering improvements as we practice continuous improvement.
  • Perform capacity planning and cost management across compute, storage, database, and networking workloads.
  • Manage various vendor relationships to ensure we can provide high quality services at a reasonable TCO.
  • Mentor engineers on reliability, operability, debugging, and distributed systems practices and help define a culture of operational excellence across the org.

Qualifications

  • +10 years experience operating high-scale production systems, including bare metal.
  • Strong fundamentals in Linux, networking, storage, databases, and distributed systems.
  • Built and operated high-scale systems where correctness, latency, throughput, and availability were critical.
  • Can write production-quality software in Go.
  • Comfortable debugging across application code, operating systems, databases, networks, and hardware.
  • Experience with observability systems, alert design, incident response, capacity planning, Kubernetes, and production automation.
  • Like working on very small, fast-moving teams at a startup.
  • Have read the AT Protocol docs, feel aligned with the mission, and want to contribute!

Requirements

  • Fully remote team with an overlap of working hours with PST required.
  • Willingness to travel to team meetups once every 3-4 months.
  • On-site component may be required for interviews.

Benefits

  • Health, dental, and vision insurance.

Additional Notes

The anticipated base salary range for this position is $200,000 - $270,000 USD, excluding equity. Equity will be considered in the total compensation package. Final base salary for this role will be based on the individual's geographic location, as well as experience level, skill set, training, licenses, and certifications.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Remote vacancy
  •  ...At Vynca, our mission is to provide comprehensive care for more quality days at home. About the job We're looking for a Site Reliability Engineer (E3) to help build and operate the infrastructure that powers Vynca's healthcare technology platform. In this role, you'll... 
    Suggested
    Full time
    Local area
    Remote work

    Vynca

    United States
    1 day ago
  • $158.5k - $172k

     ...exceptional value they deserve. About The Opportunity As a Senior Engineer on the Runtime Automation team, you will design, automate, and...  .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology... 
    Suggested
    Full time
    Work at office
    3 days per week

    Wonder

    Chicago, IL
    2 days ago
  • $160k - $180k

     ...have a big impact. See Arkestro in action at arkestro.com. About the Role Arkestro is hiring for a Senior SRE Engineer to manage our performance and reliability for our software platform and infrastructure. The right candidate will own and develop our infrastructural... 
    Suggested
    Full time
    Local area
    Remote work

    Arkestro

    United States
    5 days ago
  • $90k - $180k

     ...medicines. Our 115,000 colleagues serve people in more than 160 countries. JOB DESCRIPTION: About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division. We... 
    Suggested
    Full time
    Remote work
    Shift work

    Abbott

    Sunnyvale, CA
    5 days ago
  • $175k - $185k

     ...together. Come join our team as we develop new ways to improve the lives of working Americans. About the role: As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of our... 
    Suggested
    Daily paid
    Remote work
    Home office
    Flexible hours

    Branch

    United States
    20 hours ago
  •  ...our Series B and have grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are inventors. Ivo was first-to-...  ...expect us to hit our SLAs. What? We’re looking for an Senior Site level Reliability Engineer as part of Infrastructure team to: Own uptime,... 
    Contract work
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours

    Ivo Inc.

    San Francisco, CA
    1 day ago
  • $100k - $120k

     ...position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in United States. The Site Reliability Engineer will play a critical role in improving platform reliability,... 
    Full time
    Temporary work
    Remote work
    Flexible hours

    Jobgether

    United States
    1 day ago
  • $87.4k - $123.4k

     ...the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.*** The Site Reliability Engineer will help ensure the reliability, scalability, and performance of Empower’s financial services platform. This person will... 
    16 hours
    Full time
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Work at office
    Local area
    Remote work
    Work from home
    Work visa
    Flexible hours

    Empower

    United States
    1 day ago
  • $104.43k - $156.65k

     ...Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual...  ..., Fox, Disney, NBC, Paramount+, and many others. Our Site Reliability Engineering (SRE) team is at the heart of our mission to deliver... 
    Permanent employment
    Full time
    Work experience placement
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Comcast

    Centennial, CO
    20 hours ago
  • $150k - $190k

     ...to help shape how we operate. You'll be responsible for the reliability, performance, and availability of Develocity instances serving...  ...Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you... 
    Full time
    Remote work
    Work from home
    Shift work

    Develocity

    United States
    20 hours ago
  •  ...cloud infrastructure. You'll lead initiatives that improve reliability, scalability, and operational excellence. Key...  ...management tools. Required Qualifications • 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps. • Hands-on experience... 
    Full time
    Remote work

    IPolarity LLC

    United States
    1 day ago
  • $230k - $255k

    Join Aya Healthcare, winner of multiple Top Workplace awards! We're looking for a highly experienced Manager, Site Reliability Engineering to lead the team behind one of healthcare's most relied-on workforce platforms. In this leadership role, you'll guide and grow a... 
    Full time
    Local area
    Remote work

    Aya Healthcare

    United States
    1 day ago
  • $187k - $243k

     ...through early diagnosis and longitudinal care management of chronic conditions. We're looking for a Senior Manager of Site Reliability Engineering to join our team. You'll lead a team of ~10 SREs across North America, UK, HK, and New Zealand — owning both the day-to-... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    Shift work

    Clover Health

    United States
    1 day ago
  • $175k - $250k

     ...been customized and developed by our expert team of lawyers, engineers and research scientists. We’ve found product market fit and...  ...compensation. Role Overview As a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability... 
    Full time
    Relocation package

    Harvey

    Remote
    1 day ago
  • $140k - $230k

     ...Zoox is seeking a Site Reliability Engineer to help ensure the availability, performance, and resilience of the services that power the development and operation of our autonomous vehicles. In this role, you will own the full lifecycle of our services—from designing fault... 
    Full time

    Zoox

    Remote
    1 day ago
  • Role Description We are looking for a Site Reliability Engineer (SRE) who is passionate about infrastructure reliability, automation, and building scalable production systems. ~Own and improve production infrastructure reliability and stability ~Prepare, execute, and... 
    Full time
    Remote work

    Social Discovery Group

    Remote
    2 days ago
  • $106.5k - $177.5k

    Role Description The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing on the seamless integration, scalability, and long-term reliability... 
    Full time
    Remote work

    Noctua Technology

    Remote
    19 days ago
  •  ...to grow, adapt and use your skills consistently. Our customers rely on us in the moments that matter. Engineering delivers on that promise. The Senior Site Reliability Engineer is responsible for ensuring our SaaS products are fast, stable and optimized for our... 
    Full time
    Work experience placement

    Donnelley Financial Solutions

    Remote
    3 days ago
  • Role Description We are expanding our Site Reliability Engineering (SRE) team and seeking a highly skilled and passionate Senior SRE to join us. As a member of our growing SRE function, you will play a critical role in ensuring the reliability, scalability, and performance... 
    Full time
    Temporary work

    QAD, Inc.

    Remote
    4 days ago
  • Role Description Join us as a Senior Site Reliability Engineer on our mission to turn payments into possibilities! The Site Reliability Engineering (SRE) team ensures the reliability, availability, scalability, and performance of a mission-critical payment orchestration... 
    Full time
    Immediate start
    Remote work

    CellPoint Digital

    Remote
    4 days ago
  • Role Description We’re looking for a Senior Site Reliability Engineer who takes ownership seriously — someone who designs for reliability, ships the automation, and stands behind it in production. You’ll work across cloud-native infrastructure on systems that process millions... 
    Full time

    CertifyOS

    Remote
    4 days ago
  • Role Description As a Site Reliability Engineer on the Central AI team, you will help Health Catalyst engineer teams adopt AI responsibly and effectively. You bring deep experience solutioning and implementing AI systems, and you use that expertise to evaluate architectures... 
    Full time

    Health Catalyst

    Remote
    2 days ago
  • $130k - $170k

     ...Senior Site Reliability Engineer About Us Founded in 2014, we offer the industry’s first and only cloud‑based, fully‑customisable, end‑to‑end software solution to automate securities‑based lending from origination through the life of the loan. By combining thought... 
    Full time
    Flexible hours
    Shift work

    Supernova Technology™

    Chicago, IL
    1 day ago
  • $114k - $148k

    Role Description As a Site Reliability Engineer, you will focus on ensuring the platform and services customers rely on are reliable, performant...  ...other team members as needed. You will interact with internal staff, managers, and customers to implement and maintain... 
    Full time
    Temporary work
    Work experience placement

    OneStream Software

    Remote
    6 days ago
  • Role Description Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and services. Our SRE teams solve reliability, security, and usability at scale for... 
    Full time
    Work at office

    Akamai

    Remote
    7 days ago
  • $110k - $137.49k

    Role Description The Sr Site Reliability Engineer, Release will prototype, write, maintain, and test code in multiple stages of the release process and in multiple environments in order to rapidly deliver automated solutions to our application releases. Assesses unusual... 
    Full time
    Work experience placement
    Remote work

    Alkami Technology

    Remote
    4 days ago
  •  ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD... 
    Full time
    Remote work

    SecurityScorecard

    New York, NY
    1 day ago
  • $126k - $248k

     ...As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and... 
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    New Jersey
    4 days ago
  • $130k - $180k

     ...of building large in-house AI/ML infrastructure. Built by engineers, for engineers. From large-scale GPU orchestration to inference...  ...and AI R&D. THE ROLE Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work

    Nebius

    Des Moines, IA
    2 days ago
  • $190k - $240k

    Role Description As a Sr. Site Reliability Engineer (SRE) at ICD, you will play a critical role in ensuring the reliability and seamless operation of our global platform and AWS infrastructure to create scalable and highly reliable software systems. Job Responsibilities... 
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tradeweb

    Remote
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!