Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$185.5k - $232k

Formation Bio (Formerly TrailSpark)

Senior Site Reliability Engineer

New York, NY; Boston, MA; San Francisco, CA

About Formation Bio

Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.

Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials. Recognizing that this development bottleneck may ultimately limit the number of new medicines that can reach patients, Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. Formation Bio partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients. The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, John Doerr, Spark Capital, SV Angel Growth, and others.

At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.

About the Position

As a Senior Site Reliability Engineer at Formation Bio, you will build and operate the infrastructure, delivery systems, and operational practices that allow our engineering organization to ship reliable software quickly and safely. You will work across cloud infrastructure, developer platforms, observability, and production workloads, including product applications, internal tools, data systems, and ML and AI workloads, taking problems from initial diagnosis through implementation and production adoption.

Formation Bio is an AI-native engineering organization. You will use modern AI tools, including agentic coding systems, as part of your daily engineering practice while applying the judgment needed to validate their output and operate reliable production systems. You will partner closely with Product Engineering, Data Engineering, and Data Science to design, build, and develop the infrastructure and operational platform that supports product software, data systems, and ML and AI workloads, with the opportunity to shape how the broader organization builds and operates AI-native systems.

Responsibilities
  • Own the infrastructure and operational platform for shared engineering workloads, including compute, runtime environments, orchestration, deployment, observability, access controls, and reliability.
  • Build and operate secure, observable, reliable infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software.
  • Research, develop, and maintain core AWS infrastructure and additional cloud outposts for services across development, staging, and production, including compute, networking, databases, load balancers, and secrets management.
  • Create, review, maintain, and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns.
  • Establish strong operational practices, including SLOs, monitoring, alerting, runbooks, incident response, advanced diagnostics, root cause analysis, and post-incident follow-through.
  • Work with Product engineering, Data Engineering, and Data Science to evaluate, negotiate, and implement architecture and infrastructure for product software, data systems, model training, and inference based on functional requirements
  • Use AI tools to accelerate infrastructure development, investigate incidents, improve documentation, build automation, and make operational improvements while validating their output.
  • Participate in the support rotation and incident response. Maintain a strong bias for automation, the patience for ClickOps, and the experience to know when each is appropriate.
  • Write and review requirements, design documents, and operating procedures. Share knowledge and mentor engineers on infrastructure and SRE fundamentals.
About You
  • 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline.
  • Production experience operating cloud infrastructure and distributed systems, with strong operational and reliability judgment.
  • Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation.
  • Experience with AWS and Snowflake. Experience with Azure, GCP, and/or Vercel is a plus.
  • Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking. Terragrunt is a plus
  • Experience supporting production ML or AI workloads, MLOps infrastructure, workflow orchestration, model serving, or a related platform is a plus. Strong SRE and infrastructure fundamentals are the core requirements.
  • Experience managing shared COTS and FOSS software applications in addition to our internal software and tools
  • Daily fluency with AI tools, including LLMs and agentic coding systems, paired with strong engineering judgment and a high bar for validation.
  • Exceptional collaboration and communication skills across engineering, Data Science, Data Engineering, Security, and non-technical partners.
  • Experience operating infrastructure in a regulated or validated environment is a plus.

Total Compensation Range: $185,500 - $232,000

Compensation Individual compensation is determined by several factors, including role scope, geographic location, and skills & experience. Your offer will reflect where you fall within the range based on these considerations. In addition to base salary, we offer equity, comprehensive benefits, and generous perks. If the posted range doesn't match your expectations, we still encourage you to apply!

Where We Hire Formation Bio is prioritizing hiring in key hubs, primarily the New York City and Boston metro areas, with a hybrid model requiring 3 days per week in office. Applicants from the Research Triangle (NC) and San Francisco Bay Area may also be considered. Please apply only if you reside in these locations or are willing to relocate.

Equal Opportunity Formation Bio is committed to building a diverse and inclusive team. We are an equal opportunity employer and welcome candidates from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, national origin, ancestry, sex (including pregnancy, childbirth, breastfeeding, and related medical conditions), gender identity or expression, sexual orientation, age, disability, genetic information, marital status, military or veteran status, or any other characteristic protected by federal, state, or local law.

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in San Francisco, CA vacancy
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    2 days ago
  • $190.8k - $267.1k

     ...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser...  ...team partners closely with Ads Engineering to improve reliability, scalability, operational...  ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,... 
    Senior
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    2 days ago
  • $117k - $209.33k

    Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB...  ...a pivotal role in engineering the reliable, globally connected, multi-cloud...  ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    2 days ago
  • $165k - $227k

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  •  ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ...ownership, working together to build scalable, reliable, and secure products that empower...  ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  •  ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 
    Senior

    Alembic Limited

    San Francisco, CA
    17 hours ago
  • $175k - $250k

     ...000.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance...  ...scalability, performance, and reliability across environments. What You’ll Do Design... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    3 days ago
  •  ...of healthcare, we'd love to meet you. Apply now to join our growing team. About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    1 day ago
  •  ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated access to business data across hundreds of sources, so they can discover the entities that matter, reason over real-time... 
    Senior
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    2 days ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content...  ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering... 
    Senior

    Hive

    San Francisco, CA
    1 day ago
  • $55k - $151.47k

     ...ApplicableSpecialismIFS - Internal Firm Services - OtherManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a Site Reliability Engineer - Senior Associate, you will play a pivotal role in enhancing the reliability, scalability, and performance of our... 
    Senior
    Full time
    H1b

    PwC

    San Francisco, CA
    19 hours ago
  • $181k - $225k

     ...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time... 
    Senior
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    4 days ago
  • Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...
    Senior
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    1 day ago
  •  ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'... 
    Senior

    Alembic

    San Francisco, CA
    4 days ago
  • $167.7k - $245.2k

     ...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    4 days ago
  • $232k - $319k

     ...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and...  ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    6 hours ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...The Team Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Role Overview We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Senior
    Full time
    Work at office
    Remote work
    Worldwide

    Mongodb

    San Francisco, CA
    15 days ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    6 hours ago
  • $139.76k - $287.75k

     ...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,...  ...will be instrumental in advancing the reliability, scalability, automation, observability...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Senior
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    1 day ago
  • $106k - $130k

     ..., for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and... 
    Senior
    Hourly pay
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning

    San Francisco, CA
    2 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites. About the Role Fivetran is building... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    4 days ago
  • $262k - $364k

    Lead a team of software/systems engineers on projects for users and be directly responsible for uptime.Own end-to-end availability...  ...Engineering.Experience with machine learning infrastructure.Site Reliability Engineering (SRE) combines software and systems engineering... 
    Senior

    Google

    San Bruno, CA
    2 days ago
  • $300k

     ...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely...  ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    1 day ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at... 
    Temporary work

    TextNow

    San Francisco, CA
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!