Senior Site Reliability Engineer
$185.5k - $232kFormation Bio (Formerly TrailSpark)
Senior Site Reliability Engineer
New York, NY; Boston, MA; San Francisco, CA
About Formation Bio
Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development.
Advancements in AI and drug discovery are creating more candidate drugs than the industry can progress because of the high cost and time of clinical trials. Recognizing that this development bottleneck may ultimately limit the number of new medicines that can reach patients, Formation Bio, founded in 2016 as TrialSpark Inc., has built technology platforms, processes, and capabilities to accelerate all aspects of drug development and clinical trials. Formation Bio partners, acquires, or in-licenses drugs from pharma companies, research organizations, and biotechs to develop programs past clinical proof of concept and beyond, ultimately helping to bring new medicines to patients. The company is backed by investors across pharma and tech, including a16z, Sequoia, Sanofi, Thrive Capital, John Doerr, Spark Capital, SV Angel Growth, and others.
At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently.
About the Position
As a Senior Site Reliability Engineer at Formation Bio, you will build and operate the infrastructure, delivery systems, and operational practices that allow our engineering organization to ship reliable software quickly and safely. You will work across cloud infrastructure, developer platforms, observability, and production workloads, including product applications, internal tools, data systems, and ML and AI workloads, taking problems from initial diagnosis through implementation and production adoption.
Formation Bio is an AI-native engineering organization. You will use modern AI tools, including agentic coding systems, as part of your daily engineering practice while applying the judgment needed to validate their output and operate reliable production systems. You will partner closely with Product Engineering, Data Engineering, and Data Science to design, build, and develop the infrastructure and operational platform that supports product software, data systems, and ML and AI workloads, with the opportunity to shape how the broader organization builds and operates AI-native systems.
Responsibilities
- Own the infrastructure and operational platform for shared engineering workloads, including compute, runtime environments, orchestration, deployment, observability, access controls, and reliability.
- Build and operate secure, observable, reliable infrastructure for product applications, containerized services, internal tools, data systems, ML pipelines, inference, and agentic software.
- Research, develop, and maintain core AWS infrastructure and additional cloud outposts for services across development, staging, and production, including compute, networking, databases, load balancers, and secrets management.
- Create, review, maintain, and optimize infrastructure as code, CI/CD pipelines, and reusable platform patterns.
- Establish strong operational practices, including SLOs, monitoring, alerting, runbooks, incident response, advanced diagnostics, root cause analysis, and post-incident follow-through.
- Work with Product engineering, Data Engineering, and Data Science to evaluate, negotiate, and implement architecture and infrastructure for product software, data systems, model training, and inference based on functional requirements
- Use AI tools to accelerate infrastructure development, investigate incidents, improve documentation, build automation, and make operational improvements while validating their output.
- Participate in the support rotation and incident response. Maintain a strong bias for automation, the patience for ClickOps, and the experience to know when each is appropriate.
- Write and review requirements, design documents, and operating procedures. Share knowledge and mentor engineers on infrastructure and SRE fundamentals.
About You
- 5+ years of relevant experience in Site Reliability Engineering, infrastructure, systems, DevOps, or a similar discipline.
- Production experience operating cloud infrastructure and distributed systems, with strong operational and reliability judgment.
- Experience with advanced diagnostics, incident response, root cause analysis, observability, and automation.
- Experience with AWS and Snowflake. Experience with Azure, GCP, and/or Vercel is a plus.
- Working experience with Docker, GitHub, Kubernetes, Python, Terraform or OpenTofu, and virtual networking. Terragrunt is a plus
- Experience supporting production ML or AI workloads, MLOps infrastructure, workflow orchestration, model serving, or a related platform is a plus. Strong SRE and infrastructure fundamentals are the core requirements.
- Experience managing shared COTS and FOSS software applications in addition to our internal software and tools
- Daily fluency with AI tools, including LLMs and agentic coding systems, paired with strong engineering judgment and a high bar for validation.
- Exceptional collaboration and communication skills across engineering, Data Science, Data Engineering, Security, and non-technical partners.
- Experience operating infrastructure in a regulated or validated environment is a plus.
Total Compensation Range: $185,500 - $232,000
Compensation Individual compensation is determined by several factors, including role scope, geographic location, and skills & experience. Your offer will reflect where you fall within the range based on these considerations. In addition to base salary, we offer equity, comprehensive benefits, and generous perks. If the posted range doesn't match your expectations, we still encourage you to apply!
Where We Hire Formation Bio is prioritizing hiring in key hubs, primarily the New York City and Boston metro areas, with a hybrid model requiring 3 days per week in office. Applicants from the Research Triangle (NC) and San Francisco Bay Area may also be considered. Please apply only if you reside in these locations or are willing to relocate.
Equal Opportunity Formation Bio is committed to building a diverse and inclusive team. We are an equal opportunity employer and welcome candidates from all backgrounds. All qualified applicants will receive consideration for employment without regard to race, color, creed, religion, national origin, ancestry, sex (including pregnancy, childbirth, breastfeeding, and related medical conditions), gender identity or expression, sexual orientation, age, disability, genetic information, marital status, military or veteran status, or any other characteristic protected by federal, state, or local law.
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy automating...SeniorFull time$182.8k - $247.3k
...changing mission to develop education for our half a billion (and growing!) learners around the world.About the role...As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed...SeniorWork experience placement- The Role:GIPHY is seeking a highly experienced Site Reliability Engineer to join our SRE team. You will help design, build, operate, and evolve the infrastructure that powers GIPHY, including our cloud environment, Kubernetes clusters, and CI/CD platforms.You will also...SeniorFull timeWork experience placementRemote work
$158.5k - $172k
...the exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate,... ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire...SeniorFull timeTemporary workWork at officeFlexible hours3 days per week- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....SeniorLocal area
$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...SeniorFull timeTemporary workLocal areaRemote workWorldwide$150k - $170k
...Senior Site Reliability Engineer – Zip Co Join to apply for the Senior Site Reliability Engineer role at Zip Co At Zip, we build cloud‑native software applications that serve millions of customers and process billions of dollars in payments. We’re looking for...SeniorCasual workWork at officeRemote workFlexible hours- ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure...SeniorFull timeWork at officeRemote workFlexible hours2 days per week
$189k - $283.6k
...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience Technologies...SeniorFull timeLocal areaRemote workRelocation packageFlexible hoursShift work$500 per month
...accounts. Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to... ...significant impact, we encourage you to apply. Your Role: As a Site Reliability Engineer at Alpaca, you'll help keep our brokerage platform...SeniorHome office$191k - $226k
....S. — and using AI to scale that impact further and faster than anyone else can. About the role: We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner's products and AI/ML workloads...SeniorRemote workWork visaFlexible hours- ...Job Description A major financial services company in NYC is growing its team rapidly, and they are looking for a Senior DevOps Engineer / Site Reliability Engineer who can join. If you’re passionate about high-availability, reliability, automation, we’d be excited...Senior
- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure...Senior
$104k - $178k
...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The team is responsible for building and maintaining the reliability, scalability, and performance of DV's digital media measurement...Senior$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...eliminated manual deployment processes.At the Senior Engineer level, your influence is...SeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week- ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have: - Site Reliability Engineering and system reliability optimization - GitLab CI/CD and HashiCorp Vault secrets management -...Contract work
$400k
...in financial markets, the organization combines innovation, engineering excellence, and data-driven insights to support complex trading operations worldwide. This opportunity is for a Senior Site Reliability Engineer to join a high-performance infrastructure...SeniorPermanent employmentWorldwide$168k - $200k
...that is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable,...Senior$120k - $150k
...allows each person to achieve personal success and add value to our teams and communities.We are currently looking for a Site Reliability Engineer to join our Platform Engineering team in New York, NY.About the RoleJoin our Platform Engineering team as a Site Reliability...Full time- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad...Shift work
$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the cloud. They ensure...Work at officeLocal areaImmediate start$207k - $300k
...providing feedback to ensure best practices in reliability, security, and efficiency.Triage and... ...development initiatives. Mentor other engineers and contribute to the engineering... ...principles to cloud environments; and Managing senior stakeholders, external partners, and...Full timeWork at office- Skip To Main ContentBack to SearchRemote in United States of America: New YorkSite Reliability Engineering& 4 othersLooking for something else?Find a vacancy that works for you. Send us your CV to receive a personalized offer.Find me a jobWe are seeking a specialized Observability...
$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$190k - $260k
...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving...Full timeWork experience placementWork at officeLocal areaRemote workHome office$260k - $300k
...software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding... ...faster than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that...$207k - $300k
...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...:Master's degree in Computer Science or Engineering.Experience mentoring engineers and cultivating... ...consensus across cross-functional teams.Site Reliability Engineering (SRE) combines...- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling...Flexible hours
- ...significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment... ..., with clear accountability for outcomes.Act as a senior escalation point during critical incidents, driving...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- site reliability engineer sre New York, NY
- site reliability engineer New York, NY
- site reliability engineer remote New York, NY
- senior operations technician New York, NY
- senior operations associate New York, NY
- senior cloud service delivery manager New York, NY
- senior it service manager New York, NY
- senior project engineer New York, NY
- senior chief engineer New York, NY
- sr operations manager New York, NY



