Director of Site Reliability Engineering

$205k - $305k

Stellar

Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high-scale today. Developers and companies around the world build on it, and the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem.

SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production services.

This is a senior engineering leadership role reporting to the CTO. You will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence.

Engineering teams at SDF own the services they build. SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.

You will be successful here if you bring strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution. SDF is a small, mission-driven foundation with a broad technical surface area, so this role requires leverage, ownership, and a bias toward solving the right problems over creating processes for its own sake.

In this role, you will:

Lead, coach, and develop a distributed SRE team, setting a clear vision, charter, operating model, priorities, and success measures.
Define and roll out a Service Ownership & Maturity Framework across engineering, with expectations that vary appropriately by service criticality.
Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices.
Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence.
Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk.
Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team.
Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability.
Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering.
Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction.

You have:

10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles.
5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams.
Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk.
3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments.
3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety.
Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
Experience helping product or application engineering teams improve service ownership, operational readiness, and production accountability.
A pragmatic approach to tooling: you understand when to build, buy, adapt, simplify, or retire systems based on the actual engineering problem.
The ability to operate effectively in a small or mid-sized engineering organization where influence comes from credibility, judgment, and outcomes rather than bureaucracy.
Clear executive communication skills and the ability to partner directly with a CTO and senior engineering leaders.

Bonus Points if:

Experience leading SRE, infrastructure, or platform work in a lean, high-agency organization.
Experience supporting globally distributed teams or 24/7 operational coverage.
Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil.
Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls.
Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high-reliability technical ecosystems.
Experience evaluating vendors and infrastructure platforms with skepticism, technical rigor, and cost discipline.
Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity.

We offer competitive pay with a base salary range for this position of $205,000 - $305,000 depending on job-related knowledge, skills, experience, and location. In addition, we offer lumen-denominated grants along with the following perks and benefits:

USA Benefits/Perks:

Competitive health, dental & vision coverage with most plans covered at 100% for the employee + any dependents
Flexible time off + 15 company holidays including a company-wide holiday break
Generous paid parental leave for all parents, plus paid pregnancy disability leave for birthing parents
Gym reimbursement ($80 per month)
Life & ADD (up to $50K)
Short & Long term disability
401K with 4% match
Health & Dependent Care FSA Accounts
Commuter benefits with $250/month employer contribution
Health Savings Account (HSA) with monthly employer contribution
Family building benefits through Kindbody
Wellbeing benefits (One Medical, Rightway, Headspace)
L&D budget of $1,500/year
Daily lunch and snacks in office
Company retreats

#LI-Hybrid

About Stellar

Stellar is more than a blockchain. Powered by a decentralized, fast, scalable, and uniquely sustainable network made for financial products and services and a thriving and passionate ecosystem that includes a non-profit organization driven by a mission, Stellar is paving the path to unlock the world's economic potential through blockchain technology. Built with speed and low costs in mind, the Stellar network provides builders and financial institutions worldwide a platform to issue assets, and to send and convert currencies in real time creating real world utility. Founded in 2014, the Stellar Development Foundation (SDF) supports the continued development and growth of the Stellar network and also serves the ecosystem of NGOs, corporations, universities, small businesses, governments, and solo entrepreneurs building on the Stellar network through tooling, funding and strategic collaborations. Together, Stellar is where blockchain meets the real world.

About the Stellar Development Foundation

The Stellar Development Foundation (SDF) is a non-profit organization focused on working with and supporting change-makers to create equitable access to the global financial system through blockchain technology. SDF provides grants, investments, funding, and other awards to builders and organizations. SDF also develops resources and tooling on the Stellar network to help unlock real world utility. As a nonprofit foundation, SDF puts the health of the Stellar network and the Stellar ecosystem and its mission above all else.

We look forward to hearing from you!

Privacy Policy

By submitting your application, you are agreeing to our use and processing of your data in accordance with our Privacy Policy.

SDF is committed to diversity in its workforce and is proud to be an equal opportunity employer. SDF does not make hiring or employment decisions on the basis of race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other basis protected by applicable local, state or federal law.

Apply

Vacancy posted 1 day ago

Similar jobs that could be interesting for youBased on the Director of Site Reliability Engineering in San Francisco, CA vacancy

Site Reliability Engineer
...shape the future of healthcare, we’d love to meet you. About the role We’re hiring an SRE to join our engineering team at Plenful and take ownership of the reliability and performance of the systems that power our product. You’ll work across our distributed workflow...
Suggested
Work at office
Remote work
Flexible hours
2 days per week
Plenful
San Francisco, CA
3 days ago
Director of SRE: Lead Reliability & Platform Engineering
Stellar is seeking a Director of Site Reliability Engineering to lead a distributed SRE team and shape service operations. This role is crucial for improving the reliability and operational maturity of services within the Stellar ecosystem. The ideal candidate will have...
Suggested
Stellar
San Francisco, CA
4 days ago
Site Reliability Engineer
$150k
...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our...
Suggested
VantageScore
San Francisco, CA
4 days ago
Site Reliability Engineer
$150k - $250k
...Site Reliability Engineer role USC or GC only are considered at this time. San Francisco - Local to Bay area only but role is remote and occasion meeting required Latest update, 03/31/2026: The Site Reliability Engineer role is critical for...
Suggested
Work experience placement
Casual work
Local area
Immediate start
Remote work
3B Staffing LLC
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
...advanced algorithms that significantly outperforms individual engineers. We combine language models with human ingenuity to push the... ...quality. The Role: We are seeking an experienced Site Reliability Engineer to join our Platform Engineering team in the Bay Area...
Suggested
CodeRabbit
San Francisco, CA
1 day ago
Senior Site Reliability Engineer
$181.69k - $213.75k
...Senior Site Reliability Engineer San Francisco, California; Santa Clara, California; Seattle, WA The Company You'll Join Carta connects founders, investors, and limited partners through world-class software, purpose-built for everyone in venture capital, private...
Full time
Work at office
Carta
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate...
Alembic Technologies
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and...
Immediate start
Remote work
Worldwide
OutSystems
San Francisco, CA
1 day ago
Site Reliability Engineer
$260k - $300k
...when it does, to make sure it's resolved faster than anyone expects. You will own both the production reliability of our user-facing products and the platform engineering that lets our team ship quickly and confidently. That means SLOs, incident response, and on-call on...
Cognition AI
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
...experiment constantly as we find the right paths in an AI-native landscape. The Role You'll be the infrastructure and reliability engineer on the Data Replication team - a full-stack product team running over 3 million sync jobs a week powering thousands of data...
Local area
Airbyte
San Francisco, CA
23 hours ago
Site Reliability Engineer
...Open Source LLM Gateway Engineer LiteLLM is an open-source LLM Gateway with 34K+ stars on GitHub and trusted by companies like NASA... ...expanding and seeking our 6th Engineer focused on owning reliability, performance, and infrastructure stability for the LiteLLM proxy...
BerriAI
San Francisco, CA
23 hours ago
Site Reliability Engineer
...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging...
Worldwide
Shift work
Happy Robot
San Francisco, CA
1 day ago
Senior Site Reliability Engineer
...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient...
TechChain Talent
San Francisco, CA
23 hours ago
Senior Site Reliability Engineer
$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team... ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building...
Work at office
Immediate start
Worldwide
Monday to Friday
Flexible hours
Drata Inc
San Francisco, CA
1 day ago
Senior Software Engineer - Site Reliability Engineering
...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it...
Temporary work
Work experience placement
Phenom People
San Francisco, CA
4 days ago
Senior Site Reliability Engineer
...founders with PhDs in AI, Math, and Computer Science - is poised to redefine computing. About the Role We're seeking a Site Reliability Engineer to ensure Hyperbolic's GPU marketplace and AI infrastructure operate with exceptional reliability, performance, and...
Hyperbolic Labs
San Francisco, CA
4 days ago
Site Reliability Engineer (SRE)
$170k - $230k
...Site Reliability Engineer (SRE) Palo Alto / San Francisco Bay Area About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible and affordable for the world's leading enterprises, AI startups, and the AI research community,...
Work at office
Local area
1 day per week
Mithril
San Francisco, CA
4 days ago
Site Reliability Engineer
$230k - $310k
...millions of daily users while enabling our engineering teams to ship fast. You'll own the... ...building automation and tooling that improves reliability and partnering with engineering to... ...What You'll Bring ~5+ years in site reliability engineering, DevOps, or systems...
Full time
Work at office
Work from home
Gamma
San Francisco, CA
3 days ago
Sr. Site Reliability Engineer
$106k - $130k
...sponsorship. Overall Purpose To create and maintain the next generation of application infrastructure and to be responsible for reliability, automation and scalability using and the latest best practices. Essential Functions Implement software and tools to...
Hourly pay
Work experience placement
Work at office
Immediate start
Visa sponsorship
Work visa
Flexible hours
Early Warning Services
San Francisco, CA
3 days ago
Site Reliability Engineer
...access to life-saving treatment. What We Look for in a Great Engineer You have the intensity and technical mastery to own... ...support high-velocity feature release while maintaining the highest reliability. DevX Support: Support Developer Experience (DevX) work to...
Work at office
Latent
San Francisco, CA
1 day ago
Senior Site Reliability Engineer, Fleet Management
$127k - $249k
...The Team Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....
Work at office
Local area
Remote work
Worldwide
Flexible hours
MongoDB
San Francisco, CA
23 hours ago
Site Reliability Engineer
...$10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices. About the Role As a Site Reliability Engineer (SRE) at Mercor, you'll own production reliability across our most critical systems, partnering directly with infrastructure...
Work at office
Relocation package
Mercor Alabaster
San Francisco, CA
4 days ago
Site Reliability Engineer
...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform. You'll be building and operating the core systems that power agentic AI at scale. Your mission: keep...
Blaxel, Inc
San Francisco, CA
2 days ago
Site Reliability Engineer
...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...
Contract work
Local area
InterSources
San Francisco, CA
4 days ago
Site Reliability Engineer
...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...
Relocation package
1872 Consulting
San Francisco, CA
4 days ago
Site Reliability Engineer II
$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...
Work at office
Restaurant365
San Francisco, CA
3 days ago
Senior Site Reliability Engineer
$195k - $240k
...Senior Site Reliability Engineer San Francisco (Hybrid) At You.com, we are building the AI Search Infrastructure that powers modern AI systems. Our goal is to create the trusted knowledge layer that agents, applications, and enterprises rely on to retrieve real-time...
Full time
Immediate start
Remote work
Work from home
Flexible hours
Y.O.U.
San Francisco, CA
4 days ago
Site Reliability Engineer
...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you will help Shipt understand where we can improve stability and reliability. There will be a focus on the intersection of systems...
BayOne Solutions
San Francisco, CA
1 day ago
Senior Site Reliability Engineer
$117k - $209.33k
...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a new...
For contractors
Autodesk
San Francisco, CA
23 hours ago
Site Reliability Engineer
...complex, highest ROI healthcare workflows. We're actively hiring as we continue to scale. About the role We're hiring a Site Reliability Engineer (SRE) to ensure the reliability, performance, and scalability of Plenful's production systems as we continue to grow....
Work at office
Remote work
Flexible hours
2 days per week
Plenful
San Francisco, CA
4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director of Site Reliability Engineering. Be the first to apply!