Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Director of Site Reliability Engineering

techchaintalent

About Stellar

Stellar is a decentralised, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient than most blockchain-based systems. Since 2014, the Stellar Development Foundation has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high scale today.

About the Role

SDF is hiring a Director of Site Reliability Engineering to lead a team of 4 SREs and shape how engineering teams own, operate, and improve production services. This is a backfill reporting directly to the CTO.

The Director will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence. Engineering teams at SDF own the services they build; SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.

This is a hands-on leadership role. The ideal candidate brings strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution.

Key Responsibilities
  • Lead, coach, and develop a distributed SRE team of 4, setting a clear vision, charter, operating model, priorities, and success measures
  • Define and roll out a Service Ownership and Maturity Framework across engineering, with expectations that vary appropriately by service criticality
  • Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation
  • Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices
  • Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence
  • Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk
  • Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team
  • Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster
  • Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering
  • Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction
Requirements
  • 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles
  • 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers
  • 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments
  • 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety
Bonus Skills
  • Experience leading SRE, infrastructure, or platform work in a lean, high-agency organisation
  • Experience supporting globally distributed teams or 24/7 operational coverage
  • Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil
  • Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls
  • Experience in financial services, regulated environments, blockchain, crypto, or other high-reliability technical ecosystems
  • Experience evaluating vendors and infrastructure platforms with scepticism, technical rigour, and cost discipline
  • Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity
#J-18808-Ljbffr
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Director of Site Reliability Engineering in San Francisco, CA vacancy
  •  ...About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 
    Suggested

    Alembic Limited

    San Francisco, CA
    3 days ago
  • $175k - $250k

     ...00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On-Site only. Must live within commuting distance of...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Suggested
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    16 hours ago
  • $117k - $209.33k

     ...Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting... 
    Suggested
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    16 hours ago
  •  ...millions of daily users while enabling our engineering teams to ship fast. You'll own the...  ...building automation and tooling that improves reliability and partnering with engineering to...  ...services What you'll bring ~5+ years in Site Reliability Engineering, DevOps, or... 
    Suggested
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    1 day ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $24... 
    Suggested
    Full time

    Alembic Technologies

    San Francisco, CA
    16 hours ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • $153k - $191.3k

     ...manufacturing, data processing, and software engineering, our office is a truly inspiring mix of...  ...environments, to guarantee the reliability, scalability, and availability of our services...  ...of Industry and Security and/or the Directorate of Defense Trade Controls. Benefits... 
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    1 day ago
  • $61k - $101k

     ...Salary: $61,000 - 101,000 per year Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong familiarity with SRE culture and the practical application of reliability principles... 
    Full time

    J.P. Morgan

    San Francisco, CA
    4 days ago
  •  ...design of information and operational support systems.  Required Skills/Qualifications: BS/MS degree in Computer Science, Engineering, or a related subject. Equivalent experience accepted.   Proven working experience in installing, configuring, and troubleshooting... 
    Full time
    Work experience placement
    Remote work
    Flexible hours
    San Francisco, CA
    more than 2 months ago
  • $250k

     ...Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves working... 
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $165k - $227k

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    19 days ago
  • $150k

     ...Job Description Job Description About The Role   We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and... 

    VantageScore

    San Francisco, CA
    24 days ago
  •  ...guarantees and certifications. We're hiring staff-level SREs to help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions for how we deploy, observe, and secure the platform, and help... 
    Remote work
    Flexible hours

    Akka

    San Francisco, CA
    26 days ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    3 days ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating... 
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    San Francisco, CA
    3 days ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering... 
    Work at office

    Restaurant365

    San Francisco, CA
    3 days ago
  • $200k - $240k

     ...Senior Site Reliability Engineer (SRE) Location: San Francisco, CA Work Model: Onsite Industry: Renewable Energy Comp: $200,000 - $240,000 We're partnering with a fast-growing energy technology company looking for a Senior Site Reliability Engineer... 

    Lawrence Harvey

    San Francisco, CA
    2 days ago
  • $106k - $130k

     ...employer, at the date of hire. This position is ineligible for employment Visa sponsorship. Overall Purpose The Sr. Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services, LLC

    San Francisco, CA
    1 day ago
  •  ...had design and make. Operate is the phase that tells you what actually happened, and it is ours. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and developer autonomy as we scale our platform. In this role,... 
    Remote work

    MaintainX

    San Francisco, CA
    2 days ago
  • $113.4k - $162k

     ...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is aboutimpactatscale.... 

    TextNow

    San Francisco, CA
    3 days ago
  • $150k - $220k

     ...teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative...  ...can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping... 
    Local area

    Forge Global

    San Francisco, CA
    3 days ago
  •  ...anime content we all love. Join our team, and help us shape the future of anime! About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability,... 

    Engg

    San Francisco, CA
    4 days ago
  • $205k - $305k

     ...the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem. SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production... 
    Temporary work
    Work at office
    Local area
    Worldwide
    Flexible hours

    P2P

    San Francisco, CA
    4 days ago
  • $195k - $257.5k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and operate the infrastructure that powers our blockchain platform at... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    3 days ago
  • $200k - $260k

     ...Join the Cloud Infrastructure Team as a technical leader driving reliability, automation, and scalability across the systems running Sight...  ...and reliability practices across teams, mentor senior engineers, and be a primary escalation point for the org's hardest systems... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    16 hours ago
  • $204k - $306k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta, Inc.

    San Francisco, CA
    1 day ago
  •  ...SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role, you will pair deep...  ...Provide technical leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud Platform Define... 
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    1 day ago
  •  ...SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role, you will pair deep...  ...Provide technical leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud Platform Define... 
    Full time
    Contract work
    Work at office

    Ironclad Inc

    San Francisco, CA
    16 hours ago
  •  ...About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial...  ...Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to... 
    Full time

    Anthropic

    San Francisco, CA
    17 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Director of Site Reliability Engineering. Be the first to apply!