Director, Site Reliability Engineering
$205k - $305kStellar
Director Of Site Reliability Engineering
Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high-scale today. Developers and companies around the world build on it, and the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem.
SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production services.
This is a senior engineering leadership role reporting to the CTO. You will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence.
Engineering teams at SDF own the services they build. SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.
You will be successful here if you bring strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution. SDF is a small, mission-driven foundation with a broad technical surface area, so this role requires leverage, ownership, and a bias toward solving the right problems over creating processes for its own sake.
In this role, you will:
- Lead, coach, and develop a distributed SRE team, setting a clear vision, charter, operating model, priorities, and success measures.
- Define and roll out a Service Ownership & Maturity Framework across engineering, with expectations that vary appropriately by service criticality.
- Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
- Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices.
- Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence.
- Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk.
- Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team.
- Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability.
- Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering.
- Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction.
You have:
- 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles.
- 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
- Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams.
- Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk.
- 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments.
- 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety.
- Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
- Experience helping product or application engineering teams improve service ownership, operational readiness, and production accountability.
- A pragmatic approach to tooling: you understand when to build, buy, adapt, simplify, or retire systems based on the actual engineering problem.
- The ability to operate effectively in a small or mid-sized engineering organization where influence comes from credibility, judgment, and outcomes rather than bureaucracy.
- Clear executive communication skills and the ability to partner directly with a CTO and senior engineering leaders.
Bonus points if:
- Experience leading SRE, infrastructure, or platform work in a lean, high-agency organization.
- Experience supporting globally distributed teams or 24/7 operational coverage.
- Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil.
- Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls.
- Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high-reliability technical ecosystems.
- Experience evaluating vendors and infrastructure platforms with skepticism, technical rigor, and cost discipline.
- Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity.
We offer competitive pay with a base salary range for this position of $205,000 - $305,000 depending on job-related knowledge, skills, experience, and location. In addition, we offer lumen-denominated grants along with the following perks and benefits:
- Competitive health, dental & vision coverage with most plans covered at 100% for the employee + any dependents
- Flexible time off + 15 company holidays including a company-wide holiday break
- Generous paid parental leave for all parents, plus paid pregnancy disability leave for birthing parents
- Gym reimbursement ($80 per month)
- Life & ADD (up to $50K)
- Short & Long term disability
- 401K with 4% match
- Health & Dependent Care FSA Accounts
- Commuter benefits with $250/month employer contribution
- Health Savings Account (HSA) with monthly employer contribution
- Family building benefits through Kindbody
- Wellbeing benefits (One Medical, Rightway, Headspace)
- L&D budget of $1,500/year
- Daily lunch and snacks in office
- Company retreats
About Stellar
Stellar is more than a blockchain. Powered by a decentralized, fast, scalable, and uniquely sustainable network made for financial products and services and a thriving and passionate ecosystem that includes a non-profit organization driven by a mission, Stellar is paving the path to unlock the world's economic potential through blockchain technology. Built with speed and low costs in mind, the Stellar network provides builders and financial institutions worldwide a platform to issue assets, and to send and convert currencies in real time creating real world utility. Founded in 2014, the Stellar Development Foundation (SDF) supports the continued development and growth of the Stellar network and also serves the ecosystem of NGOs, corporations, universities, small businesses, governments, and solo entrepreneurs building on the Stellar network through tooling, funding and strategic collaborations. Together, Stellar is where blockchain meets the real world.
About the Stellar Development Foundation
The Stellar Development Foundation (SDF) is a non-profit organization focused on working with and supporting change-makers to create equitable access to the global financial system through blockchain technology. SDF provides grants, investments, funding, and other awards to builders and organizations. SDF also develops resources and tooling on the Stellar network to help unlock real world utility. As a nonprofit foundation, SDF puts the health of the Stellar network and the Stellar ecosystem and its mission above all else.
We look forward to hearing from you!
Privacy Policy
By submitting your application, you are agreeing to our use and processing of your data in accordance with our Privacy Policy.
SDF is committed to diversity in its workforce and is proud to be an equal opportunity employer. SDF does not make hiring or employment decisions on the basis of race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other basis protected by applicable local, state or federal law.
$120k - $142k
...on making journalism so good that it’s worth paying for. Mission Overview & Responsibilities: At The New York Times, our Site Reliability Engineering (SRE) team is central to how we design, test, and operate the systems that support our most critical customer...SuggestedLocal areaFlexible hours$151k - $297k
As a TPM for SRE, you will partner with SRE leaders and engineers to scale the platform that underpins all of MongoDB’s cloud products. You will drive program execution, strengthen production reliability practices, and coordinate cross-functional efforts across US and EMEA...SuggestedLocal areaRemote workWorldwideFlexible hours$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SuggestedFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$141k - $216.6k
...—it means helping shape the future of emergency response and building a safer, more connected world.Position OverviewAs a Site Reliability Engineer, you'll own the reliability, observability, and operational excellence of our Unified Call (UC) platform—the mission-critical...SuggestedWork experience placementWork at office$158.5k - $172k
...exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate, and... .... This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire technology...SuggestedFull timeTemporary workWork at officeFlexible hours3 days per week$120k - $200k
...PermContact: Kunal DaveContact Email: ****@*****.*** Reliability Engineer(SRE) ResponsibilitiesGlobal Architecture & Disaster Recovery... ...practices (e.g., Chaos Engineering, resilience testing, automated recovery)SkillsBilingual Mandarin Site Reliability Engineer(SRE)Overseas- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Commercial & Investment Bank, Production management team, you will solve complex and broad...Shift work
$200k - $250k
Hudson River Trading (HRT) is seeking a Senior Site Reliability Engineer focused on storage to join our growing Enterprise SRE team. This team is responsible for developing and maintaining productivity service infrastructure for the entire firm, both on-prem and in the...Work at officeLocal areaImmediate start$110k - $120k
...largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.Job DescriptionJob Title: Site Reliability Engineer (SRE) / L3 Support EngineerGetting to know us:As a leading financial services and healthcare technology company based on...Ongoing contractFull timeCasual workRemote workFlexible hours$139k - $257.55k
The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...Full timeTemporary workLocal areaRemote workWorldwide- ...globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial & Investment Bank, Production Management team, you hold a leadership...
- ...customers. Cohere is a team of researchers, engineers, designers, and more, who are all... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...advanced NLP applications? We are looking for a Site Reliability Engineer to join the Model Serving...Full timeWork experience placementWork at officeLocal areaRemote workHome office
$194k - $267k
...career-defining work. We're all in on this mission. If you are too, let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable, scalable, and secure cloud services...Local areaWorldwideFlexible hours$194k - $267k
...do something more than once, automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$150k - $250k
What We DoAt Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people... ...markets.Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience,...Full timeTemporary workPart time$131k - $164k
Position OverviewWe are seeking a highly skilled Staff Site Reliability Engineer with deep technical expertise across VMware, Linux, and automation frameworks, to join our global Infrastructure & Operations team. This role is a hands-on senior engineering position responsible...Work at officeLocal areaVisa sponsorshipFlexible hours$194k - $267k
...all in on this mission. If you are too, let's talk.Position Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk ecosystem. In this role, you will move beyond simple monitoring to...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$182k - $250.8k
...Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great... ...for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with...Permanent employmentLocal areaRemote workWorldwideFlexible hoursWeekend workWeekday work- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications. The ideal candidate...Flexible hours
- ...self-healing, deployment/rollback automation). Establish reliability standards: SLOs/SLIs, error budgets, production readiness reviews... ..., and release risk controls. Performance and reliability engineering: capacity planning, load/performance analysis, resilience...
- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support front-office trading systems in a production environment. This role focuses on troubleshooting complex trading infrastructure...
- ...Chariot Engineering Hire Chariot's engineering hire will be responsible for taking the Chariot platform to the next level. You will lead the evolution of our banking product, DAF product and technology strategy, along with building a growing team of software engineers...Work experience placementWork at officeWork from homeMonday to FridayMonday to Thursday
$150k - $175k
...Site Reliability Engineer At ASAPP, our mission is simple: deliver the best AI-powered customer experience—faster than anyone else. To achieve that, we're guided by principles that shape how we think, build, and execute. We value customer obsession, purposeful speed...Remote work- ...DevOps Engineer Responsible for reliability and support of container platform on-prem and external clouds (Azure /AWS /Google) Monitor and troubleshoot container platform environment performance issues, connectivity issues, security issues, etc. Perform deep dives...
- ...human risk—the leading cause of cybersecurity breaches—and build safer, more resilient organizations. The Role: As a Senior Site Reliability Engineer (SRE) at Dune Security, you will play a critical role in ensuring our platform's stability, scalability, and security. You...Full timeWork at office
$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ..., and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...Local area$141.8k - $195k
...work, grow fast, and bring their full selves to the herd. Why You'll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team in the U.S. Cribl provides...Temporary workRemote work- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
- ...subscriptions at scale, combining the agility of a high-growth business with the backing of a global organization. As the Site Reliability Engineer, you will help ensure the reliability, scalability, and observability of CloudBlue’s multi-tenant SaaS platforms used by...Remote workWorldwideFlexible hours
$139k - $257.55k
...The Challenge The Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses...Temporary workLocal areaRemote workWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Director, Site Reliability Engineering. Be the first to apply!
- principal network engineer New York, NY
- principal application developer New York, NY
- director sales engineering New York, NY
- civil engineer project manager New York, NY
- principal security engineer New York, NY
- principal engineer New York, NY
- principal battery engineer New York, NY
- principal infrastructure engineer New York, NY
- director quality engineering New York, NY
- senior civil engineer project manager New York, NY

