Manager, Site Reliability Engineer
$150k - $220kForge Global
At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values of being bold, accountable, and humble. We experience the value that our vision brings to the world every day, helping the teams behind the greatest innovations of our generation, from space travel to artificial intelligence, and more.
With liquidity solutions, exclusive data and insights, a custody offering, and a vibrant marketplace, Forge's goal is to build the best-in-class technology infrastructure to power a global private market that is transparent, accessible, and seamless for companies, their employees, and investors. Through Forge, employees can sell their private shares, employers can reward shareholders with pre-IPO liquidity and individual and institutional investors can participate in private unicorn growth.
Forge's differentiated global marketplace addresses rising demand among individual and institutional investors for exposure to private company stocks and is building a growing network effect.
Our ability to offer these powerful financial solutions has generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies, teams, and innovators in this way.
The Role:
- Manage Forge's Site Reliability Engineering team responsible for keeping Forge systems highly available for customers.
- Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning.
- Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics.
- Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation.
- Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards.
- Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
- Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability.
- Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
- 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
- 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
- Bachelor's degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
- Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
- Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support.
- Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
- Strong technical judgment, communication skills, and ability to influence across engineering and non-engineering stakeholders.
- Experience in FinTech, financial services, or another regulated industry.
- Experience with AWS and/or Azure cloud platforms.
- Familiarity with Kubernetes, container platforms, infrastructure-as-code, Terraform, Ansible, or similar automation tooling.
- Experience with observability platforms such as Datadog, CloudWatch, or similar tools.
- Experience improving developer experience through paved-road platforms, standardization, and self-service infrastructure capabilities.
- Experience supporting growth-stage companies where speed, scale, reliability, and operational discipline must be balanced.
Forge is proud to be an equal opportunity employer committed to supporting a diverse and inclusive workplace. Our employment decisions are made without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender, gender identity, gender expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, sexual orientation, veteran status, or any other characteristic protected by federal, state, or local laws.
- ...Google Cloud in San Francisco, CA seeks a Manager, Software Engineer for Site Reliability Engineering to lead a team responsible for uptime, availability, and reliability at scale. This role blends hands-on software engineering with people leadership and strategic roadmapping...Suggested
$165k - $225.6k
...infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering, this role will help build, improve, and...SuggestedPermanent employmentLocal areaWorldwideFlexible hours- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ..., GPU utilization, concurrency, and model lifecycle management. Define and instrument SLOs and SLIs across customer workloads...SuggestedFlexible hours
- ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with... ...senior engineering candidate to join the Site Reliability organization in San Francisco. Working... ...operational efficiency. Incident Management: Lead the coordinated response to incidents...SuggestedWorldwideWeekend work
$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense.... ...expects. You will own both the production reliability of our user-facing products and the... ...that matters. Infrastructure as Code: Manage cloud infrastructure through code. Build...Suggested- ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure... ...) ~ Strong debugging, problem-solving, and incident-management skills Preferred Experience with...
- ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally...Contract workLocal area
$230k - $310k
...daily users while enabling our engineering teams to ship fast. You'll... ...automation and tooling that improves reliability and partnering with... ...that scale with the product Manage and optimize our compute, networking... ...'ll bring ~5+ years in site reliability engineering,...Full timeWork at officeWork from home- ...culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of... ...6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-...Immediate startRemote workWorldwide
$170k - $220k
...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes.... ...of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also...Work at officeRemote workFlexible hours- ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging...WorldwideShift work
- ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability... ..., deployment automation, rollback mechanisms, and config management Implement and maintain monitoring, alerting, and...
$117k - $209.33k
...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure... ...such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. Success...Full timeFor contractors- ...products that empower people across the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and...Temporary workWorldwide
- ...Superhuman Docs's collaborative workspaces, Mail's inbox management, and Go, the proactive AI assistant that... ...responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth...WorldwideHome officeFlexible hours
- ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you... ...Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production environments at scale. Data Proficiency:...
- ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the... ...infrastructure for our users that scales, is reliable, and makes the complexities of operating... ...of the challenges: streaming, token management, rate limits, model-specific quirks....Permanent employmentShift work
- ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI... ...Systems Builder — Close the Loop Build and maintain fleet management systems: OTA update pipelines, device health tracking,...Remote work
$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations... ...Key Responsibilities Capacity Ingestion and Management: Takes proactive steps to design and architect infrastructure...Temporary workImmediate startFlexible hoursShift work$195k - $257.5k
...Staff Site Reliability Engineer Circle (NYSE: CRCL) is one of the world's leading internet financial platform companies, building the foundation... ...Operate and scale production blockchain infrastructure, managing full nodes across networks such as Arc, Ethereum, Solana,...Flexible hours$194k - $267k
...automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$200k - $260k
...Infrastructure Team as a technical leader driving reliability, automation, and scalability across the... ...practices across teams, mentor senior engineers, and be a primary escalation point for... ...engineers, without needing formal management authority to do it ~ Strong bias for...Casual workWork at officeRemote workFlexible hours- ...for our future. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability,... ...continuity plans Infrastructure & Automation * Architect and manage cloud infrastructure on AWS using Infrastructure as Code (...Full timePart time
$181k - $263k
...line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability... ...expertise: internals, autoscaling, multi-tenant workload management, and rightsizing ~ Advanced experience with real-time and...Full timeWork from homeWorldwideFlexible hoursNight shift- ...Job: Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary expert responsible for our entire compute ecosystem. Your key responsibilities will include: As a Staff SRE, you...
$120.6k - $150.9k
...Staff Site Reliability Engineer (SRE) We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our... ...initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization...Flexible hours$221.2k - $300k
...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link corporate_fare Google place San Francisco, CA, USA Advanced Experience owning outcomes and decision making, solving ambiguous...Full timeWork at office- ...Capital One, and CERN. We also run Akka Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s... ...and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions...Remote workFlexible hours
$61k - $101k
...Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong... ...: cloud or SaaS experience. Preferred: memory management and dump analysis experience, ideally Java heap dump analysis...Full time- ...world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,... ...can be provided) Cloud/SaaS experience Memory management and dump analysis (Java heap dump analysis preferred) ITSM...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Site Reliability Engineer. Be the first to apply!
- extraction manager San Francisco, CA
- certification manager San Francisco, CA
- senior manager tax San Francisco, CA
- ranch manager San Francisco, CA
- valuation manager San Francisco, CA
- lean manager San Francisco, CA
- employment manager San Francisco, CA
- senior preconstruction manager San Francisco, CA
- e-learning manager San Francisco, CA
- refrigeration manager San Francisco, CA



