Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineer

$150k - $220k

Forge Global

At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values of being bold, accountable, and humble. We experience the value that our vision brings to the world every day, helping the teams behind the greatest innovations of our generation, from space travel to artificial intelligence, and more.


With liquidity solutions, exclusive data and insights, a custody offering, and a vibrant marketplace, Forge's goal is to build the best-in-class technology infrastructure to power a global private market that is transparent, accessible, and seamless for companies, their employees, and investors. Through Forge, employees can sell their private shares, employers can reward shareholders with pre-IPO liquidity and individual and institutional investors can participate in private unicorn growth.


Forge's differentiated global marketplace addresses rising demand among individual and institutional investors for exposure to private company stocks and is building a growing network effect.


Our ability to offer these powerful financial solutions has generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies, teams, and innovators in this way.


The Role:

As an engineering organization, we pride ourselves on engineering as a creative activity. Engineering managers enable engineers to do their best work by maintaining a culture and environment where engineers can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge's SRE team responsible for keeping Forge systems highly available for customers, while partnering closely with Platform, Engineering, Security, Compliance, and Product teams to improve reliability, observability, incident response, and operational maturity. This is an opportunity for a hands-on technical leader who can coach engineers, improve production operations, and help Forge build and run secure, scalable, and highly reliable products.

Responsibilities:
  • Manage Forge's Site Reliability Engineering team responsible for keeping Forge systems highly available for customers.
  • Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning.
  • Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics.
  • Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation.
  • Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards.
  • Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
  • Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability.
    • Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
    • Partner with Security, Compliance, and Risk partners to ensure reliability and infrastructure practices meet the needs of a regulated business.
Qualifications:
  • 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
  • 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
  • Bachelor's degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
  • Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
  • Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support.
  • Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
  • Strong technical judgment, communication skills, and ability to influence across engineering and non-engineering stakeholders.
Preferred Qualifications:
  • Experience in FinTech, financial services, or another regulated industry.
  • Experience with AWS and/or Azure cloud platforms.
  • Familiarity with Kubernetes, container platforms, infrastructure-as-code, Terraform, Ansible, or similar automation tooling.
  • Experience with observability platforms such as Datadog, CloudWatch, or similar tools.
  • Experience improving developer experience through paved-road platforms, standardization, and self-service infrastructure capabilities.
  • Experience supporting growth-stage companies where speed, scale, reliability, and operational discipline must be balanced.

For residents of San Francisco/Bay Area, CA or New York, NY the annual salary range for this role is $150,000-$220,000 + annual bonus. Final offers may vary from the amount listed based on geography, candidate experience and expertise, annual bonus, and other factors.

Upon offer, we conduct background checks that include employment and education verification, state, and county criminal history searches as well as fingerprint and drug test.


Forge is proud to be an equal opportunity employer committed to supporting a diverse and inclusive workplace. Our employment decisions are made without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender, gender identity, gender expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, sexual orientation, veteran status, or any other characteristic protected by federal, state, or local laws.
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineer in San Francisco, CA vacancy
  •  ...Google Cloud in San Francisco, CA seeks a Manager, Software Engineer for Site Reliability Engineering to lead a team responsible for uptime, availability, and reliability at scale. This role blends hands-on software engineering with people leadership and strategic roadmapping... 
    Suggested

    Jobleads-US

    San Francisco, CA
    5 days ago
  • $165k - $225.6k

     ...infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology. The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering, this role will help build, improve, and... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    5 days ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ..., GPU utilization, concurrency, and model lifecycle management. Define and instrument SLOs and SLIs across customer workloads... 
    Suggested
    Flexible hours

    Baseten

    San Francisco, CA
    4 days ago
  •  ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with...  ...senior engineering candidate to join the Site Reliability organization in San Francisco. Working...  ...operational efficiency. Incident Management: Lead the coordinated response to incidents... 
    Suggested
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer. Our team is extremely talent-dense....  ...expects. You will own both the production reliability of our user-facing products and the...  ...that matters. Infrastructure as Code: Manage cloud infrastructure through code. Build... 
    Suggested

    Cognition AI

    San Francisco, CA
    1 day ago
  •  ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...) ~ Strong debugging, problem-solving, and incident-management skills Preferred Experience with... 

    Blaxel, Inc

    San Francisco, CA
    2 days ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally... 
    Contract work
    Local area

    InterSources

    San Francisco, CA
    3 days ago
  • $230k - $310k

     ...daily users while enabling our engineering teams to ship fast. You'll...  ...automation and tooling that improves reliability and partnering with...  ...that scale with the product Manage and optimize our compute, networking...  ...'ll bring ~5+ years in site reliability engineering,... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    3 days ago
  •  ...culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of...  ...6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    17 hours ago
  • $170k - $220k

     ...Senior Site Reliability Engineer Supio is a trusted AI platform purpose-built for law firms, reshaping how data drives impactful outcomes....  ...of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also... 
    Work at office
    Remote work
    Flexible hours

    Supio

    San Francisco, CA
    5 days ago
  •  ...enterprise that runs the real economy. Learn more about our vision in our manifesto. About the Role We're looking for a Site Reliability Engineer to take the lead on scaling our operational resilience as we grow. You'll own the stability, observability, and debugging... 
    Worldwide
    Shift work

    Happy Robot

    San Francisco, CA
    1 day ago
  •  ...Site Reliability Engineer (SRE) We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability...  ..., deployment automation, rollback mechanisms, and config management Implement and maintain monitoring, alerting, and... 

    Alembic Technologies

    San Francisco, CA
    4 days ago
  • $117k - $209.33k

     ...Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure...  ...such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. Success... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    5 days ago
  •  ...products that empower people across the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability, and... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    1 day ago
  •  ...Superhuman Docs's collaborative workspaces, Mail's inbox management, and Go, the proactive AI assistant that...  ...responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    3 days ago
  •  ...JOB DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you...  ...Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production environments at scale. Data Proficiency:... 

    BayOne Solutions

    San Francisco, CA
    1 day ago
  •  ...Arena Intelligence Engineer Arena Intelligence is looking for an engineer to build the...  ...infrastructure for our users that scales, is reliable, and makes the complexities of operating...  ...of the challenges: streaming, token management, rate limits, model-specific quirks.... 
    Permanent employment
    Shift work

    Arena AI

    San Francisco, CA
    2 days ago
  •  ...Site Reliability Engineer Specter's mission is to help automate the physical world. Today, we build video sensors with state-of-the-art AI...  ...Systems Builder — Close the Loop Build and maintain fleet management systems: OTA update pipelines, device health tracking,... 
    Remote work

    Specter Services LLC

    San Francisco, CA
    18 hours ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations...  ...Key Responsibilities Capacity Ingestion and Management: Takes proactive steps to design and architect infrastructure... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    5 days ago
  • $195k - $257.5k

     ...Staff Site Reliability Engineer Circle (NYSE: CRCL) is one of the world's leading internet financial platform companies, building the foundation...  ...Operate and scale production blockchain infrastructure, managing full nodes across networks such as Arc, Ethereum, Solana,... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $194k - $267k

     ...automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    16 hours ago
  • $200k - $260k

     ...Infrastructure Team as a technical leader driving reliability, automation, and scalability across the...  ...practices across teams, mentor senior engineers, and be a primary escalation point for...  ...engineers, without needing formal management authority to do it ~ Strong bias for... 
    Casual work
    Work at office
    Remote work
    Flexible hours

    Sight Machine

    San Francisco, CA
    2 days ago
  •  ...for our future. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability,...  ...continuity plans Infrastructure & Automation * Architect and manage cloud infrastructure on AWS using Infrastructure as Code (... 
    Full time
    Part time

    Federal Reserve Bank of San Francisco

    San Francisco, CA
    3 days ago
  • $181k - $263k

     ...line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability...  ...expertise: internals, autoscaling, multi-tenant workload management, and rightsizing ~ Advanced experience with real-time and... 
    Full time
    Work from home
    Worldwide
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    1 day ago
  •  ...Job: Staff Site Reliability Engineer (SRE) Location: San Francisco, CA Job Responsibilities As our Staff SRE, you'll be the primary expert responsible for our entire compute ecosystem. Your key responsibilities will include: As a Staff SRE, you... 

    United IT Solutions

    San Francisco, CA
    3 days ago
  • $120.6k - $150.9k

     ...Staff Site Reliability Engineer (SRE) We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our...  ...initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization... 
    Flexible hours

    WEX

    San Francisco, CA
    2 days ago
  • $221.2k - $300k

     ...Manager, Software Engineer, Site Reliability Engineering Share Manager, Software Engineer, Site Reliability Engineering ~ link Copy link corporate_fare Google place San Francisco, CA, USA Advanced Experience owning outcomes and decision making, solving ambiguous... 
    Full time
    Work at office

    Google Inc.

    San Francisco, CA
    1 hour ago
  •  ...Capital One, and CERN. We also run Akka Automated Operations, our managed platform for customer workloads on dedicated, BYOC, and BYOK8s...  ...and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture decisions... 
    Remote work
    Flexible hours

    Akka

    San Francisco, CA
    12 days ago
  • $61k - $101k

     ...Requirements: We expect formal training or certification in site reliability engineering, plus 3+ years of hands-on experience. We want strong...  ...: cloud or SaaS experience. Preferred: memory management and dump analysis experience, ideally Java heap dump analysis... 
    Full time

    J.P. Morgan

    San Francisco, CA
    4 days ago
  •  ...world's most complex and mission-critical systems.   As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,...  ...can be provided) Cloud/SaaS experience Memory management and dump analysis (Java heap dump analysis preferred) ITSM... 

    JPMorgan Chase & Co.

    San Francisco, CA
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineer. Be the first to apply!