Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineer

$150k - $220k
Full-time

Forge Global

At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values of being bold, accountable, and humble. We experience the value that our vision brings to the world every day, helping the teams behind the greatest innovations of our generation, from space travel to artificial intelligence, and more.

With liquidity solutions, exclusive data and insights, a custody offering, and a vibrant marketplace, Forge’s goal is to build the best-in-class technology infrastructure to power a global private market that is transparent, accessible, and seamless for companies, their employees, and investors. Through Forge, employees can sell their private shares, employers can reward shareholders with pre-IPO liquidity and individual and institutional investors can participate in private unicorn growth.

Forge's differentiated global marketplace addresses rising demand among individual and institutional investors for exposure to private company stocks and is building a growing network effect.

Our ability to offer these powerful financial solutions has generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies, teams, and innovators in this way.

The Role:

As an engineering organization, we pride ourselves on engineering as a creative activity. Engineering managers enable engineers to do their best work by maintaining a culture and environment where engineers can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team responsible for keeping Forge systems highly available for customers, while partnering closely with Platform, Engineering, Security, Compliance, and Product teams to improve reliability, observability, incident response, and operational maturity. This is an opportunity for a hands-on technical leader who can coach engineers, improve production operations, and help Forge build and run secure, scalable, and highly reliable products.

Responsibilities:

  • Manage Forge’s Site Reliability Engineering team responsible for keeping Forge systems highly available for customers.

  • Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning.

  • Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics.
  • Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation.
  • Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards.
  • Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
  • Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability.

    • Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
    • Partner with Security, Compliance, and Risk partners to ensure reliability and infrastructure practices meet the needs of a regulated business.

Qualifications:

  • 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
  • 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
  • Bachelor’s degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
  • Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
  • Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support.
  • Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
  • Strong technical judgment, communication skills, and ability to influence across engineering and non-engineering stakeholders.

Preferred Qualifications:

  • Experience in FinTech, financial services, or another regulated industry.
  • Experience with AWS and/or Azure cloud platforms.
  • Familiarity with Kubernetes, container platforms, infrastructure-as-code, Terraform, Ansible, or similar automation tooling.
  • Experience with observability platforms such as Datadog, CloudWatch, or similar tools.
  • Experience improving developer experience through paved-road platforms, standardization, and self-service infrastructure capabilities.
  • Experience supporting growth-stage companies where speed, scale, reliability, and operational discipline must be balanced.

For residents of San Francisco/Bay Area, CA or New York, NY the annual salary range for this role is $150,000-$220,000 + annual bonus. Final offers may vary from the amount listed based on geography, candidate experience and expertise, annual bonus, and other factors.

Upon offer, we conduct background checks that include employment and education verification, state, and county criminal history searches as well as fingerprint and drug test.

Forge is proud to be an equal opportunity employer committed to supporting a diverse and inclusive workplace. Our employment decisions are made without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender, gender identity, gender expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, sexual orientation, veteran status, or any other characteristic protected by federal, state, or local laws.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineer in San Francisco, CA vacancy
  • $150k - $220k

     ...in this way. The Role: As an engineering organization, we pride ourselves on engineering...  ...as a creative activity. Engineering managers enable engineers to do their best work...  ..., mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team... 
    Suggested
    Local area

    Forge Global

    San Francisco, CA
    3 hours ago
  •  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll...  ...habits and tooling.Architect and manage the SLO and error-budget framework, empowering... 
    Suggested
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    8 hours ago
  • $165k - $225.6k

     ...infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build, improve, and... 
    Suggested
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common...  ..., observability, access controls, auditability, and cost management. You will troubleshoot production issues, document operational... 
    Suggested
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $147k - $227k

     ...mission. If you are too, let's talk. The Engineering Opportunity We are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group...  ..., TLS, service networking, and traffic management. Experience with observability platforms... 
    Suggested
    Full time
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  •  ...fully integrated solutions to manage everything from business...  ...what’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...together to build scalable, reliable, and secure products that empower...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    1 day ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range...  ...observability and alerting systems.The Fleet Management team provides the core runtime...  ...critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    3 days ago
  •  ...Superhuman Docs's collaborative workspaces, Mail’s inbox management, and Go, the proactive AI assistant that...  ...responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    1 day ago
  • $139.76k - $287.75k

     ...business.We are seeking a Senior Site ReliabilityEngineer to help...  ...in advancing the reliability, scalability, automation, observability...  ...candidate is a highly hands-on engineer with strong production...  ...infrastructure provisioning and change management through Terraform/... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    8 hours ago
  • $167.7k - $245.2k

     ...performance, efficiency, change management, monitoring, emergency...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    8 hours ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. Success... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    2 days ago
  • $113.4k - $162k

     ...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...GitHub, Terraform, Ansible, or similar tools to build and manage cloud infrastructure efficiently.Incident Management Expert... 
    Temporary work

    TextNow

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate...  ...experience, including authoring reusable modules, managing state and environments, and delivering infrastructure changes... 
    Flexible hours

    Circle

    San Francisco, CA
    8 hours ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and...  ...cross-functionallyNice-to-HaveExperience with cloud and managed services (e.g. AWS)Experience supporting data-intensive platforms... 

    Alembic

    San Francisco, CA
    1 day ago
  • $148.5k - $223.9k

     ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with...  ...eliminate toil and improve operational efficiency.Incident Management: Lead the coordinated response to incidents as an... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    8 hours ago
  • $220k - $235k

     ...Magic Quadrant for Contract Lifecycle Management, a Fortune Great Place to Work, and one...  ...of our cloud platform and champion engineering excellence across Ironclad. In this role...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    8 hours ago
  • $204k - $306k

     ...We're all in on this mission. If you are too, let's talk.Manager, Site Reliability EngineeringSan Francisco, CaliforniaSecure Every Identity,...  ...week in our San Francisco Office. The IDaaS Site Reliability Engineering GroupOkta authenticates, authorizes and provisions... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Okta

    San Francisco, CA
    1 day ago
  • $217k - $303.9k

     ...Reddit grow its business. The reliability of our Ads systems directly...  ...partners closely with Ads Engineering teams to improve reliability...  ...ecosystem.We're looking for a Staff Site Reliability Engineer who...  ...SLOs, automation, incident management, and performance... 
    For contractors
    Work experience placement
    Remote work
    Flexible hours

    Reddit

    San Francisco, CA
    3 hours ago
  • $194k - $267k

     ..., automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • $195k - $257.5k

     ...is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and...  ...on:Operate and scale production blockchain infrastructure, managing full nodes across networks such as Arc, Ethereum, Solana,... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  • $194k - $267k

     ...let's talk.The TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission...  ...balancing, ingress, TLS, service networking, and traffic management.Strategic experience designing comprehensive observability... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $146.7k - $234.3k

     ...team for our future. We are seeking an experienced Lead Site Reliability Engineer to join our engineering team and drive the reliability, scalability...  ...plans Infrastructure & Automation • Architect and manage cloud infrastructure on AWS using Infrastructure as Code (... 
    Full time
    Temporary work
    Part time
    Shift work

    Federal Reserve System

    San Francisco, CA
    2 days ago
  • $174k - $239k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal, Inc. is looking for...  ...networking concepts, such as BGP and IPsec management, and has leveraged AWS networking services,... 
    Work experience placement
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...Overview:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...  ....Required Skills & Experience (The Essentials)Log Management: Minimum 5+ Experience scaling and managing Splunk Cloud at... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    8 hours ago
  •  ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ..., GPU utilization, concurrency, and model lifecycle management. Define and instrument SLOs and SLIs across customer workloads... 
    Flexible hours

    Baseten

    San Francisco, CA
    5 days ago
  • The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...) ~ Strong debugging, problem-solving, and incident-management skills Preferred Experience with infrastructure-as... 

    Blaxel, Inc

    San Francisco, CA
    3 days ago
  • $260k - $300k

     ...makers of Devin, the first AI software engineer. Our team is extremely talent-dense....  ...expects. You will own both the production reliability of our user-facing products and the...  ...that matters. Infrastructure as Code: Manage cloud infrastructure through code. Build... 

    Cognition AI

    San Francisco, CA
    5 days ago
  •  ...DESCRIPTION Project Outline: We are looking for a Site Reliability Engineer with experience in incident response. In this role, you...  ...Background: 4+ years in SRE, DevOps, or Systems Engineering roles managing production environments at scale. - Data Proficiency:... 

    BayOne Solutions

    San Francisco, CA
    2 days ago
  •  ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on....  ...workloads.Own network device configuration management end to end, ensuring consistency and... 

    Alembic

    San Francisco, CA
    3 days ago
  • $167.7k - $245.2k

     ...Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineer. Be the first to apply!