Manager, Site Reliability Engineer
$150k - $220kForge USA
Manager, Site Reliability Engineer
At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values of being bold, accountable, and humble. We experience the value that our vision brings to the world every day, helping the teams behind the greatest innovations of our generation, from space travel to artificial intelligence, and more.
With liquidity solutions, exclusive data and insights, a custody offering, and a vibrant marketplace, Forge's goal is to build the best-in-class technology infrastructure to power a global private market that is transparent, accessible, and seamless for companies, their employees, and investors. Through Forge, employees can sell their private shares, employers can reward shareholders with pre-IPO liquidity and individual and institutional investors can participate in private unicorn growth.
Forge's differentiated global marketplace addresses rising demand among individual and institutional investors for exposure to private company stocks and is building a growing network effect.
Our ability to offer these powerful financial solutions has generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies, teams, and innovators in this way.
The Role:
As an engineering organization, we pride ourselves on engineering as a creative activity. Engineering managers enable engineers to do their best work by maintaining a culture and environment where engineers can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge's SRE team responsible for keeping Forge systems highly available for customers, while partnering closely with Platform, Engineering, Security, Compliance, and Product teams to improve reliability, observability, incident response, and operational maturity. This is an opportunity for a hands-on technical leader who can coach engineers, improve production operations, and help Forge build and run secure, scalable, and highly reliable products.
Responsibilities:
- Manage Forge's Site Reliability Engineering team responsible for keeping Forge systems highly available for customers.
- Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning.
- Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics.
- Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation.
- Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards.
- Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
- Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability.
- Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
- Partner with Security, Compliance, and Risk partners to ensure reliability and infrastructure practices meet the needs of a regulated business.
Qualifications:
- 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
- 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
- Bachelor's degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
- Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
- Hands-on experience with observability, monitoring, alerting, incident response, troubleshooting, and production support.
- Experience with CI/CD, infrastructure automation, cloud platforms, and operational tooling.
- Strong technical judgment, communication skills, and ability to influence across engineering and non-engineering stakeholders.
Preferred Qualifications:
- Experience in FinTech, financial services, or another regulated industry.
- Experience with AWS and/or Azure cloud platforms.
- Familiarity with Kubernetes, container platforms, infrastructure-as-code, Terraform, Ansible, or similar automation tooling.
- Experience with observability platforms such as Datadog, CloudWatch, or similar tools.
- Experience improving developer experience through paved-road platforms, standardization, and self-service infrastructure capabilities.
- Experience supporting growth-stage companies where speed, scale, reliability, and operational discipline must be balanced.
For residents of San Francisco/Bay Area, CA or New York, NY the annual salary range for this role is $150,000-$220,000 + annual bonus. Final offers may vary from the amount listed based on geography, candidate experience and expertise, annual bonus, and other factors.
Forge is proud to be an equal opportunity employer committed to supporting a diverse and inclusive workplace. Our employment decisions are made without regard to race, color, religion, sex (including pregnancy, childbirth, or related medical conditions), gender, gender identity, gender expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, sexual orientation, veteran status, or any other characteristic protected by federal, state, or local laws.
- ...EIT) organization is expanding, and we are seeking a Senior Site Reliability Engineer to help drive a major architectural modernization. In this... ...: • Everything as Code: Drive repository-led management across our public and private cloud environments to establish...SuggestedPermanent employmentFull timeH1bLocal areaRemote workShift work
$153k - $210k
...Senior Software Engineer, Site Reliability Engineering Reno, NV; San Ramon, CA; NYC - Hybrid Are you passionate about building resilient... ...(SLOs), and error budget practices to proactively manage reliability. Identify capacity constraints and reliability...SuggestedFull time$45 - $85 per hour
...Description The Site Reliability Engineering groups goal is to ensure Customers can always use the service reliably. We're looking for... ...Devops, Cloud, Aws, Terraform, Automation, incident management Top Skills Details Devops,Cloud,Aws,Terraform,Automation...SuggestedContract workTemporary work$100k - $250k
...financial markets. Role Roadmap As a member of Kalshi's engineering team, you'll help build the next-generation financial... ...own, and evolve. What You'll Do Improve observability, reliability, and service availability by defining and measuring key metrics...SuggestedLocal area- ...Deployment Responsible for reliability and support of Container... ...reliability issues, Incident management, problem management Identifying... ...blameless RCA, partner with engineering and operation teams across... ...Automation Process Engineer,Site Reliability Engineer,Full...Suggested
$130k - $200k
...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups and global enterprises, from bare metal up through the platform services teams actually build on. Our culture runs on ownership...Shift work- ...Site Reliability Engineer I, Abhishek, would like to share a job opportunity as Site Reliability Engineer in Jacksonville, FL, Cary, NC or New York, NY (Onsite) location for a Fulltime position. In case, if you are not comfortable with this location, please share your...Full timeWork visa
$104k - $178k
...Sr. Site Reliability Engineer I You will join the Site Reliability Engineering (SRE) team within DoubleVerify's Technology organization. The... ...Responding to incidents and driving them to resolution, managing Sev1/Sev2 situations. Reducing MTTR (mean time to resolution...- ...Site Reliability Engineer Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We... ...of high-performance computing (HPC) systems and workload managers (Slurm) • worked with modern AI-oriented solutions (Fluidstack...Relocation package
- ...Site Reliability Engineer (SRE) Job Title Site Reliability Engineer (SRE) Job Summary We are seeking a skilled Site... ...RCA), and implement preventive measures. Configure and manage observability solutions including logging, monitoring, tracing...Flexible hours
- ...Chariot Engineering Hire Chariot's engineering hire will be responsible for taking the Chariot platform to the next level. You will... ...world while working as part of a small, fast-moving team. Manage and build out the backend server applications and components that...Work experience placementWork at officeWork from homeMonday to FridayMonday to Thursday
- ...Senior Site Reliability Engineer (SRE) Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating real systems at scale — not just building infrastructure...Full timeWork at officeRemote workFlexible hours2 days per week
- ...Senior Site Reliability Engineer (SRE) Our client is seeking a Senior Site Reliability Engineer (SRE) with 10–15 years of experience to support... ...on troubleshooting complex trading infrastructure, managing observability, and collaborating across trading and technology...
- ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies....Local area
- ...SRE Engineer Location: New York, NY, USA Experience: 8-12 Years Client: Amex Job Description: This is an SRE role supporting... ..., who has great analytical skills and is a good incident manager as well. Skills: SRE REST Web Services Kafka Kubernetes...
$260k - $300k
...makers of Devin, the first AI software engineer. Our team is extremely talent-dense.... ...expects. You will own both the production reliability of our user-facing products and the... ...that matters. Infrastructure as Code: Manage cloud infrastructure through code. Build...$139k - $257.55k
...Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning,... ...them Own patch, vulnerability, and golden-image lifecycle management across the fleet - triage, remediate, automate Contribute...Temporary workLocal areaRemote workWorldwide- ...Site Reliability Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence... ..., GPU utilization, concurrency, and model lifecycle management. Define and instrument SLOs and SLIs across customer workloads...Flexible hours
$189k - $283.6k
...Afterpay is transforming the way customers manage their spending over time. TIDAL is a... ...proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...strong desire to perform and grow as an engineer ~5+ years of software development...Full timeLocal areaRemote workRelocation packageFlexible hoursShift work- ...Software Reliability Engineer Good software has to run where customers need it. For many of Retool's largest customers, that means running... ...team owns the systems that make this possible: Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud)...
$120k - $180k
...people, and works with high-profile manufacturers including leaders in space and defense. You will be the first dedicated Site Reliability Engineer and own critical infrastructure end to end. This is a greenfield opportunity to architect the path from AWS to on-premises...Permanent employmentFull timeRelocation package$182.8k - $247.3k
...to develop education for our half a billion (and growing!) learners around the world. About the role... As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems...Work experience placement$115k - $125k
...Site Reliability Engineer New York City, NY Pico fuels the global capital markets community by providing exceptional market data services and customized managed infrastructure solutions. As financial industry experts at the center of markets and technology, we help...Work experience placementWork at officeWork from homeMonday to FridayFlexible hoursShift workWeekend workAfternoon shiftEarly shift$105k - $300k
...Site Reliability Engineer At Citadel, a leading investor in the world's financial markets, we aim to win together as one team to earn the... ...solutions for issues based on root cause analyses Own incident management and drive tactical and strategic solutions Lead by...- ...Staff Site Reliability Engineer Tabs is the leading AI-native revenue platform for modern finance and accounting teams. Tabs agents automate the entire contract-to-cash lifecycle, including billing, collections, revenue recognition, and reporting, to help teams eliminate...Full timeContract workWork at office
- ...raising the bar. This is the place. The Role As a Staff Site Reliability Engineer you'll play a lead role on the founding SRE team at our... ...production readiness standards org-wide Lead incident management practices, escalation design, and drive systemic improvements...Work at office
$131k - $164k
...Staff Site Reliability Engineer New York, New York, United States Role Overview You're a seasoned Site Reliability Engineer who loves owning... ...that host enterprise and SaaS workloads. Integrate and manage Active Directory for authentication, access control, and...Work at officeLocal areaVisa sponsorshipFlexible hours$194k - $267k
...automate it" and who can rapidly self-educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$179k - $250k
...and fintechs trust Alloy for smarter risk management that drives growth. Through our... ...Infrastructure Team is a small team (6 engineers) responsible for a large and growing infrastructure... ...isn’t just scale—it’s making that scale reliable, secure, and operable with less manual...Work at officeLocal areaImmediate startWork from homeHome officeMonday to FridayFlexible hours- ...Principal Site Reliability Engineer Location: New York, NY (Onsite) Job type: Contract Job Description: Job Requirements Must Have:... ...optimization - GitLab CI/CD and HashiCorp Vault secrets management - Implementation of AIOps use cases and self-healing rules...Contract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Manager, Site Reliability Engineer. Be the first to apply!
- extraction manager New York, NY
- certification manager New York, NY
- senior manager tax New York, NY
- ranch manager New York, NY
- sterile manager New York, NY
- collision manager New York, NY
- valuation manager New York, NY
- lean manager New York, NY
- employment manager New York, NY
- senior preconstruction manager New York, NY


