Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Manager, Site Reliability Engineering

Form Energy, Inc.

Are you ready to build America's energy future? Form Energy is an American manufacturing and energy technology company. We're revolutionizing energy storage with cost-effective, multi-day technology designed to keep the electric grid secure and reliable, even during extended periods of stress. By strengthening the electric system and reimagining what's possible, we're giving clean energy a whole new form!

In recent years, Form Energy has earned a number of accolades, including being named by TIME as a "Best Invention", MIT Technology Review as a "Top Climate Tech Company To Watch", and Fast Company as "One of the Next Big Things In Tech". We are making rapid progress on our mission of delivering energy storage for a better world, and our team is growing just as rapidly to meet demand. We have signed contracts with leading electric utilities across the United States and production of our iron-air batteries is underway at our first high-volume manufacturing facility in West Virginia.

Working for Form Energy is more than just a job, it's a chance to be part of something extraordinary. And now - right as we significantly scale up battery manufacturing - might be the most exciting moment in the company's history to join. We are assembling a team of highly talented and driven individuals across the country. Driven by our core values of humanity, excellence, and creativity, our team is determined to deliver on our mission and transform the energy landscape for the better.

Feeling energized to make a meaningful impact on the world? Then keep reading - you've come to the right place.

Role Description

Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function responsible for maintaining the reliability, availability, and supportability of our deployed energy storage systems. This role will own the mechanisms used to monitor fleet health, respond to incidents, triage field issues, coordinate engineering escalation, and ensure new products and releases are operationally ready. As Form Energy moves from early deployments to a growing commercial fleet, a key objective of this role is to scale operational capacity through automation, tooling, and product design improvements. The successful candidate will work across software, firmware, controls, hardware, field service, and customer-facing teams to build the scalable processes and capabilities required to enable reliable fleet operations.

Relocation assistance is available.

What you'll do:
  • Lead Product Operations Assurance with a DevOps first principals approach to fleet monitoring, incident response, field triage, engineering escalation, and production support.

  • Build a highly automated operating model that enables the team to support a rapidly growing deployed fleet, managing operational work through automation and driving product design improvements that reduce sustaining engineering overhead.

  • Establish incident management processes, including severity definitions, escalation paths, incident command, communications, and post-incident reviews.

  • Coordinate cross-functional engineering response to complex issues spanning software, firmware, controls, networking, hardware, and site infrastructure.

  • Develop diagnostic playbooks, troubleshooting procedures, and operational tooling that improve first-line response and reduce dependence on individual experts.

  • Partner with Data, Analytics, and Cloud Applications teams to define the telemetry, dashboards, alerts, and workflows required to effectively monitor and support the fleet.

  • Define operational readiness requirements for new product releases and deployments, including monitoring, diagnostics, recovery procedures, escalation paths, and support documentation.

  • Analyze incidents and fleet data to identify recurring failure modes and drive reliability, diagnosability, and serviceability improvements back into the product.

  • Establish and track operational metrics such as fleet availability, incident frequency, time to detection, time to containment, time to recovery, and recurrence.

  • Build the team, processes, on-call model, and automation required to scale fleet operations as the installed base grows.

What you'll bring:
  • 12+ years of experience supporting complex production, industrial, energy, infrastructure, automotive, robotics, or other cyber-physical systems, including technical leadership or people management experience.

  • Demonstrated experience leading production incident response, technical troubleshooting, escalation management, and root-cause investigation in deployed systems.

  • Experience leading SRE, DevOps, or production operations teams in highly automated environments, with a track record of scaling operational capacity through software, tooling, and product improvements.

  • Strong ability to diagnose and coordinate resolution of problems spanning software, firmware, controls, networking, hardware, and field operations.

  • Experience with operational monitoring, observability, alerting, remote diagnostics, reliability practices, and the operational processes needed to support deployed products.

  • Proven ability to lead cross-functional teams through high-priority technical issues, communicate risk and recovery plans clearly, and translate operational learnings into improvements in product reliability, diagnosability, and serviceability.

  • Bachelor's degree in engineering, computer science, or a related technical discipline, or equivalent practical experience.

#LI-TR1

Humanity is a cornerstone of Form Energy's culture, and we make sure our compensation and benefits reflect that. Form Energy offers competitive salaries, stock options, and a holistic benefits package to ensure all employees have what they need to thrive while working here.

When it comes to you and your family's health, we cover 100% of medical, dental, and vision premiums for full-time employees - and 80% of healthcare premiums for dependents. This starts from day one. We also offer at least 12 weeks of paid leave for new parents (up to 20 weeks for birthing parents), and generous vacation policies to give employees time to recharge when needed.

To build America's energy future, we need everyone at the table. We are proud to be an equal opportunity employer, and encourage candidates from all backgrounds to apply to our open jobs.

If you may require reasonable accommodations to participate in our interview process, please contact View email address on click.appcast.io. Requests for accommodations will be treated with discretion.
Form Energy is committed to maintaining the privacy of our applicants. Please be aware that we will never solicit sensitive personal information such as Social Security numbers or bank account details during the recruiting or hiring process.

#J-18808-Ljbffr
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Manager, Site Reliability Engineering in Berkeley, CA vacancy
  • $198.05k - $267.95k

     ...Lead Site Reliability Engineer Company The Boeing Company The Boeing Company is looking for a Lead Site Reliability Engineer to...  ...automation, Infrastructure as Code, Ansible, configuration management, and repeatable operational patterns that reduce toil and... 
    Suggested
    Permanent employment
    Work experience placement
    Relocation
    Visa sponsorship
    Work visa
    Relocation package
    Flexible hours
    Shift work

    Boeing

    Berkeley, CA
    1 day ago
  •  ...attract incredibly creative scientists and engineers from leading academic institutions and...  ...Summary We are looking for a Site Reliability Engineer to own the digital...  ...resources based on demand. Access Management: Ensure the right people have access to... 
    Suggested
    Visa sponsorship

    Astera

    Emeryville, CA
    1 day ago
  • $174.92k - $209.91k

     ...access to data as simple and reliable as electricity. With Fivetran...  ...and ready to query, with no engineering or maintenance required. We’re...  ...our teams, systems, and career sites. About the Role...  ...with engineering teams, product managers, as well as support and sales... 
    Suggested
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    1 day ago
  •  ...Site Reliability Engineer The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site...  ...advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work... 
    Suggested
    Work at office
    Night shift

    Bay Systems Consulting Inc

    Berkeley, CA
    12 hours ago
  •  ...Site Reliability Professional We are looking for a highly motivated site reliability professional...  ...by collaborating with multiple engineering teams. Analytical mindset and good...  ...Experience in Docker orchestration and management. ~ Experience with Kubernetes.... 
    Suggested
    Work experience placement

    My3Tech Inc

    Oakland, CA
    1 day ago
  • $80 per hour

     ...possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems...  ..., Prometheus/VictoriaMetrics, Alertmanager, and building management/cooling systems Network security fundamentals (ACLs, firewalls... 
    Contract work
    Shift work

    LTD Global, LLC

    Berkeley, CA
    12 hours ago
  • $80 per hour

     ...Overview Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific...  ...increase. Utilize ServiceNow to support incident management, trouble-ticketing, operational workflows, and service management... 
    Hourly pay
    Full time
    Work at office
    Local area
    Shift work
    Night shift

    Essnova Solutions, Inc.

    Berkeley, CA
    14 days ago
  • $150k - $220k

     ...in this way. The Role: As an engineering organization, we pride ourselves on engineering...  ...as a creative activity. Engineering managers enable engineers to do their best work...  ..., mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge’s SRE team... 
    Local area

    Forge Global

    San Francisco, CA
    12 hours ago
  •  ...the world's most complex and mission-critical systems. As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology,...  ...exposure (training can be provided) Cloud/SaaS experience Memory management and dump analysis (Java heap dump analysis preferred) ITSM... 

    J.P. Morgan

    San Francisco, CA
    4 days ago
  • $194k - $237k

     ...employment Visa sponsorship. Role Summary The Principal Site Reliability Engineer applies software engineering and systems engineering...  ...as Code, automation, testing, incident response, capacity management, resilience, and operational readiness. Identify recurring... 
    Hourly pay
    Work at office
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning Services

    San Francisco, CA
    23 hours ago
  •  ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves... 

    Breakout Tools

    San Francisco, CA
    5 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals...  ...services. You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role includes... 

    Socket

    San Francisco, CA
    2 days ago
  • $200.7k - $250.9k

     ...washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in...  ...opportunities for improvements in reliability/observability/performance/preparedness and...  ...candidate for the role: Has past Site Reliability Engineering or DevOps experience... 

    Embedded Shishya

    San Francisco, CA
    1 day ago
  • $120k - $168.49k

     ...Site Reliability Engineer, Cloud Infrastructure About Quizlet At Quizlet, our mission is to help every learner achieve their outcomes in the...  ...automation scripts and tools for deployments, infrastructure management, and operational tasks to reduce manual effort (toil).... 
    Internship
    Work at office
    3 days per week

    Quizlet

    San Francisco, CA
    5 days ago
  •  ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage... 

    Lambda

    San Francisco, CA
    1 day ago
  • $200k - $240k

     ...Senior Site Reliability Engineer (SRE)Location: San Francisco, CAWork Model: OnsiteIndustry: Renewable EnergyComp: $200,000 - $240,000 We’re...  ...infrastructure across AWS and physical environments. Build, manage, and automate infrastructure using Terraform and... 

    Lawrence Harvey Search & Selection

    San Francisco, CA
    1 day ago
  •  ...troubleshooting, providing rubric-based written feedback. This role requires hands-on Kubernetes expertise in EKS/GKE/AKS or self-managed clusters, with strong scripting in Go, Python, or TypeScript, and ability to document findings clearly for engineering #J-18808-Ljbffr

    Obsidian

    San Francisco, CA
    3 days ago
  • $350k

     ...and infrastructure providers to build reliable, high-performance platforms...  .... This opportunity is for a Staff Site Reliability Engineer to lead the reliability of large-scale...  ...scale GPU fleets, including lifecycle management, validation, firmware rollouts, upgrades... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    23 hours ago
  • $200k - $240k

     ...across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to solve challenging...  ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    12 hours ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based...  ...implementing improvements to prevent recurrence. Create and manage multiple cloud instances (dev, staging, test), optimize... 
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    4 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round...  ...scalable applications in a fast-paced environment. Manage and optimize Kubernetes clusters for high availability and... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  •  ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with...  ...senior engineering candidate to join the Site Reliability organization in San Francisco. Working...  ...operational efficiency. Incident Management: Lead the coordinated response to incidents... 
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    4 days ago
  •  ...small team of former Google and Stripe engineers, including the founding team of Google...  ...looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you...  ...systems Define and track SLIs/SLOs, manage error budgets, and proactively monitor... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    4 days ago
  •  ...A leading technology firm is looking for a Manager to expand their Cloud Site Reliability team. The ideal candidate will have extensive Linux administration experience, a passion for automation, and be comfortable in a remote, diverse workplace. This position emphasizes... 
    Remote work

    mbhsobana

    San Francisco, CA
    5 days ago
  •  ...products that empower people across the globe. Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability, availability,... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    1 day ago
  •  ...unplanned downtime and run better operations, across 13.9 million managed assets and 79.5 million completed work orders. In August 2...  ...actually happened, and it is ours. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability,... 
    Remote work

    MaintainX

    San Francisco, CA
    3 days ago
  •  ...treatment. What We Look for in a Great Engineer You have the intensity and...  ...deep expertise in Kubernetes and Helm to manage deployment, scaling, and operational health...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer... 
    Work at office

    LATENT

    San Francisco, CA
    46 minutes ago
  •  ...The role We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure...  ...) ~ Strong debugging, problem-solving, and incident-management skills Preferred Experience with... 

    Blaxel, Inc

    San Francisco, CA
    2 days ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations...  ...Key Responsibilities Capacity Ingestion and Management: Takes proactive steps to design and architect infrastructure... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    12 hours ago
  •  ...scripting languages such as Python and Bash Strong experience with infrastructure as code, specifically Terraform or Azure Resource Manager Templates, CloudFormation Strong experience with Azure, GCP or AWS services You have deployed a production application to... 
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Manager, Site Reliability Engineering. Be the first to apply!