Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Reliability Engineer

Bay Systems

Job Description

Job Description

Company: Lawrence Berkeley National Laboratory ( Through Bay Systems Consulting Inc. )

1 year full-time Contract.

Extension is based on company budget and performance

National Energy Research Scientific Computing Center (NERSC) - Site Reliability Engineer.

As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment.

DUTIES & RESPONSIBILITIES

  • Onsite 5-day weekly schedule consisting of Owl (midnight–8 am) shifts to monitor the NERSC HPC Facility.
  • Review and respond to alerts from computer systems, storage, network, and other data
  • center/facility-related systems by triaging or calling the appropriate on-call staff.
  • Create appropriate solutions to improve processes, prevent issue recurrence, and automate
  • responses to all routine service conditions.
  • Identify issues and propose solutions that will improve monitoring capabilities or provide better
  • automation for triage.
  • Possess expertise in ServiceNow and its usage to develop and implement customized service
  • management solutions.
  • Respond to alerts from multiple systems to ensure that data collection continues 24/7, providing
  • real-time information for diagnoses.
  • Develop and maintain tools within the monitoring pipeline in collaboration with the Operations
  • Team.
  • Create new software to provide alerts and notifications from HPC system APIs into the
  • monitoring pipeline.
  • Builds and maintains application/tool configurations to ensure software runs reliably as data and
  • user demands grow.
  • Collaborate with other groups at NERSC to ensure that communication and workflows are clearly
  • understood.
  • Work closely with other NERSC groups to coordinate center-wide maintenance activities and
  • manage diagnostic and notification software during maintenance periods.
  • Perform regular physical and logical walkthroughs of the data center floor to monitor
  • environmental health, power distribution units, and cooling infrastructure to ensure peak
  • operational efficiency.
  • Provide accurate information in the trouble ticketing system for outages, maintenance updates,
  • and other incidents so that workflows and protocols can be appropriately tracked by others.
  • Work on and resolve problems of diverse scope where data analysis requires the evaluation of
  • identifiable factors.
  • Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
  • Work on and resolve complex issues where the analysis of situations or data requires an in-depth
  • evaluation of variable factors.

REQUIREMENTS

  • Experience in or willingness to work within a 24/7 onsite team environment to support
  • large-scale data centers or critical installations.
  • Experience on Linux shell and working in a command-line (e.g. SSH) environment.
  • Experience with developing tools using various programming languages such as C, C++, Perl,
  • Java, or Python or a scripting language with knowledge of standard software development
  • practices.
  • Motivated, self-starter who can learn technologies that improve data center management in
  • areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management
  • software, evaporative
  • cooling, and power utilization.
  • Experience with network security: configuring/maintaining ACLs, knowledge of firewalls
  • Experience collaborating across technical teams to resolve operational bottlenecks and ensure
  • system reliability and alignment with service-level objectives.
  • Good to Have : Practical experience in developing and deploying Agentic AI or autonomous
  • automation tools to streamline technical tasks.
  • Experience with ServiceNow implementation is a plus
  • Familiarity with ITSM best practices and an understanding of how to align service lifecycles with
  • business goals is preferred.

SKILLS

  • Strong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH)
  • environment.
  • Strong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python
  • or a scripting language with knowledge of standard software development practices.
  • Knowledge of and ability to work on large data communications networks/ Network Protocols and IT
  • infrastructure supporting highly available systems and applications.
  • Strong communication skills and ability to work effectively across multiple technical teams.
  • Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to
  • automate decision-making, optimize complex workflows, and enhance proactive system monitoring.

Vacancy posted 3 hours ago
Similar jobs that could be interesting for youBased on the Reliability Engineer in Berkeley, CA vacancy
  • $139.4k - $205k

     ...lifetime deliveries. We’re focused on how to do the next 10B even better.About the RoleWe are seeking a highly motivated Senior Reliability & Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the... 
    Suggested
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Oakland, CA
    3 days ago
  • $80 per hour

     ...Job Description Hybrid — Berkeley, CA Assignment: 10/26/2026 – 10/27/2027 $80/hr   Role Summary   As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a national-scale HPC facility accessible... 
    Suggested
    Shift work

    LTD GLOBAL, LLC

    Berkeley, CA
    1 day ago
  • $15k

     ...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills... 
    Suggested
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  •  ...multi-day technology designed to keep the electric grid secure and reliable, even during extended periods of stress. By strengthening the...  ...Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function responsible for maintaining... 
    Suggested
    Full time
    Remote work
    Relocation package

    Form Energy, Inc.

    Berkeley, CA
    5 days ago
  • $81.5k - $141.3k

     ...effectively and comfortably, with life-changing products that provide accurate data to drive better-informed decisions.The Reliability Engineer II is an individual contributor with deep expertise in electrical engineering, software development principles, and project... 
    Suggested
    Worldwide
    Shift work

    Abbott

    Alameda, CA
    5 days ago
  •  ...Pacific Gas and Electric Company is seeking an experienced Senior Data Engineer / Manager-level to lead reliability data architecture and governance. You will drive enterprise data capabilities, oversee multi-source integration, and guide data pipelines to support regulatory... 
    Work at office

    Jobleads-US

    Oakland, CA
    3 days ago
  • $68k - $108k

     ...technical teams to resolve operational bottlenecks and support system reliability and service-level objectives. Practical experience...  ...More: We are recruiting for a long-term Site Reliability Engineer contract with our direct client in Berkeley, California. The role... 
    Long term contract
    Full time
    Work at office
    Night shift

    Digital Technology Solutions

    Berkeley, CA
    4 days ago
  • $75 per hour

     ...organizations with the technical talent they need to solve complex problems. For this role, you’ll join the team supporting NERSC, where reliable computing systems help scientists advance research in energy, physics, materials, and more. Location: Onsite at Lawrence... 
    Hourly pay
    Contract work
    Night shift

    Essnova Solutions, Inc.

    Berkeley, CA
    6 days ago
  •  ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco · Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once... 
    Full time
    Remote work

    Andromeda Cluster

    San Francisco, CA
    2 days ago
  •  ...standards you set become the standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs,...  ...technical depth and no spin. Turn recurring customer pain into engineering fixes with the production teams. What We’re Looking For... 

    Fluidstack

    San Francisco, CA
    5 days ago
  • $170k - $250k

     ...About the RoleWe are seeking a highly motivated Senior/Staff Test Engineer to join our team. This individual will play a key role in the...  ...and equipment.Ability to write Python scripts to automate reliability tests.Experience with CAD and shop tools to design and build fixtures... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    1 day ago
  • $165k - $210k

    Keep the systems we operate running, and push what you learn on call back into how we design the next one. Team Operate Type Full-time Band $165k - $210k Location San Francisco or remote (UTC-8 to UTC+2) What we look for Production on-call experience...
    Full time
    Remote work

    Brainfullstack Ltd.

    San Francisco, CA
    2 days ago
  •  ...vehicles powered by clean energy Building more resilient homes with reliable backup Designing a flexible and distributed electrical grid The Role SPAN is seeking a Staff Reliability Engineer to be the technical leader for our reliability team, improve lab... 
    Work at office
    Flexible hours

    SPAN Inc

    San Francisco, CA
    4 days ago
  •  ...that our team members have what they need to do their best work — both in and out of the office. We're looking for a Hardware Reliability Engineer to join us on this journey. You'll take the lead on planning and driving hardware reliability work for Oura's wearable... 
    Contract work
    Work at office
    Local area
    Remote work
    Flexible hours

    Oura

    San Francisco, CA
    5 days ago
  • $160k - $220k

     ...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.Senior Database Reliability Engineer (DBRE) Experience Level: Mid-Senior (4+ years PostgreSQL experience)About the RoleWe are looking for a highly skilled Database... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    4 days ago
  • $180k - $230k

     ...About the Role – Staff Reliability Engineer Peak Energy is looking for a Staff Reliability Engineer to help scale ESS hardware reliability with a focus on HV electronics, devices, and power conversion. You will work cross‑functionally to realize industry‑leading reliability... 
    Immediate start
    Flexible hours

    Peak Energy Services

    San Francisco, CA
    3 days ago
  •  ...high bar, move fast, and care deeply about each other and our customers. About the Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of Scribe’s data tier. Our engineering org is doubling — which means... 
    Full time
    Work at office
    Remote work
    Home office
    Flexible hours
    3 days per week

    scribehow.com

    San Francisco, CA
    2 days ago
  •  ...Senior Database Reliability EngineerSan Francisco, CA, United StatesAbout CrunchyrollFounded by fans, Crunchyroll delivers the art and...  ...millions of anime fans around the world. The Database Operations Engineering team provides a seamless infrastructure foundation to our... 
    Flexible hours

    Crunchyroll

    San Francisco, CA
    2 days ago
  • $150k - $180k

     ...isn’t it.   The Role As we continue to develop and deploy cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance, durability, and robustness of critical hardware systems. This... 
    Full time
    Immediate start
    Worldwide
    Flexible hours
    Night shift

    Eight Sleep

    San Francisco, CA
    a month ago
  •  ...Job Description Job Description Mechanical Reliability Engineer — Test Infrastructure & Thermal‑Fluid SystemsRole snapshot Help keep sophisticated mechanical test assets running at peak performance. In this sustaining-focused role, you’ll support, maintain, and continuously... 

    Technical Talent Group

    Berkeley, CA
    1 day ago
  •  ...platform to launch future Movewear products and transform millions of lives in the coming years. The Role As our Hardware Reliability Engineer, you will be the person who makes sure our products don't just work in the lab -- they work on real humans, in real... 
    Full time
    Relocation
    Flexible hours

    Skip

    San Francisco, CA
    5 days ago
  • $139.76k - $287.75k

     ...GitOps workflows. This role will be instrumental in advancing the reliability, scalability, automation, observability, and operational...  ...and delivery ecosystem.The ideal candidate is a highly hands-on engineer with strong production experience and a proven ability to build... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    3 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $106k - $130k

     ...any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational... 
    Hourly pay
    Full time
    Immediate start
    Visa sponsorship
    Work visa
    Flexible hours

    Early Warning

    San Francisco, CA
    5 days ago
  • $190.8k - $267.1k

     ...engaged communities while helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser success,...  ...experience.The Ads Reliability team partners closely with Ads Engineering to improve reliability, scalability, operational excellence,... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    4 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad... 

    JP Morgan Chase

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical... 
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...Read ouroperating principles to see it in full.About the teamThe Engineering team at Airwallex is a diverse group of innovators, builders,...  ...sense of ownership, working together to build scalable, reliable, and secure products that empower businesses of all sizes to grow... 
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 

    Alembic

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!