Reliability Engineer
Bay Systems
Job Description
Job Description
Company: Lawrence Berkeley National Laboratory ( Through Bay Systems Consulting Inc. )
1 year full-time Contract.
Extension is based on company budget and performance
National Energy Research Scientific Computing Center (NERSC) - Site Reliability Engineer.
As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment.
DUTIES & RESPONSIBILITIES
- Onsite 5-day weekly schedule consisting of Owl (midnight–8 am) shifts to monitor the NERSC HPC Facility.
- Review and respond to alerts from computer systems, storage, network, and other data
- center/facility-related systems by triaging or calling the appropriate on-call staff.
- Create appropriate solutions to improve processes, prevent issue recurrence, and automate
- responses to all routine service conditions.
- Identify issues and propose solutions that will improve monitoring capabilities or provide better
- automation for triage.
- Possess expertise in ServiceNow and its usage to develop and implement customized service
- management solutions.
- Respond to alerts from multiple systems to ensure that data collection continues 24/7, providing
- real-time information for diagnoses.
- Develop and maintain tools within the monitoring pipeline in collaboration with the Operations
- Team.
- Create new software to provide alerts and notifications from HPC system APIs into the
- monitoring pipeline.
- Builds and maintains application/tool configurations to ensure software runs reliably as data and
- user demands grow.
- Collaborate with other groups at NERSC to ensure that communication and workflows are clearly
- understood.
- Work closely with other NERSC groups to coordinate center-wide maintenance activities and
- manage diagnostic and notification software during maintenance periods.
- Perform regular physical and logical walkthroughs of the data center floor to monitor
- environmental health, power distribution units, and cooling infrastructure to ensure peak
- operational efficiency.
- Provide accurate information in the trouble ticketing system for outages, maintenance updates,
- and other incidents so that workflows and protocols can be appropriately tracked by others.
- Work on and resolve problems of diverse scope where data analysis requires the evaluation of
- identifiable factors.
- Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
- Work on and resolve complex issues where the analysis of situations or data requires an in-depth
- evaluation of variable factors.
REQUIREMENTS
- Experience in or willingness to work within a 24/7 onsite team environment to support
- large-scale data centers or critical installations.
- Experience on Linux shell and working in a command-line (e.g. SSH) environment.
- Experience with developing tools using various programming languages such as C, C++, Perl,
- Java, or Python or a scripting language with knowledge of standard software development
- practices.
- Motivated, self-starter who can learn technologies that improve data center management in
- areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management
- software, evaporative
- cooling, and power utilization.
- Experience with network security: configuring/maintaining ACLs, knowledge of firewalls
- Experience collaborating across technical teams to resolve operational bottlenecks and ensure
- system reliability and alignment with service-level objectives.
- Good to Have : Practical experience in developing and deploying Agentic AI or autonomous
- automation tools to streamline technical tasks.
- Experience with ServiceNow implementation is a plus
- Familiarity with ITSM best practices and an understanding of how to align service lifecycles with
- business goals is preferred.
SKILLS
- Strong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH)
- environment.
- Strong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python
- or a scripting language with knowledge of standard software development practices.
- Knowledge of and ability to work on large data communications networks/ Network Protocols and IT
- infrastructure supporting highly available systems and applications.
- Strong communication skills and ability to work effectively across multiple technical teams.
- Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to
- automate decision-making, optimize complex workflows, and enhance proactive system monitoring.
$139.4k - $205k
...lifetime deliveries. We’re focused on how to do the next 10B even better.About the RoleWe are seeking a highly motivated Senior Reliability & Test Engineer to join our team. This individual will play a key role in the development and validation of our unmanned platforms at the...SuggestedHourly payWork at officeLocal areaRemote workFlexible hours$80 per hour
...Job Description Hybrid — Berkeley, CA Assignment: 10/26/2026 – 10/27/2027 $80/hr Role Summary As a Site Reliability Engineer on the Operations Technology team, you'll be part of a round-the-clock crew keeping a national-scale HPC facility accessible...SuggestedShift work$15k
...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills...SuggestedWork at officeLocal areaRemote work- ...multi-day technology designed to keep the electric grid secure and reliable, even during extended periods of stress. By strengthening the... ...Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function responsible for maintaining...SuggestedFull timeRemote workRelocation package
$81.5k - $141.3k
...effectively and comfortably, with life-changing products that provide accurate data to drive better-informed decisions.The Reliability Engineer II is an individual contributor with deep expertise in electrical engineering, software development principles, and project...SuggestedWorldwideShift work- ...Pacific Gas and Electric Company is seeking an experienced Senior Data Engineer / Manager-level to lead reliability data architecture and governance. You will drive enterprise data capabilities, oversee multi-source integration, and guide data pipelines to support regulatory...Work at office
$68k - $108k
...technical teams to resolve operational bottlenecks and support system reliability and service-level objectives. Practical experience... ...More: We are recruiting for a long-term Site Reliability Engineer contract with our direct client in Berkeley, California. The role...Long term contractFull timeWork at officeNight shift$75 per hour
...organizations with the technical talent they need to solve complex problems. For this role, you’ll join the team supporting NERSC, where reliable computing systems help scientists advance research in energy, physics, materials, and more. Location: Onsite at Lawrence...Hourly payContract workNight shift- ...Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco · Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once...Full timeRemote work
- ...standards you set become the standard. Role Scope Own reliability for named customer workloads: their clusters, their SLAs,... ...technical depth and no spin. Turn recurring customer pain into engineering fixes with the production teams. What We’re Looking For...
$170k - $250k
...About the RoleWe are seeking a highly motivated Senior/Staff Test Engineer to join our team. This individual will play a key role in the... ...and equipment.Ability to write Python scripts to automate reliability tests.Experience with CAD and shop tools to design and build fixtures...Hourly payWork at officeLocal areaRemote workFlexible hours$165k - $210k
Keep the systems we operate running, and push what you learn on call back into how we design the next one. Team Operate Type Full-time Band $165k - $210k Location San Francisco or remote (UTC-8 to UTC+2) What we look for Production on-call experience...Full timeRemote work- ...vehicles powered by clean energy Building more resilient homes with reliable backup Designing a flexible and distributed electrical grid The Role SPAN is seeking a Staff Reliability Engineer to be the technical leader for our reliability team, improve lab...Work at officeFlexible hours
- ...that our team members have what they need to do their best work — both in and out of the office. We're looking for a Hardware Reliability Engineer to join us on this journey. You'll take the lead on planning and driving hardware reliability work for Oura's wearable...Contract workWork at officeLocal areaRemote workFlexible hours
$160k - $220k
...This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.Senior Database Reliability Engineer (DBRE) Experience Level: Mid-Senior (4+ years PostgreSQL experience)About the RoleWe are looking for a highly skilled Database...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$180k - $230k
...About the Role – Staff Reliability Engineer Peak Energy is looking for a Staff Reliability Engineer to help scale ESS hardware reliability with a focus on HV electronics, devices, and power conversion. You will work cross‑functionally to realize industry‑leading reliability...Immediate startFlexible hours- ...high bar, move fast, and care deeply about each other and our customers. About the Role We’re hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of Scribe’s data tier. Our engineering org is doubling — which means...Full timeWork at officeRemote workHome officeFlexible hours3 days per week
- ...Senior Database Reliability EngineerSan Francisco, CA, United StatesAbout CrunchyrollFounded by fans, Crunchyroll delivers the art and... ...millions of anime fans around the world. The Database Operations Engineering team provides a seamless infrastructure foundation to our...Flexible hours
$150k - $180k
...isn’t it. The Role As we continue to develop and deploy cutting-edge autonomous technologies, we are seeking a Senior Reliability Engineer (REL) to lead efforts in ensuring the long-term performance, durability, and robustness of critical hardware systems. This...Full timeImmediate startWorldwideFlexible hoursNight shift- ...Job Description Job Description Mechanical Reliability Engineer — Test Infrastructure & Thermal‑Fluid SystemsRole snapshot Help keep sophisticated mechanical test assets running at peak performance. In this sustaining-focused role, you’ll support, maintain, and continuously...
- ...platform to launch future Movewear products and transform millions of lives in the coming years. The Role As our Hardware Reliability Engineer, you will be the person who makes sure our products don't just work in the lab -- they work on real humans, in real...Full timeRelocationFlexible hours
$139.76k - $287.75k
...GitOps workflows. This role will be instrumental in advancing the reliability, scalability, automation, observability, and operational... ...and delivery ecosystem.The ideal candidate is a highly hands-on engineer with strong production experience and a proven ability to build...Work at officeLocal areaRelocationRelocation package$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper). As...Work at officeLocal areaRemote workWorldwideFlexible hours$106k - $130k
...any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.Role Summary The Senior Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational...Hourly payFull timeImmediate startVisa sponsorshipWork visaFlexible hours$190.8k - $267.1k
...engaged communities while helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser success,... ...experience.The Ads Reliability team partners closely with Ads Engineering to improve reliability, scalability, operational excellence,...For contractorsWork experience placement- ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure Platforms team, you will solve complex and broad...
$152.5k - $205k
...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical...Flexible hours- ...Read ouroperating principles to see it in full.About the teamThe Engineering team at Airwallex is a diverse group of innovators, builders,... ...sense of ownership, working together to build scalable, reliable, and secure products that empower businesses of all sizes to grow...Temporary workLocal area
$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,...Full timeWorldwideWeekend work- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Reliability Engineer. Be the first to apply!
- database reliability engineer
- reliability maintenance engineering technician
- principal reliability engineer
- fixed equipment reliability engineer
- reliability engineer
- reliability engineering manager
- maintenance & reliability engineer
- sr reliability engineer
- senior reliability engineer
- network reliability engineer




