Site Reliability Engineer
Bay Systems
Job Description
Job Description
The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE Office of Science programs. NERSC provides critical HPC and data systems and support for NERSC’s 11,000 plus users researching energy, physics, materials science, and chemistry and other DOE mission areas. As a Site Reliability Engineer in the Operations Technology Group, you will be a member of a 24x7 team that helps ensure NERSC is accessible, reliable, and secure for our scientific users. This position leverages advanced data collection and monitoring systems to proactively manage the health of our environment. Ultimately, your work ensures that NERSC’s computational power remains an uninterrupted catalyst supporting fundamental scientific research related to energy.
ESSENTIAL DUTIES & RESPONSIBILITIES
- Works an onsite 5-day weekly schedule consisting of Owl (midnight–8 am) shifts to monitor the NERSC HPC Facility.
- Review and respond to alerts from computer systems, storage, network, and other data center/facility-related systems by triaging or calling the appropriate on-call staff.
- Create appropriate solutions to improve processes, prevent issue recurrence, and automate responses to all routine service conditions.
- Identify issues and propose solutions that will improve monitoring capabilities or provide better automation for triage.
- Possess expertise in ServiceNow and its usage to develop and implement customized service management solutions.
- Respond to alerts from multiple systems to ensure that data collection continues 24/7, providing real-time information for diagnoses.
- Develop and maintain tools within the monitoring pipeline in collaboration with the Operations Team.
- Create new software to provide alerts and notifications from HPC system APIs into the monitoring pipeline.
- Builds and maintains application/tool configurations to ensure software runs reliably as data and user demands grow.
- Collaborate with other groups at NERSC to ensure that communication and workflows are clearly understood.
- Work closely with other NERSC groups to coordinate center-wide maintenance activities and manage diagnostic and notification software during maintenance periods.
- Perform regular physical and logical walkthroughs of the data center floor to monitor environmental health, power distribution units, and cooling infrastructure to ensure peak operational efficiency.
- Provide accurate information in the trouble ticketing system for outages, maintenance updates, and other incidents so that workflows and protocols can be appropriately tracked by others.
- Work on and resolve problems of diverse scope where data analysis requires the evaluation of identifiable factors.
- Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
- Work on and resolve complex issues where the analysis of situations or data requires an in-depth evaluation of variable factors.
POSITION REQUIREMENTS
- Experience in or willingness to work within a 24/7 onsite team environment to support large-scale data centers or critical installations.
- Experience on Linux shell and working in a command-line (e.g. SSH) environment.
- Experience with developing tools using various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
- Motivated, self-starter who can learn technologies that improve data center management in areas like Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative
- cooling, and power utilization.
- Experience with network security: configuring/maintaining ACLs, knowledge of firewalls
- Experience collaborating across technical teams to resolve operational bottlenecks and ensure system reliability and alignment with service-level objectives.
Good to Have :
- Practical experience in developing and deploying Agentic AI or autonomous automation tools to streamline technical tasks.
- Experience with ServiceNow implementation is a plus
- Familiarity with ITSM best practices and an understanding of how to align service lifecycles with business goals is preferred.
Knowledge, Skills & Abilities
- Strong hands-on knowledge of the Linux shell and working in a command-line (e.g. SSH) environment.
- Strong hands-on knowledge on various programming languages such as C, C++, Perl, Java, or Python or a scripting language with knowledge of standard software development practices.
- Knowledge of and ability to work on large data communications networks/ Network Protocols and IT infrastructure supporting highly available systems and applications.
- Strong communication skills and ability to work effectively across multiple technical teams.
- Good to Have : Ability to build, and deploy Agentic AI solutions that utilize autonomous agents to automate decision-making, optimize complex workflows, and enhance proactive system monitoring.
$15k
...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills...SuggestedWork at officeLocal areaRemote work$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites.About the RoleFivetran is building data...SuggestedFull timeWork at officeRemote work$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites. About the Role Fivetran is building...SuggestedFull timeWork at officeRemote work- ...ambitious goals and attract incredibly creative scientists and engineers from leading academic institutions and from frontier AI labs... ...human brain. Position Summary We are looking for a Site Reliability Engineer to own the digital infrastructure that powers our...SuggestedVisa sponsorship
$80 per hour
...Must be authorized to work in the United States. Position Overview Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance...SuggestedHourly payFull timeWork at officeLocal areaShift workNight shift$80 per hour
...facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want your work...Contract workTemporary workShift work$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...Full timeFor contractors$114.3k - $235.32k
...verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business.We are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native platform built on AWS, Kubernetes/EKS, and ArgoCD-driven...Work at officeLocal areaRelocationRelocation package$113.4k - $162k
...break down barriers to communication and free the flow of conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd, reliability and everything in between!This role is about impact at...Temporary work- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...
$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations,...Full timeWorldwideWeekend work$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...Permanent employmentLocal areaWorldwideFlexible hours- ...an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth. You will work with our existing production...WorldwideHome officeFlexible hours
$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...Flexible hours$152.5k - $205k
...work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind critical...Flexible hours- ...let’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ..., working together to build scalable, reliable, and secure products that empower businesses... ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work closely...Temporary workLocal areaWorldwide
$190.8k - $267.1k
...while helping Reddit grow its business. The reliability of our Ads systems directly impacts... ...Reliability team partners closely with Ads Engineering teams to improve reliability,... ...advertising ecosystem.We're looking for a Staff Site Reliability Engineer who will define and...For contractorsWork experience placementRemote workFlexible hours$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable...Local areaWorldwideFlexible hours- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...Work at officeWeekend work
$200k - $300k
...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture Round - Healthcare, AI Office Type: Onsite Salary: $200K-$300K Company Description We're representing a dynamic...Work at office- ...Site Reliability Engineer We are looking for a dynamic engineer to join our rapidly growing SRE team. As an SRE, you will report to our VP of Technical Operations and be responsible for operating an extremely high performance and scalable, low latency platform built...Relocation package
$181k - $225k
...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time...Work at officeImmediate start3 days per week$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description...Work at officeVisa sponsorshipFlexible hours$163.71k - $306k
...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content... ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering...- ...Infrastructure team builds the platforms and tooling that help engineering teams develop, deploy, and operate production systems safely... ...safe shipping the default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll define and scale...Permanent employmentWork experience placementWork at officeLocal area
- ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring through a focused Engineering Hiring Sprint... ...: Platform Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers AI Agents Engineers...Work at officeLocal areaFlexible hours
$150k
...Site Reliability Engineer San Francisco, CA About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security...$98.58k - $138.02k
...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering...Work at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
- junior website developer Berkeley, CA
- construction site safety Berkeley, CA
- on-site clinical research associate (traveling/remote) Berkeley, CA
- historic site Berkeley, CA
- official site Berkeley, CA
- site leader Berkeley, CA
- site safety Berkeley, CA
- IT site lead Berkeley, CA
- junior site reliability engineer
- site reliability engineer


