Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

$80 per hour

Essnova Solutions


Work Location: Onsite – California
Schedule: Full-Time | 5 Days Per Week | Midnight–8:00 AM (Owl Shift)
This position does not offer sponsorship. Must be authorized to work in the United States.

Position Overview

Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance computing (HPC) and data environment supporting scientific research for the U.S. Department of Energy (DOE) Office of Science.

The Site Reliability Engineer will work as part of a 24/7 operations environment responsible for maintaining the accessibility, reliability, security, and operational health of large-scale computing and data systems.

This is a highly hands-on position combining Linux systems administration, infrastructure monitoring, incident response, programming and scripting, automation, networking, ServiceNow, and physical data center operations.

IMPORTANT SCHEDULE REQUIREMENT: This position requires working onsite five days per week on the midnight–8:00 AM shift. Candidates must be willing and able to consistently work this overnight schedule.

Key Responsibilities
  • Monitor high-performance computing systems, storage infrastructure, networks, and other data center and facility-related systems.
  • Review and respond to infrastructure and system alerts, perform initial triage, and engage appropriate on-call personnel when escalation is required.
  • Respond to alerts across multiple systems to help ensure monitoring and data collection remain operational 24/7.
  • Troubleshoot system, application, network, monitoring, and infrastructure issues affecting system reliability.
  • Develop solutions that improve operational processes, prevent recurring issues, and automate responses to routine service conditions.
  • Identify opportunities to improve monitoring capabilities, alerting, incident triage, and operational automation.
  • Develop and maintain tools within the monitoring pipeline in collaboration with operations personnel.
  • Develop software and integrations capable of generating alerts and notifications from HPC system APIs into monitoring pipelines.
  • Build and maintain application and tool configurations to ensure reliable operation as data volumes and user demands increase.
  • Utilize ServiceNow to support incident management, trouble-ticketing, operational workflows, and service management activities.
  • Collaborate across technical teams to identify and resolve operational bottlenecks and maintain system reliability.
  • Coordinate with technical groups during center-wide maintenance activities.
  • Manage diagnostic, monitoring, and notification software during planned maintenance periods.
  • Perform regular physical and logical walkthroughs of the data center floor.
  • Monitor environmental conditions, power distribution units (PDUs), cooling infrastructure, and other facility systems supporting reliable data center operations.
  • Maintain accurate trouble-ticket documentation for outages, incidents, maintenance activities, troubleshooting actions, and operational updates.
  • Analyze problems of varying complexity and evaluate technical data to determine appropriate troubleshooting and remediation methods.
  • Exercise independent technical judgment when selecting methods and approaches for resolving operational issues.
Compensation

$80.00 per hour

The anticipated pay rate for this position is $80.00 per hour . Actual compensation may be determined based on job-related factors including experience, qualifications, skills, contractual requirements, and applicable law.

Equal Employment Opportunity

Essnova Solutions, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, creed, sex, pregnancy, childbirth or related medical conditions, sexual orientation, gender, gender identity or expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, military or veteran status, or any other characteristic protected by applicable federal, state, or local law.

Essnova Solutions, Inc. is committed to providing reasonable accommodations to qualified individuals with disabilities and applicants with disabilities throughout the recruitment and employment process.

Requirements

Required Qualifications
  • 5+ years of relevant professional experience in Site Reliability Engineering, systems/infrastructure engineering, DevOps, data center operations, HPC operations, network/system operations, or a closely related technical environment.
  • Strong hands-on experience working with Linux , including Linux shell and command-line environments such as SSH.
  • Programming and/or scripting experience using one or more languages such as:
    • Python
    • C
    • C++
    • Perl
    • Java
    • Comparable scripting or programming languages
  • Knowledge of standard software development practices.
  • Experience supporting large-scale IT infrastructure, highly available systems, data centers, critical installations, or comparable technical environments.
  • Knowledge of large data communications networks and common network protocols.
  • Network security experience, including knowledge of firewalls and access control lists (ACLs) .
  • Experience troubleshooting infrastructure, application, system, network, or operational issues.
  • Experience responding to monitoring alerts and performing technical incident triage.
  • Ability to analyze operational and system data to identify problems and determine appropriate solutions.
  • Experience collaborating across multiple technical teams to resolve operational issues and maintain system reliability.
  • Strong written and verbal communication skills.
  • Ability to independently learn and apply new technologies in a complex technical environment.
  • Ability and willingness to work within a 24/7 operational environment .
  • Ability and willingness to work onsite five days per week from midnight–8:00 AM.
Education
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline preferred.
  • An equivalent combination of education, technical training, certifications, and relevant professional experience may be considered.
Technical Environment

Candidates may work with technologies and platforms including:

  • Linux / SSH
  • Python
  • C / C++
  • Perl
  • Java
  • ServiceNow
  • Kubernetes
  • Prometheus
  • VictoriaMetrics
  • Alertmanager
  • HPC systems
  • Monitoring and alerting pipelines
  • Network protocols
  • Firewalls and ACLs
  • Building management systems
  • Data center power and cooling infrastructure
  • Infrastructure and operational automation

Candidates are not necessarily expected to have prior experience with every technology listed above but should possess the technical foundation and learning ability necessary to work effectively within a complex computing and data center environment.

Preferred Qualifications
  • Experience implementing, configuring, or customizing ServiceNow .
  • Familiarity with IT Service Management (ITSM) best practices and service lifecycle management.
  • Experience supporting high-performance computing (HPC) environments.
  • Experience supporting scientific computing, large-scale data centers, critical infrastructure, or other highly available environments.
  • Hands-on experience with Kubernetes .
  • Experience with monitoring technologies such as Prometheus, VictoriaMetrics, Alertmanager , or comparable platforms.
  • Experience developing monitoring, alerting, infrastructure automation, or incident-response tools.
  • Experience developing integrations with system or infrastructure APIs.
  • Understanding of data center environmental monitoring, cooling systems, power utilization, and/or building management systems.
  • Practical experience developing or deploying Agentic AI or autonomous automation tools to streamline technical operations.
  • Experience building autonomous-agent solutions capable of automating technical decision-making, optimizing workflows, or enhancing proactive system monitoring.
Vacancy posted 14 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in Berkeley, CA vacancy
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description... 
    Suggested
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    San Francisco, CA
    3 days ago
  •  ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge... 
    Suggested
    Work at office
    Weekend work

    Fluix AI

    San Francisco, CA
    14 hours ago
  •  ...startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will...  ...into one of our partner startups or added to our vetted SRE network for future projects. This role is ideal for engineers... 
    Suggested
    Local area

    Breakout Tools

    San Francisco, CA
    3 days ago
  • $15k

     ...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills... 
    Suggested
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    5 days ago
  • $163.71k - $306k

     ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 
    Suggested

    Retool

    San Francisco, CA
    2 days ago
  • We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering... 

    Robert Half

    San Francisco, CA
    2 days ago
  • $207k - $300k

     ...pushing for changes that improve reliability and velocity.Define the...  ...reliability strategy for Home SRE.Minimum qualifications:Bachelor...  ...in Computer Science or Engineering, or a related field.Experience...  ...large engineering organizations.Site Reliability Engineering (SRE)... 
    Worldwide

    Google

    San Francisco, CA
    2 days ago
  • $80 per hour

     ...HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want... 
    Contract work
    Temporary work
    Shift work

    LTD Global

    Berkeley, CA
    8 days ago
  •  ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes...  ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    17 days ago
  • $300k

     ...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and...  .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...OpportunityTo achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth.... 
    Worldwide
    Home office
    Flexible hours

    Superhuman

    San Francisco, CA
    2 days ago
  • $167.7k - $245.2k

     ...portfolios Your ImpactThe FedRAMP SRE team is focused on our...  ...effective.We’re looking for talented engineers with a software or operations...  ...teams to ensure the reliability, performance and security of...  ...Please see the Cisco careers site to discover more benefits and... 
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    3 days ago
  • $113.4k - $162k

     ...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced... 
    Temporary work

    TextNow

    San Francisco, CA
    5 days ago
  • $152.5k - $205k

     ...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate...  ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed... 
    Flexible hours

    Circle

    San Francisco, CA
    1 day ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability...  ...engineering teams to champion DevOps and SRE best practices, deliver excellent internal... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 

    Alembic

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts...  ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  • $165k - $227k

     ...this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group...  ...will serve as a key contributor within the EPG SRE organization, partnering closely with software... 
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  •  ...what’s next.About the teamThe Engineering team at Airwallex is a diverse...  ...working together to build scalable, reliable, and secure products that...  ...to grow without borders.Our SRE team is breaking new engineering...  ....What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    2 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites.About the RoleFivetran is building data... 
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    4 days ago
  •  ...Description The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE... 
    Work at office
    Night shift

    Bay Systems Inc

    Berkeley, CA
    4 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites. About the Role Fivetran is building... 
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    4 days ago
  • $150k

     ...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our... 

    VantageScore

    San Francisco, CA
    1 day ago
  • $130k - $200k

     ...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups...  ...infrastructure that makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch... 
    Shift work

    Nscale

    San Francisco, CA
    1 day ago
  •  ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure...  ...treatment. What We Look for in a Great Engineer You have the intensity and technical...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer... 
    Work at office

    Latent

    San Francisco, CA
    3 days ago
  •  ...would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet,...  ...Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability,... 
    Remote work
    1 day per week

    Runloop AI, Inc

    San Francisco, CA
    14 hours ago
  •  ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring...  ...Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers...  ...infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes... 
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    2 days ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure...  ...enterprise platform support experience and practical SRE/observability fundamentals relevant to ServiceNow10+ years... 

    JP Morgan Chase

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support...  ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!