Site Reliability Engineer (SRE)
$80 per hourEssnova Solutions
Work Location: Onsite – California
Schedule: Full-Time | 5 Days Per Week | Midnight–8:00 AM (Owl Shift)
This position does not offer sponsorship. Must be authorized to work in the United States.
Position Overview
Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance computing (HPC) and data environment supporting scientific research for the U.S. Department of Energy (DOE) Office of Science.
The Site Reliability Engineer will work as part of a 24/7 operations environment responsible for maintaining the accessibility, reliability, security, and operational health of large-scale computing and data systems.
This is a highly hands-on position combining Linux systems administration, infrastructure monitoring, incident response, programming and scripting, automation, networking, ServiceNow, and physical data center operations.
IMPORTANT SCHEDULE REQUIREMENT: This position requires working onsite five days per week on the midnight–8:00 AM shift. Candidates must be willing and able to consistently work this overnight schedule.
Key Responsibilities
- Monitor high-performance computing systems, storage infrastructure, networks, and other data center and facility-related systems.
- Review and respond to infrastructure and system alerts, perform initial triage, and engage appropriate on-call personnel when escalation is required.
- Respond to alerts across multiple systems to help ensure monitoring and data collection remain operational 24/7.
- Troubleshoot system, application, network, monitoring, and infrastructure issues affecting system reliability.
- Develop solutions that improve operational processes, prevent recurring issues, and automate responses to routine service conditions.
- Identify opportunities to improve monitoring capabilities, alerting, incident triage, and operational automation.
- Develop and maintain tools within the monitoring pipeline in collaboration with operations personnel.
- Develop software and integrations capable of generating alerts and notifications from HPC system APIs into monitoring pipelines.
- Build and maintain application and tool configurations to ensure reliable operation as data volumes and user demands increase.
- Utilize ServiceNow to support incident management, trouble-ticketing, operational workflows, and service management activities.
- Collaborate across technical teams to identify and resolve operational bottlenecks and maintain system reliability.
- Coordinate with technical groups during center-wide maintenance activities.
- Manage diagnostic, monitoring, and notification software during planned maintenance periods.
- Perform regular physical and logical walkthroughs of the data center floor.
- Monitor environmental conditions, power distribution units (PDUs), cooling infrastructure, and other facility systems supporting reliable data center operations.
- Maintain accurate trouble-ticket documentation for outages, incidents, maintenance activities, troubleshooting actions, and operational updates.
- Analyze problems of varying complexity and evaluate technical data to determine appropriate troubleshooting and remediation methods.
- Exercise independent technical judgment when selecting methods and approaches for resolving operational issues.
Compensation
$80.00 per hour
The anticipated pay rate for this position is $80.00 per hour . Actual compensation may be determined based on job-related factors including experience, qualifications, skills, contractual requirements, and applicable law.
Equal Employment Opportunity
Essnova Solutions, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, creed, sex, pregnancy, childbirth or related medical conditions, sexual orientation, gender, gender identity or expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, military or veteran status, or any other characteristic protected by applicable federal, state, or local law.
Essnova Solutions, Inc. is committed to providing reasonable accommodations to qualified individuals with disabilities and applicants with disabilities throughout the recruitment and employment process.
Requirements
Required Qualifications
- 5+ years of relevant professional experience in Site Reliability Engineering, systems/infrastructure engineering, DevOps, data center operations, HPC operations, network/system operations, or a closely related technical environment.
- Strong hands-on experience working with Linux , including Linux shell and command-line environments such as SSH.
- Programming and/or scripting experience using one or more languages such as:
- Python
- C
- C++
- Perl
- Java
- Comparable scripting or programming languages
- Knowledge of standard software development practices.
- Experience supporting large-scale IT infrastructure, highly available systems, data centers, critical installations, or comparable technical environments.
- Knowledge of large data communications networks and common network protocols.
- Network security experience, including knowledge of firewalls and access control lists (ACLs) .
- Experience troubleshooting infrastructure, application, system, network, or operational issues.
- Experience responding to monitoring alerts and performing technical incident triage.
- Ability to analyze operational and system data to identify problems and determine appropriate solutions.
- Experience collaborating across multiple technical teams to resolve operational issues and maintain system reliability.
- Strong written and verbal communication skills.
- Ability to independently learn and apply new technologies in a complex technical environment.
- Ability and willingness to work within a 24/7 operational environment .
- Ability and willingness to work onsite five days per week from midnight–8:00 AM.
Education
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline preferred.
- An equivalent combination of education, technical training, certifications, and relevant professional experience may be considered.
Technical Environment
Candidates may work with technologies and platforms including:
- Linux / SSH
- Python
- C / C++
- Perl
- Java
- ServiceNow
- Kubernetes
- Prometheus
- VictoriaMetrics
- Alertmanager
- HPC systems
- Monitoring and alerting pipelines
- Network protocols
- Firewalls and ACLs
- Building management systems
- Data center power and cooling infrastructure
- Infrastructure and operational automation
Candidates are not necessarily expected to have prior experience with every technology listed above but should possess the technical foundation and learning ability necessary to work effectively within a complex computing and data center environment.
Preferred Qualifications
- Experience implementing, configuring, or customizing ServiceNow .
- Familiarity with IT Service Management (ITSM) best practices and service lifecycle management.
- Experience supporting high-performance computing (HPC) environments.
- Experience supporting scientific computing, large-scale data centers, critical infrastructure, or other highly available environments.
- Hands-on experience with Kubernetes .
- Experience with monitoring technologies such as Prometheus, VictoriaMetrics, Alertmanager , or comparable platforms.
- Experience developing monitoring, alerting, infrastructure automation, or incident-response tools.
- Experience developing integrations with system or infrastructure APIs.
- Understanding of data center environmental monitoring, cooling systems, power utilization, and/or building management systems.
- Practical experience developing or deploying Agentic AI or autonomous automation tools to streamline technical operations.
- Experience building autonomous-agent solutions capable of automating technical decision-making, optimizing workflows, or enhancing proactive system monitoring.
$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company ($80M Raised) Office Type: Onsite (4 Days Per Week) Salary: $170,000-$250,000 + Competitive Equity Company Description...SuggestedWork at officeVisa sponsorshipFlexible hours- ...Site Reliability Engineer (SRE) FLUIX is building the AI operating system that plans, designs, and optimizes AI infrastructure. We are based in Silicon Valley. We specialize in providing AI-driven solutions for data centers and power providers, leveraging cutting-edge...SuggestedWork at officeWeekend work
- ...startups across the US. We’re building a pool of world-class Site Reliability Engineers for current roles and for upcoming opportunities. You will... ...into one of our partner startups or added to our vetted SRE network for future projects. This role is ideal for engineers...SuggestedLocal area
$15k
...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills...SuggestedWork at officeLocal areaRemote work$163.71k - $306k
...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...Suggested- We are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity of our platform in San Francisco, California. This role will focus on improving service health, refining observability, and partnering...
$207k - $300k
...pushing for changes that improve reliability and velocity.Define the... ...reliability strategy for Home SRE.Minimum qualifications:Bachelor... ...in Computer Science or Engineering, or a related field.Experience... ...large engineering organizations.Site Reliability Engineering (SRE)...Worldwide$80 per hour
...HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want...Contract workTemporary workShift work- ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$300k
...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and... .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting...Permanent employment- ...OpportunityTo achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth....WorldwideHome officeFlexible hours
$167.7k - $245.2k
...portfolios Your ImpactThe FedRAMP SRE team is focused on our... ...effective.We’re looking for talented engineers with a software or operations... ...teams to ensure the reliability, performance and security of... ...Please see the Cisco careers site to discover more benefits and...Full timeTemporary workWork at officeLocal areaFlexible hours1 day per week$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique...Full timeFor contractors$113.4k - $162k
...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,... ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced...Temporary work$152.5k - $205k
...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate... ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed...Flexible hours$165k - $225.6k
...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability... ...engineering teams to champion DevOps and SRE best practices, deliver excellent internal...Permanent employmentLocal areaWorldwideFlexible hours- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...
$148.5k - $223.9k
...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts... ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving...Full timeWorldwideWeekend work$165k - $227k
...this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group... ...will serve as a key contributor within the EPG SRE organization, partnering closely with software...Local areaWorldwideFlexible hours- ...what’s next.About the teamThe Engineering team at Airwallex is a diverse... ...working together to build scalable, reliable, and secure products that... ...to grow without borders.Our SRE team is breaking new engineering... ....What you’ll doAs a Senior Site Reliability Engineer, you’ll work...Temporary workLocal areaWorldwide
$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites.About the RoleFivetran is building data...Full timeWork at officeRemote work- ...Description The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE...Work at officeNight shift
$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites. About the Role Fivetran is building...Full timeWork at officeRemote work$150k
...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our...$130k - $200k
...Site Reliability Engineer Nscale is the GPU cloud built for AI. We run high-performance, cost-efficient infrastructure for AI-native startups... ...infrastructure that makes AI work. The Role This is a career-level SRE role for someone who wants to own systems, not just watch...Shift work- ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure... ...treatment. What We Look for in a Great Engineer You have the intensity and technical... ...feature release while maintaining the highest reliability. DevX Support: Support Developer...Work at office
- ...would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet,... ...Role We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability,...Remote work1 day per week
- ...Engineering Hiring Sprint We're growing our engineering team and are accelerating hiring... ...Engineers Database Engineers Site Reliability Engineers Extensibility API Engineers... ...infrastructure, platform engineering, SRE, or DevOps. ~ Hands-on ownership of Kubernetes...Work at officeLocal areaFlexible hours
- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure... ...enterprise platform support experience and practical SRE/observability fundamentals relevant to ServiceNow10+ years...
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support... ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and...Work at officeLocal areaRemote workWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- junior website developer Berkeley, CA
- construction site safety Berkeley, CA
- on-site clinical research associate (traveling/remote) Berkeley, CA
- historic site Berkeley, CA
- official site Berkeley, CA
- site leader Berkeley, CA
- site safety Berkeley, CA
- IT site lead Berkeley, CA
- junior site reliability engineer
- site reliability engineer


