Site Reliability Engineer (SRE)
$68k - $108kDigital Technology Solutions
Salary: $68,000 - 108,000 per year Requirements:
- Experience working, or willingness to work, in a 24/7 onsite team supporting large-scale data centers or critical installations.
- Experience using the Linux shell and command-line environments such as SSH.
- Experience developing tools in languages such as C, C++, Perl, Java, Python, or another scripting language, along with knowledge of standard software development practices.
- Self-motivation and ability to learn technologies for data center management, such as Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.
- Experience with network security, including configuring or maintaining ACLs and understanding firewalls.
- Experience collaborating across technical teams to resolve operational bottlenecks and support system reliability and service-level objectives.
- Practical experience developing and deploying Agentic AI or autonomous automation tools is beneficial.
- ServiceNow implementation experience is a plus; familiarity with ITSM best practices and aligning service lifecycles with business goals is preferred.
- Work onsite five days per week on overnight Owl shifts, from midnight to 8 a.m., monitoring our high-performance computing facility.
- Review and respond to alerts from computing, storage, networking, and facility systems; triage issues or contact the appropriate on-call staff.
- Develop process improvements, prevent recurring issues, and automate responses to routine service conditions.
- Identify ways to improve monitoring and automate triage.
- Use ServiceNow to develop and implement customized service management solutions.
- Respond to system alerts and help ensure continuous data collection and real-time diagnostic information.
- Develop and maintain monitoring-pipeline tools with the Operations Team, including software that brings alerts and notifications from HPC system APIs into the pipeline.
- Configure and maintain applications and tools so they operate reliably as data volumes and user demands grow.
- Coordinate with other NERSC groups to clarify communications and workflows, plan center-wide maintenance, and manage diagnostic and notification software during maintenance periods.
- Conduct regular physical and logical walkthroughs of the data center to check environmental conditions, power distribution units, and cooling infrastructure.
- Maintain accurate trouble-ticket records for outages, maintenance updates, and other incidents so teams can track workflows and protocols.
- Analyze and resolve problems of varied scope and complexity, selecting appropriate methods and exercising sound judgment.
- Agentic AI
- AI
- Support
- ITSM
- Java
- Kubernetes
- Linux
- Network
- Perl
- Prometheus
- Python
- Security
- ServiceNow
More:
We are recruiting for a long-term Site Reliability Engineer contract with our direct client in Berkeley, California. The role supports the National Energy Research Scientific Computing Center (NERSC), whose mission is to accelerate scientific discovery through high-performance computing and data analysis for U.S. Department of Energy Office of Science programs. NERSC provides critical computing and data systems and support to more than 11,000 users conducting research across energy, physics, materials science, chemistry, and other DOE mission areas. As part of the Operations Technology Group, you will help keep these services accessible, reliable, and secure, supporting scientific research through continuous, proactive monitoring.
last updated 40 week of 2026
- ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity...Suggested
- ...arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and... ...certificate lifecycles in line with HIPAA. What You Bring 5+years SRE/DevOps experience running production workloads on AWS, GCP or Azure...Suggested
- ...Job Description Job Description DTS is looking for Site Reliability Engineer (SRE) for a long term contract with our direct client Position in Berkeley, CA Job Description: The National Energy Research Scientific Computing Center (NERSC) is inviting applications...SuggestedLong term contractWork at officeNight shift
$80 per hour
...Job Description Job Description Site Reliability Engineer (SRE II) | 100% Onsite | Berkeley, CA | $80/hr | 1-Year Contract Important Notes: ~ Permanent overnight schedule of midnight to 8:00 a.m., five days per week ~100% onsite in Berkeley, California...SuggestedPermanent employmentContract workNight shift$15k
...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills...SuggestedWork at officeLocal areaRemote work- ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system... ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially...
$80 per hour
...HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want...Shift work- ...to keep the electric grid secure and reliable, even during extended periods of stress... ...Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function... ...systems. Experience leading SRE, DevOps, or production operations teams...Full timeRemote workRelocation package
- ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,... ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes... ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform...Remote jobFor contractors
$300k
...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and... .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting...Permanent employment$190.8k - $267.1k
...Reddit grow its business. The reliability of our Ads systems directly... ...team partners closely with Ads Engineering to improve reliability,... .... We’re looking for a Senior Site Reliability Engineer to build... ...operations.Drive adoption of SRE best practices including SLIs...For contractorsWork experience placement$152.5k - $205k
...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate... ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed...Flexible hours- ...it in full.About the teamThe Engineering team at Airwallex is a diverse... ...working together to build scalable, reliable, and secure products that... ...sizes to grow without borders.Our SRE team is breaking new... ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...Temporary workLocal area
$113.4k - $162k
...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,... ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced...Temporary work$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB’s broader... ...a pivotal role in engineering the reliable, globally connected, multi-cloud network... ...OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a...Local areaRemote workWorldwideFlexible hours$117k - $209.33k
...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique...Full timeFor contractors- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the...
$148.5k - $223.9k
...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts... ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving...Full timeWorldwideWeekend work$210k - $240k
...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay... ...We’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability,...Full time- ...San Francisco Bay Area · Global Offices or Remote Available What we're looking for As an SRE at Wordbricks, you will keep our systems fast, reliable, and boring. You'll own the infrastructure and operations behind our products so the rest of the team can ship without...Remote workFlexible hours
$189k - $283.6k
...to everyone. The Role As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience...Full timeRelocation packageFlexible hoursShift work- ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and operations...Immediate startRemote workWorldwide
- ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure... ...treatment. What We Look for in a Great Engineer Tool Proficiency: You are highly... ...feature release while maintaining the highest reliability. DevX Support: Support Developer...Work at office
$150k
...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our cloud infrastructure...$164k - $205k
...alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on operational... ...following, please apply: ~4+ years of experience in SRE or infrastructure roles ~ Genuine excitement about AI...Work experience placementSummer holidayLive outWork at officeLocal areaFlexible hoursShift work2 days per week- ...looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability... .... Preferred background Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading...Full timeLocal areaRemote work
- ...To achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth....WorldwideHome officeFlexible hours
- ...builds the platforms and tooling that help engineering teams develop, deploy, and operate... ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll... ...experience in backend systems, SRE, or platform engineering roles.Proven track...Permanent employmentWork experience placementWork at officeLocal area
$127k - $249k
Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the... ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and...Work at officeLocal areaRemote workWorldwideFlexible hours- ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure... ...enterprise platform support experience and practical SRE/observability fundamentals relevant to ServiceNow10+ years...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!
- construction site safety Berkeley, CA
- IT site lead Berkeley, CA
- site leader Berkeley, CA
- site safety Berkeley, CA
- junior website developer Berkeley, CA
- on-site clinical research associate (traveling/remote) Berkeley, CA
- site reliability engineering manager
- site reliability engineer sre
- site reliability engineer
- junior site reliability engineer


