Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE)

$68k - $108k
Full-time

Digital Technology Solutions

Salary: $68,000 - 108,000 per year Requirements:

  • Experience working, or willingness to work, in a 24/7 onsite team supporting large-scale data centers or critical installations.
  • Experience using the Linux shell and command-line environments such as SSH.
  • Experience developing tools in languages such as C, C++, Perl, Java, Python, or another scripting language, along with knowledge of standard software development practices.
  • Self-motivation and ability to learn technologies for data center management, such as Kubernetes, Prometheus/VictoriaMetrics, Alertmanager, building management software, evaporative cooling, and power utilization.
  • Experience with network security, including configuring or maintaining ACLs and understanding firewalls.
  • Experience collaborating across technical teams to resolve operational bottlenecks and support system reliability and service-level objectives.
  • Practical experience developing and deploying Agentic AI or autonomous automation tools is beneficial.
  • ServiceNow implementation experience is a plus; familiarity with ITSM best practices and aligning service lifecycles with business goals is preferred.
Responsibilities:
  • Work onsite five days per week on overnight Owl shifts, from midnight to 8 a.m., monitoring our high-performance computing facility.
  • Review and respond to alerts from computing, storage, networking, and facility systems; triage issues or contact the appropriate on-call staff.
  • Develop process improvements, prevent recurring issues, and automate responses to routine service conditions.
  • Identify ways to improve monitoring and automate triage.
  • Use ServiceNow to develop and implement customized service management solutions.
  • Respond to system alerts and help ensure continuous data collection and real-time diagnostic information.
  • Develop and maintain monitoring-pipeline tools with the Operations Team, including software that brings alerts and notifications from HPC system APIs into the pipeline.
  • Configure and maintain applications and tools so they operate reliably as data volumes and user demands grow.
  • Coordinate with other NERSC groups to clarify communications and workflows, plan center-wide maintenance, and manage diagnostic and notification software during maintenance periods.
  • Conduct regular physical and logical walkthroughs of the data center to check environmental conditions, power distribution units, and cooling infrastructure.
  • Maintain accurate trouble-ticket records for outages, maintenance updates, and other incidents so teams can track workflows and protocols.
  • Analyze and resolve problems of varied scope and complexity, selecting appropriate methods and exercising sound judgment.
Technologies:
  • Agentic AI
  • AI
  • Support
  • ITSM
  • Java
  • Kubernetes
  • Linux
  • Network
  • Perl
  • Prometheus
  • Python
  • Security
  • ServiceNow

More:

We are recruiting for a long-term Site Reliability Engineer contract with our direct client in Berkeley, California. The role supports the National Energy Research Scientific Computing Center (NERSC), whose mission is to accelerate scientific discovery through high-performance computing and data analysis for U.S. Department of Energy Office of Science programs. NERSC provides critical computing and data systems and support to more than 11,000 users conducting research across energy, physics, materials science, chemistry, and other DOE mission areas. As part of the Operations Technology Group, you will help keep these services accessible, reliable, and secure, supporting scientific research through continuous, proactive monitoring.

last updated 40 week of 2026

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) in Berkeley, CA vacancy
  •  ...Open role Site Reliability Engineer (SRE) San Francisco, CA (On-site) Responsibilities Develop and maintain advanced monitoring, alerting, and self-healing mechanisms that detect and address issues before they impact customers. Perform regular capacity... 
    Suggested

    Methodic

    San Francisco, CA
    3 days ago
  •  ...arc of the patient journey. The Opportunity: Machine Learning Engineer Patients count on our platform 24/7. You'll build and...  ...certificate lifecycles in line with HIPAA. What You Bring 5+years SRE/DevOps experience running production workloads on AWS, GCP or Azure... 
    Suggested

    Tala Health

    San Francisco, CA
    2 days ago
  •  ...Job Description Job Description DTS is looking for Site Reliability Engineer (SRE) for a long term contract with our direct client Position in Berkeley, CA     Job Description: The National Energy Research Scientific Computing Center (NERSC) is inviting applications... 
    Suggested
    Long term contract
    Work at office
    Night shift

    Digital Technology Solutions

    Berkeley, CA
    1 day ago
  • $80 per hour

     ...Job Description Job Description Site Reliability Engineer (SRE II) | 100% Onsite | Berkeley, CA | $80/hr | 1-Year Contract Important Notes: ~ Permanent overnight schedule of midnight to 8:00 a.m., five days per week ~100% onsite in Berkeley, California... 
    Suggested
    Permanent employment
    Contract work
    Night shift

    Perfect Timing Personnel Services, Inc.

    Berkeley, CA
    1 day ago
  • $15k

     ...packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills... 
    Suggested
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  •  ...their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system...  ..., Support, and TAMs to trust. Partner with product engineers on infrastructure requirements for new Retool products, especially... 

    re-tool®

    San Francisco, CA
    3 days ago
  • $80 per hour

     ...HPC facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want... 
    Shift work

    LTD Global, LLC

    Berkeley, CA
    3 days ago
  •  ...to keep the electric grid secure and reliable, even during extended periods of stress...  ...Description Form Energy is hiring a Manager, Site Reliability Engineer to lead the operational function...  ...systems. Experience leading SRE, DevOps, or production operations teams... 
    Full time
    Remote work
    Relocation package

    Form Energy, Inc.

    Berkeley, CA
    5 days ago
  •  ...development, cloud infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...workflows for accuracy and reliability. Work with AWS, Azure, GCP, Kubernetes...  ...DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    San Francisco, CA
    10 days ago
  • $300k

     ...experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and...  .... Skills / Must Have: ~7+ years of experience in SRE, DevOps, or Infrastructure Engineering roles supporting... 
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $190.8k - $267.1k

     ...Reddit grow its business. The reliability of our Ads systems directly...  ...team partners closely with Ads Engineering to improve reliability,...  .... We’re looking for a Senior Site Reliability Engineer to build...  ...operations.Drive adoption of SRE best practices including SLIs... 
    For contractors
    Work experience placement

    Reddit

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate...  ...public-cloud environments. This role is for an experienced SRE or infrastructure engineer who enjoys solving hard distributed... 
    Flexible hours

    Circle

    San Francisco, CA
    4 days ago
  •  ...it in full.About the teamThe Engineering team at Airwallex is a diverse...  ...working together to build scalable, reliable, and secure products that...  ...sizes to grow without borders.Our SRE team is breaking new...  ...services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Temporary work
    Local area

    Airwallex

    San Francisco, CA
    4 days ago
  • $113.4k - $162k

     ...conversation for people everywhere.TextNow is looking for motivated Site Reliability Engineer to own infrastructure, monitoring, logging, ci/cd,...  ...practices. Contribute to the design and implementation of new SRE best practices.You'll be a great fit if you have:Experienced... 
    Temporary work

    TextNow

    San Francisco, CA
    2 days ago
  • $127k - $249k

    The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB’s broader...  ...a pivotal role in engineering the reliable, globally connected, multi-cloud network...  ...OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a... 
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    7 hours ago
  • $117k - $209.33k

     ...Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ...cloud services for Autodesk GovCloud products.As part of a new SRE team supporting Autodesk GovCloud, you will have a unique... 
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    7 hours ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the... 

    Alembic

    San Francisco, CA
    4 days ago
  • $148.5k - $223.9k

     ...future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts...  ...and our customers protected. The ExperienceAs an SRE, you will be a technical leader of the team driving... 
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay...  ...We’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability,... 
    Full time

    Alembic Technologies

    San Francisco, CA
    5 days ago
  •  ...San Francisco Bay Area · Global Offices or Remote Available What we're looking for As an SRE at Wordbricks, you will keep our systems fast, reliable, and boring. You'll own the infrastructure and operations behind our products so the rest of the team can ship without... 
    Remote work
    Flexible hours

    Wordbricks, Inc.

    San Francisco, CA
    2 days ago
  • $189k - $283.6k

     ...to everyone. The Role As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure...  ...~ A strong desire to perform and grow as an engineer ~5+ years of software development experience... 
    Full time
    Relocation package
    Flexible hours
    Shift work

    Block Inc

    San Francisco, CA
    2 days ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and operations... 
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    2 days ago
  •  ...SRE Location: San Francisco, CA (5 Days In-Office) You are the infrastructure...  ...treatment. What We Look for in a Great Engineer Tool Proficiency: You are highly...  ...feature release while maintaining the highest reliability. DevX Support: Support Developer... 
    Work at office

    Latent

    San Francisco, CA
    4 days ago
  • $150k

     ...About The Role We are seeking an experienced Site Reliability Engineer (SRE) with a strong focus on DevSecOps to join our growing engineering team. In this role, you will oversee and maintain the reliability, security posture, and operational hygiene of our cloud infrastructure... 

    VantageScore®

    San Francisco, CA
    4 days ago
  • $164k - $205k

     ...alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on operational...  ...following, please apply: ~4+ years of experience in SRE or infrastructure roles ~ Genuine excitement about AI... 
    Work experience placement
    Summer holiday
    Live out
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    SupportFinity

    San Francisco, CA
    2 days ago
  •  ...looking for a midlevel or senior IC to join our Backend Engineering team as a Site Reliability Engineer. You'll own the uptime, performance, and observability...  .... Preferred background Full-time experience in SRE, DevOps, or backend engineering, preferably at a trading... 
    Full time
    Local area
    Remote work

    Databento

    San Francisco, CA
    2 days ago
  •  ...To achieve our ambitious goals, we’re looking for an SRE to join our infrastructure team. This role will be responsible for building software to ensure the reliability of our back-end systems, working with engineers who develop them, and planning for our future growth.... 
    Worldwide
    Home office
    Flexible hours

    Coda

    San Francisco, CA
    4 days ago
  •  ...builds the platforms and tooling that help engineering teams develop, deploy, and operate...  ...default for every product team.As a Staff Site Reliability Engineer on Release Engineering, you'll...  ...experience in backend systems, SRE, or platform engineering roles.Proven track... 
    Permanent employment
    Work experience placement
    Work at office
    Local area

    Plaid Financial

    San Francisco, CA
    3 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions that support the...  ...alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  •  ...the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the Enterprise Technology, Infrastructure...  ...enterprise platform support experience and practical SRE/observability fundamentals relevant to ServiceNow10+ years... 

    JP Morgan Chase

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE). Be the first to apply!