Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Cluster Site Reliability Engineer [Remote]

$205k - $235k

The Voleon Group

Berkeley, CA
  • Remote job

Voleon is a technology company that applies state-of-the-art machine learning techniques to real-world problems in finance. For more than a decade, we have led our industry and worked at the frontier of applying machine learning to investment management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future. 

As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant.  Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale.  You will support both on-prem and cloud infrastructure, and work to provide the best experience to our technical staff.  You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.

The Cluster Operations team works on the frontline to triage and mitigate real-time operational issues. You will be an integral member of this team, solving day-to-day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs. 

Responsibilities

  • Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise.
  • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability.
  • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams.
  • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do.
  • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies.
  • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability.

Requirements

  • 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead.
  • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod).
  • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)
  • Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible).
  • Experience with cloud infrastructure (AWS or GCP).
  • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry).
  • Experience with distributed storage technologies (Lustre, Ceph, S3).
  • Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation.
  • Bachelor degree in computer science or equivalent experience.

Preferred Qualifications

  • Hands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark).
  • Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed).
  • Familiarity with hybrid/on-prem environments.
  • Experience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environments.
  • Experience with HPC networking (InfiniBand, RDMA).
  • Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust).

The base salary range for this position is $205,000 to $235,000 in the location(s) of this posting. Individual salaries are determined through a variety of factors, including, but not limited to, education, experience, knowledge, skills, and geography. Base salary does not include other forms of total compensation such as bonus compensation and other benefits. Our benefits package includes medical, dental and vision coverage, life and AD&D insurance, 20 days of paid time off, 9 sick days, and a 401(k) plan with a company match.

“Friends of Voleon” Candidate Referral Program

If you have a great candidate in mind for this role and would like to have the potential to earn $15,000 if your referred candidate is successfully hired and employed by The Voleon Group, please use this  form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the  Voleon Referral Bonus Program .

Equal Opportunity Employer

The Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Senior Cluster Site Reliability Engineer [Remote] in Berkeley, CA vacancy
  • $250k

     ...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...  ..., and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated...  ...'re equally comfortable in a Terraform file, a Kubernetes cluster, and a postmortem doc. We expect engineers here to... 
    Senior
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    22 hours ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites. About the Role Fivetran is building... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    3 days ago
  •  ...attract incredibly creative scientists and engineers from leading academic institutions and...  ...Summary We are looking for a Site Reliability Engineer to own the digital infrastructure...  ...into resource utilization and cluster health. Auto-Scaling: Enable automatic... 
    Suggested
    Visa sponsorship

    Astera

    Emeryville, CA
    2 days ago
  •  ...of healthcare, we'd love to meet you. Apply now to join our growing team. About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    5 days ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content...  ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering... 
    Senior

    Hive

    San Francisco, CA
    4 days ago
  • $181k - $225k

     ...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time... 
    Senior
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    3 days ago
  • Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...
    Senior
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    4 days ago
  • $300k

     ..., full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation...  ...the operational backbone of one of the largest GPU clusters in private deployment. If you want to build and... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $174.92k - $209.91k

     ...access to data as simple and reliable as electricity. With...  ...ready to query, with no engineering or maintenance required....  ..., systems, and career sites. We're looking for a talented Senior Software Engineer with a...  ...engineering, data security, and cluster orchestration. You don't... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    4 days ago
  • $152.5k - $205k

     ...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and...  ...designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.Strong Terraform... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    5 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    8 days ago
  • $174.92k - $209.91k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We’re proud...  ...integrate our teams, systems, and career sites.About the RoleFivetran is building data... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    8 days ago
  •  ...stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research. This... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $99.45k - $134.55k

     ...great opportunity for professional growth. Find your future with us. The Boeing Company is looking for a Site Reliability Engineer (Associate, Experienced or Senior) to join the Air Dominance Site Reliability Engineering team located in Berkeley, MO. We are seeking a... 
    Senior
    Work experience placement
    Interim role
    Currently hiring
    Flexible hours
    Berkeley, CA
    6 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture...  ...fast-paced environment. Manage and optimize Kubernetes clusters for high availability and performance. Develop and maintain... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    4 days ago
  • $170k - $250k

     ...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company...  ...Experience supporting AI infrastructure, GPU clusters, machine learning platforms, or accelerated compute environments... 
    Work at office
    Visa sponsorship
    Flexible hours

    Recruiting from Scratch

    San Francisco, CA
    2 days ago
  • $80 per hour

     ...Must be authorized to work in the United States. Position Overview Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance... 
    Hourly pay
    Full time
    Work at office
    Local area
    Shift work
    Night shift

    Essnova Solutions

    Berkeley, CA
    4 days ago
  • $240k - $265k

     ...state-of-the-art AI. As an early Senior Backend Software Engineer. This role is focused on scaling our...  ...CI/CD pipelines, testing strategy, reliability targets, and observability...  ...observability, and running resilient multi-TB clusters with replication and failover. ~... 
    Senior
    Full time

    Unify

    San Francisco, CA
    1 day ago
  •  ...Compute team's mission: any engineer, using AI, should be able to...  ...skills that make this possible. Reliable, observable, and self-service...  ...design failure to fix. A Senior engineer will help define the...  ...haves Deep K8s internals: cluster lifecycle, admission controllers... 
    Senior
    Full time
    For contractors
    Internship

    Persona

    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior Backend Engineer, you will play a key role in building...  ...as on-demand + managed Kubernetes and Slurm clusters. This platform serves both our internal StaaS... 
    Senior
    Full time
    Remote work

    Together Ai

    San Francisco, CA
    1 day ago
  •  ...optimize distributed systems, ensuring reliability, performance, and scalability under real...  ...tuning: analyzers, BM25, shard sizing, cluster health, and relevance optimization....  ...individual will need to lead platform engineering managers across multiple teams, present... 
    Senior
    Full time

    Maintainx

    San Francisco, CA
    1 day ago
  •  ...About the Role We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s...  ...crucial role in shaping the architecture, reliability, and automation of our Kubernetes-...  ...in Go/Python for end-to-end cluster lifecycle management — provisioning,... 
    Senior
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  •  ...inference possible. The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage,...  ...). We're looking for an experienced Senior Software Engineer to join our storage team. You... 
    Senior
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  • $80 per hour

     ...facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want your work... 
    Contract work
    Temporary work
    Shift work

    LTD Global

    Berkeley, CA
    12 days ago
  •  ...Description The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE... 
    Work at office
    Night shift

    Bay Systems

    Berkeley, CA
    10 days ago
  •  ...help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture...  ...an operator's logs and a custom resource's status. Customer clusters on AWS, Azure, and GCP are provisioned with Crossplane compositions... 
    Remote work
    Flexible hours

    Akka

    San Francisco, CA
    13 days ago
  • $204k - $259k

     ...to more cities.  In this hybrid role, you will report to an Engineering Manager. You will: Develop business logic software to...  ...and simulated driving, understanding, characterizing and clustering the performance of the Waymo Driver Interact with ML models... 
    Senior
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • $194k - $267k

     ...educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing...  ...scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    6 days ago
  •  ...infrastructure company is hiring a Frontend Engineer to design and build the interface for...  ...and real-time UIs that translate complex cluster state into clear, actionable experiences....  ...a high-impact, early-team role based on-site in San Francisco. What You'll Do Own... 
    Remote work

    Clera

    San Francisco, CA
    17 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Cluster Site Reliability Engineer [Remote]. Be the first to apply!