Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Cluster Site Reliability Engineer

The Voleon Group

Voleon is a technology company that applies state‑of‑the‑art machine learning techniques to real‑world problems in finance. For nearly two decades, we have led our industry and worked at the frontier of applying machine learning to investment management. We have become a multibillion‑dollar asset manager, and we have ambitious goals for the future.

As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant. Your work will provide a world‑class HPC platform for researchers to focus on cutting‑edge machine learning problems at scale. You will support both on‑prem and cloud infrastructure, and work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.

The Cluster Operations team works on the frontline to triage and mitigate real‑time operational issues. You will be an integral member of this team, solving day‑to‑day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs.

Responsibilities
  • Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise.
  • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability.
  • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams.
  • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off‑the‑shelf ones won't do.
  • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies.
  • Assist in forecasting cluster growth, and help select appropriate scale‑up strategies. Help optimize operations across dimensions of cost and usability.
Requirements
  • 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead.
  • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod).
  • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.).
  • Familiarity with infrastructure‑as‑code and configuration management tools (Terraform, Ansible).
  • Experience with cloud infrastructure (AWS or GCP).
  • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry).
  • Experience with distributed storage technologies (Lustre, Ceph, S3).
  • Embodies a system engineer rather than a system administrator mindset, thinking systematically and leveraging automation.
  • Bachelor degree in computer science.
Preferred Qualifications
  • Hands‑on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes‑based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark).
  • Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed).
  • Familiarity with hybrid/on‑prem environments.
  • Experience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environments.
  • Experience with HPC networking (InfiniBand, RDMA).
  • Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust).
Equal Opportunity Employer

The Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Cluster Site Reliability Engineer in Berkeley, CA vacancy
  • $250k

     ...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments...  ..., and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • $175k - $250k

     ...250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,...  ...unavailable. Modality: On-Site only. Must live within commuting...  ..., performance, and reliability across environments. What...  ...and automate GPU compute clusters using tools such as... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    3 days ago
  • $164k - $205k

     ...our entire environment Manage and scale Kubernetes clusters that power BetterUp's platform, ensuring high availability...  ...alerting and observability systems Collaborate with engineering teams to embed reliability into the development lifecycle, shifting left on... 
    Senior
    Work experience placement
    Summer holiday
    Live out
    Work at office
    Local area
    Flexible hours
    Shift work
    2 days per week

    SupportFinity

    San Francisco, CA
    9 hours ago
  •  ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated...  ...'re equally comfortable in a Terraform file, a Kubernetes cluster, and a postmortem doc. We expect engineers here to... 
    Senior
    Work at office
    Local area
    Flexible hours

    Airbyte

    San Francisco, CA
    2 days ago
  • $153k - $191.3k

     ...manufacturing, data processing, and software engineering, our office is a truly inspiring mix of...  ...environments, to guarantee the reliability, scalability, and availability of our services...  ...resource optimization, management, and cluster tuning in a constrained environment... 
    Senior
    Full time
    Temporary work
    For contractors
    Work at office
    Local area
    Remote work
    Home office
    3 days per week

    Planet Labs PBC

    San Francisco, CA
    16 hours ago
  • $232.34k - $290.42k

     ...same: to make access to data as simple and reliable as electricity. With Fivetran, customer...  ..., canonical and ready to query, with no engineering or maintenance required. We're proud...  ...integrate our teams, systems, and career sites. About the Role Fivetran and dbt Labs... 
    Senior
    Full time
    Work at office
    Remote work

    dbt Labs

    Oakland, CA
    3 days ago
  •  ...Site Reliability Engineer We are looking for a Site Reliability Engineer to own the digital infrastructure that powers our research. This...  ...Visibility: Provide clear visibility into resource utilization and cluster health. Auto-Scaling: Enable automatic scaling of... 
    Visa sponsorship

    Astera Institute

    Emeryville, CA
    1 day ago
  •  ...US Corp. is seeking a Lead Site Reliability Engineer to spearhead our mission of delivering highly available and performant systems. With an average of over 12 years of industry experience, the successful candidate will bridge the gap between software development and systems... 
    Senior

    Axiom Pursuits

    San Francisco, CA
    9 hours ago
  •  ...About the Role We're looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You'll partner with engineers and data scientists to build, automate, and maintain... 
    Senior

    Alembic Limited

    San Francisco, CA
    16 hours ago
  • $210k - $240k

     ...Join to apply for the Senior Site Reliability Engineer role at Alembic Technologies This range is provided by Alembic Technologies. Your actual pay will be based on your skills and experience — talk with your recruiter to learn more. Base pay range $210,000.00/yr - $2... 
    Senior
    Full time

    Alembic Technologies

    San Francisco, CA
    3 days ago
  •  ...Senior Engineering Role at Salesforce Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here,...  ...Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with... 
    Senior
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...Job Requisition ID # 26WD99273 Position Overview Want to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products. As part of a... 
    Senior
    For contractors

    Autodesk

    San Francisco, CA
    2 days ago
  •  ...of healthcare, we'd love to meet you. Apply now to join our growing team. About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours
    2 days per week

    Plenful

    San Francisco, CA
    1 day ago
  •  ...come shape the future and be part of a truly unique global culture at OutSystems! Hybrid Onsite in Menlo Park, CA Site Reliability Engineering (SRE) is a discipline that incorporates aspects of software engineering and applies them to infrastructure and... 
    Senior
    Immediate start
    Remote work
    Worldwide

    OutSystems

    San Francisco, CA
    3 days ago
  • $81.1k - $187k

     ...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    San Francisco, CA
    2 days ago
  • $160k - $250k

     ...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content...  ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering... 
    Senior

    Hive

    San Francisco, CA
    1 day ago
  • $185.5k - $232k

     ...Senior Site Reliability Engineer New York, NY; Boston, MA; San Francisco, CA About Formation Bio Formation Bio is a tech and AI driven pharma company differentiated by radically more efficient drug development. Advancements in AI and drug discovery are creating... 
    Senior
    Work experience placement
    Work at office
    Local area
    Relocation
    3 days per week

    Formation Bio (Formerly TrailSpark)

    San Francisco, CA
    1 day ago
  • $181k - $225k

     ...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time... 
    Senior
    Work at office
    Immediate start
    3 days per week

    Altruist

    San Francisco, CA
    4 days ago
  • $189k - $283.6k

     ...the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics...  ...accountability ~ A strong desire to perform and grow as an engineer ~5+ years of software development experience... 
    Senior
    Full time
    Relocation package
    Flexible hours
    Shift work

    Block Inc

    San Francisco, CA
    9 hours ago
  • Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...
    Senior
    Temporary work
    Work experience placement

    Phenom People

    San Francisco, CA
    1 day ago
  • $220k - $235k

     ...are seeking a strategic, high‑output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,...  ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud... 
    Senior
    Full time
    Work at office

    Ironclad Inc

    San Francisco, CA
    9 hours ago
  • $181k - $263k

     ...and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Staff Site Reliability Engineer who will set the technical direction for reliability engineering across LiveRamp's global infrastructure. This is a... 
    Senior
    Full time
    Work from home
    Worldwide
    Flexible hours
    Night shift

    LiveRamp

    San Francisco, CA
    3 days ago
  • $300k

     ..., full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation...  ...the operational backbone of one of the largest GPU clusters in private deployment. If you want to build and... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  • $166.59k - $199.91k

     ...access to data as simple and reliable as electricity. With Fivetran...  ...and ready to query, with no engineering or maintenance required. We’re...  ...We’re looking for a talented Senior Full Stack Engineer with a passion...  ..., data security, and cluster orchestration. You don’t necessarily... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Oakland, CA
    3 days ago
  • $99.45k - $134.55k

     ...great opportunity for professional growth. Find your future with us. The Boeing Company is looking for a Site Reliability Engineer (Associate, Experienced or Senior) to join the Air Dominance Site Reliability Engineering team located in Berkeley, MO. We are seeking a... 
    Senior
    Work experience placement
    Interim role
    Currently hiring
    Flexible hours
    Berkeley, CA
    23 days ago
  • $174.92k - $209.91k

    Senior Software Engineer From Fivetran's founding until now, our mission has remained...  ...to data as simple and reliable as electricity. With...  ...teams, systems, and career sites. We're looking for a talented...  ...engineering, data security, and cluster orchestration. You don't... 
    Senior
    Full time
    Work at office
    Remote work

    dbt Labs

    Oakland, CA
    3 days ago
  • $80 per hour

     ...facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want your work to... 
    Contract work
    Shift work

    LTD Global, LLC

    Berkeley, CA
    1 day ago
  •  ...Site Reliability Engineer The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC's mission is to accelerate scientific discovery through high performance computing and data analysis... 
    Work at office
    Night shift

    Bay Systems Consulting Inc

    Berkeley, CA
    1 day ago
  •  ...Site Reliability Engineer Site Reliability Engineer with well-developed organizational, analytical and problem-solving skills for a multi-year engagement with a foremost Healthcare IT Solutions group based in the San Francisco Bay Area. A highly motivated professional... 
    Work experience placement

    Samprasoft

    Oakland, CA
    2 days ago
  • $200k - $300k

     ...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture...  ...fast-paced environment. Manage and optimize Kubernetes clusters for high availability and performance. Develop and maintain... 
    Work at office

    Recruiting from Scratch

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Cluster Site Reliability Engineer. Be the first to apply!