Senior Cluster Site Reliability Engineer [Remote]
$205k - $235kThe Voleon Group
- Remote job
Voleon is a technology company that applies state-of-the-art machine learning techniques to real-world problems in finance. For more than a decade, we have led our industry and worked at the frontier of applying machine learning to investment management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.
As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant. Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale. You will support both on-prem and cloud infrastructure, and work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.
The Cluster Operations team works on the frontline to triage and mitigate real-time operational issues. You will be an integral member of this team, solving day-to-day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs.
Responsibilities
- Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise.
- Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability.
- Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams.
- Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do.
- Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies.
- Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability.
Requirements
- 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead.
- Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod).
- Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)
- Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible).
- Experience with cloud infrastructure (AWS or GCP).
- Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry).
- Experience with distributed storage technologies (Lustre, Ceph, S3).
- Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation.
- Bachelor degree in computer science or equivalent experience.
Preferred Qualifications
- Hands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark).
- Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed).
- Familiarity with hybrid/on-prem environments.
- Experience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environments.
- Experience with HPC networking (InfiniBand, RDMA).
- Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust).
The base salary range for this position is $205,000 to $235,000 in the location(s) of this posting. Individual salaries are determined through a variety of factors, including, but not limited to, education, experience, knowledge, skills, and geography. Base salary does not include other forms of total compensation such as bonus compensation and other benefits. Our benefits package includes medical, dental and vision coverage, life and AD&D insurance, 20 days of paid time off, 9 sick days, and a 401(k) plan with a company match.
“Friends of Voleon” Candidate Referral Program
If you have a great candidate in mind for this role and would like to have the potential to earn $15,000 if your referred candidate is successfully hired and employed by The Voleon Group, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Voleon Referral Bonus Program .
Equal Opportunity Employer
The Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
$250k
...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ..., and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering...SeniorFull timeRemote work- ...Airbyte Infrastructure And Reliability Engineer Airbyte is the data and action layer for AI agents. We give agents fast, accurate, authenticated... ...'re equally comfortable in a Terraform file, a Kubernetes cluster, and a postmortem doc. We expect engineers here to...SeniorWork at officeLocal areaFlexible hours
$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites. About the Role Fivetran is building...SeniorFull timeWork at officeRemote work- ...attract incredibly creative scientists and engineers from leading academic institutions and... ...Summary We are looking for a Site Reliability Engineer to own the digital infrastructure... ...into resource utilization and cluster health. Auto-Scaling: Enable automatic...SuggestedVisa sponsorship
- ...of healthcare, we'd love to meet you. Apply now to join our growing team. About the Role Plenful is hiring a Senior Site Reliability Engineer (SRE) to keep our production systems reliable, performant, and scalable as we grow. This role is centered on operating...SeniorFull timeWork at officeRemote workFlexible hours2 days per week
$160k - $250k
...DevOps And Systems Engineer Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content... ...machine learning models, we also need to grow our DevOps and Site Reliability team to maintain the reliability of our enterprise SaaS offering...Senior$181k - $225k
...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time...SeniorWork at officeImmediate start3 days per week- Job Title At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...SeniorTemporary workWork experience placement
$300k
..., full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation... ...the operational backbone of one of the largest GPU clusters in private deployment. If you want to build and...SeniorPermanent employment$174.92k - $209.91k
...access to data as simple and reliable as electricity. With... ...ready to query, with no engineering or maintenance required.... ..., systems, and career sites. We're looking for a talented Senior Software Engineer with a... ...engineering, data security, and cluster orchestration. You don't...SeniorFull timeWork at officeRemote work$152.5k - $205k
...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and... ...designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.Strong Terraform...SeniorFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$174.92k - $209.91k
...same: to make access to data as simple and reliable as electricity. With Fivetran, customer... ..., canonical and ready to query, with no engineering or maintenance required. We’re proud... ...integrate our teams, systems, and career sites.About the RoleFivetran is building data...SeniorFull timeWork at officeRemote work- ...stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research. This...Full time
$99.45k - $134.55k
...great opportunity for professional growth. Find your future with us. The Boeing Company is looking for a Site Reliability Engineer (Associate, Experienced or Senior) to join the Air Dominance Site Reliability Engineering team located in Berkeley, MO. We are seeking a...SeniorWork experience placementInterim roleCurrently hiringFlexible hours$200k - $300k
...Site Reliability Engineer Title of Role: Site Reliability Engineer Location: San Francisco, onsite Company Stage of Funding: Venture... ...fast-paced environment. Manage and optimize Kubernetes clusters for high availability and performance. Develop and maintain...Work at office$170k - $250k
...Site Reliability Engineer (SRE) Location: San Francisco, CA / Palo Alto, CA Company Stage of Funding: Growth-Stage AI Infrastructure Company... ...Experience supporting AI infrastructure, GPU clusters, machine learning platforms, or accelerated compute environments...Work at officeVisa sponsorshipFlexible hours$80 per hour
...Must be authorized to work in the United States. Position Overview Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance...Hourly payFull timeWork at officeLocal areaShift workNight shift$240k - $265k
...state-of-the-art AI. As an early Senior Backend Software Engineer. This role is focused on scaling our... ...CI/CD pipelines, testing strategy, reliability targets, and observability... ...observability, and running resilient multi-TB clusters with replication and failover. ~...SeniorFull time- ...Compute team's mission: any engineer, using AI, should be able to... ...skills that make this possible. Reliable, observable, and self-service... ...design failure to fix. A Senior engineer will help define the... ...haves Deep K8s internals: cluster lifecycle, admission controllers...SeniorFull timeFor contractorsInternship
$160k - $230k
...lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior Backend Engineer, you will play a key role in building... ...as on-demand + managed Kubernetes and Slurm clusters. This platform serves both our internal StaaS...SeniorFull timeRemote work- ...optimize distributed systems, ensuring reliability, performance, and scalability under real... ...tuning: analyzers, BM25, shard sizing, cluster health, and relevance optimization.... ...individual will need to lead platform engineering managers across multiple teams, present...SeniorFull time
- ...About the Role We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s... ...crucial role in shaping the architecture, reliability, and automation of our Kubernetes-... ...in Go/Python for end-to-end cluster lifecycle management — provisioning,...SeniorFull timeWork at officeLocal areaWork from homeFlexible hours
- ...inference possible. The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage,... ...). We're looking for an experienced Senior Software Engineer to join our storage team. You...SeniorFull timeWork at officeLocal areaWork from homeFlexible hours
$80 per hour
...facility supports 11,000+ scientists pushing the boundaries of what's possible, and we need a sharp, self-motivated SRE to help keep that engine running without interruption. If you love solving real problems on live infrastructure, thrive on ownership, and want your work...Contract workTemporary workShift work- ...Description The National Energy Research Scientific Computing Center (NERSC) is inviting applications for the position of Site Reliability Engineer. NERSC’s mission is to accelerate scientific discovery through high performance computing and data analysis for the DOE...Work at officeNight shift
- ...help run and evolve that infrastructure, working alongside the senior engineers already on the team. You'll contribute to architecture... ...an operator's logs and a custom resource's status. Customer clusters on AWS, Azure, and GCP are provisioned with Crossplane compositions...Remote workFlexible hours
$204k - $259k
...to more cities. In this hybrid role, you will report to an Engineering Manager. You will: Develop business logic software to... ...and simulated driving, understanding, characterizing and clustering the performance of the Waymo Driver Interact with ML models...SeniorFull timeRemote work$194k - $267k
...educate on new concepts and tools. Position Overview: The Site Reliability Engineer (SRE) will play a key role in building and managing... ...scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience...Permanent employmentWork at officeLocal areaWorldwideFlexible hours- ...infrastructure company is hiring a Frontend Engineer to design and build the interface for... ...and real-time UIs that translate complex cluster state into clear, actionable experiences.... ...a high-impact, early-team role based on-site in San Francisco. What You'll Do Own...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Cluster Site Reliability Engineer [Remote]. Be the first to apply!
- senior developer Berkeley, CA
- senior aws cloud engineer Berkeley, CA
- remote senior salesforce administrator Berkeley, CA
- senior manager tax Berkeley, CA
- senior creative Berkeley, CA
- senior brand designer Berkeley, CA
- senior accountant remote Berkeley, CA
- senior manager accenture Berkeley, CA
- senior information technology project manager Berkeley, CA
- senior implementation engineer Berkeley, CA



