Senior Cluster Site Reliability Engineer
$15kThe Voleon Group
Voleon is a technology company that applies state-of-the-art AI and machine learning techniques to real-world problems in finance. For nearly two decades, we have led our industry and worked at the frontier of applying AI/ML to investment management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.Your colleagues will include internationally recognized experts in artificial intelligence and machine learning research as well as highly experienced finance and technology professionals. In addition to our enriching and collegial working environment, we offer highly competitive compensation and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant. Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale. You will support both on-prem and cloud infrastructure, and work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.The Cluster Operations team works on the frontline to triage and mitigate real-time operational issues. You will be an integral member of this team, solving day-to-day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs. ResponsibilitiesBe a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they ariseEnsure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliabilityDiagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teamsDevelop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't doHelp software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policiesAssist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usabilityRequirements5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech leadKnowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)Experience with cloud infrastructure (AWS or GCP)Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)Experience with distributed storage technologies (Lustre, Ceph, S3)Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automationBachelor degree in computer sciencePreferred QualificationsHands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark)Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed)Familiarity with hybrid/on-prem environmentsExperience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environmentsExperience with HPC networking (InfiniBand, RDMA)Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust)“Friends of Voleon” Candidate Referral ProgramIf you have a great candidate in mind for this role and would like to have the potential to earn $15,000 if your referred candidate is successfully hired and employed by The Voleon Group, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Voleon Referral Bonus Program.Equal Opportunity EmployerThe Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation TypeRemoteDepartmentSoftwareCompensationBase Salary $205K – $235K • Offers BonusThe listed base salary range for this position is based upon the location(s) of this posting. Individual salaries are determined through a variety of factors, including, but not limited to, education, experience, knowledge, skills, and geography. Base salary does not include other forms of total compensation such as bonus compensation and other benefits. Our benefits package includes medical, dental, and vision coverage, life and AD&D insurance, 20 days of paid time off, 9 sick days, and a 401(k) plan with a company match.
$250k
...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ..., and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering...SeniorFull timeRemote work$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$175k - $250k
...50,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,... ...unavailable. Modality: On-Site only. Must live within commuting... ..., performance, and reliability across environments.... ...and automate GPU compute clusters using tools such as Python...SeniorFull timeRemote workRelocationRelocation package- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
$190.8k - $267.1k
...helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser... ...team partners closely with Ads Engineering to improve reliability, scalability, operational... ...advertiser trust. We’re looking for a Senior Site Reliability Engineer to build, operate,...SeniorFor contractorsWork experience placement$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...a pivotal role in engineering the reliable, globally connected, multi-cloud... ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...SeniorLocal areaRemote workWorldwideFlexible hours$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible forThe Site Reliability Engineer builds and maintains shared platform capabilities, common libraries, and infrastructure that help Circle teams ship secure...SeniorFlexible hours$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SeniorFull timeFor contractors$139.76k - $287.75k
...their business.We are seeking a Senior Site ReliabilityEngineer to help... ...in advancing the reliability, scalability, automation, observability... ...is a highly hands-on engineer with strong production experience... ...in Kubernetes, including cluster operations, troubleshooting,...SeniorWork at officeLocal areaRelocationRelocation package$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations...SeniorFull timeWorldwideWeekend work$165k - $227k
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.The Engineering OpportunityWe are looking for an experienced Senior Site Reliability Engineer to join Okta's Emerging Products Group (EPG). Our mission is to build highly reliable...SeniorLocal areaWorldwideFlexible hours- ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ...ownership, working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal areaWorldwide
$167.7k - $245.2k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...Lambda Inc. in San Francisco is seeking a Storage Engineer to own the reliability, performance, and capacity of our production storage fleet across multiple data centers, using a software-defined data plane. You will build monitoring, dashboards, and alerting for storage...Senior
- ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'...Senior
$200k - $260k
...the future of professional services is being written today — and we're just getting started.Role OverviewAs a Software Engineer on the Site Reliability team at Harvey, you will ensure the reliability, scalability, and performance of our legal AI platform. You'll join a...SeniorRelocation package$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$200.7k - $250.9k
...washed away in a flood in 1942, the Royal Engineers rebuilt it. Then it washed away again in... ...opportunities for improvements in reliability/observability/performance/preparedness and... ...candidate for the role: Has past Site Reliability Engineering or DevOps experience...Senior- Job TitleAt U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition...Senior
$200k - $240k
...systems across all product teams. You will collaborate closely with engineering leadership, product managers, and cross-functional teams to... ...and Helm ~ Understand the importance of performant and reliable systems ~ Education - Ideally looking for a B.A. / B.S. degree...SeniorWork at officeImmediate start3 days per week$167.7k - $245.2k
...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$15k
...ambitious goals for the future.As a Senior Software Engineer in Strategy Research Analytics, you will... ...shape technical direction, establish reliability standards, and drive consolidation... ...cloud technologies and on-prem compute clusters (e.g., Slurm, SSH, Unix)Exposure to...SeniorLocal area$232k - $319k
...to help us continue to scale the service with great people and reliable, cost-effective, and efficient infrastructure, processes, and... ...enabled with self-serviceAccelerate the velocity of SRE and product engineering by developing robust platforms, powerful tooling, and...SeniorPermanent employmentLocal areaWorldwideFlexible hours$300k
..., full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation... ...the operational backbone of one of the largest GPU clusters in private deployment. If you want to build and...SeniorPermanent employment$174.92k - $209.91k
...Senior Software Engineer From Fivetran's founding until now, our mission has... ...access to data as simple and reliable as electricity. With... ...teams, systems, and career sites. We're looking for a talented... ...engineering, data security, and cluster orchestration. You don't...SeniorFull timeWork at officeRemote work$15k
...goals for the future.Your TeamAs a Senior or Staff Software Engineer on our Data Engineering team, you will... ...promote our research effort through reliable delivery of high-quality data.Build... ...PostgreSQL, Artifactory, Ceph, Redis)Cluster management and containerization...SeniorLocal area- ...stabilize and keep these hyperscale supercomputers reliable and efficient during the training of the frontier models. About the Role We are looking for engineers to operate the next generation of compute clusters that power OpenAI’s frontier research. This...Full time
$194k - $267k
...self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing... ...scalable, and fault-tolerant Kubernetes platforms. Ensure clusters are optimized for production workloads, providing high resilience...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$99.45k - $134.55k
...great opportunity for professional growth. Find your future with us. The Boeing Company is looking for a Site Reliability Engineer (Associate, Experienced or Senior) to join the Air Dominance Site Reliability Engineering team located in Berkeley, MO. We are seeking a...SeniorWork experience placementInterim roleCurrently hiringFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Cluster Site Reliability Engineer. Be the first to apply!
- senior grant accountant Berkeley, CA
- senior consulting engineer Berkeley, CA
- sr electrical engineer Berkeley, CA
- senior assistant Berkeley, CA
- senior brand strategist Berkeley, CA
- sr accountant Berkeley, CA
- senior medical science liaison Berkeley, CA
- senior property accountant Berkeley, CA
- senior software engineer remote Berkeley, CA
- senior mulesoft developer Berkeley, CA



