Senior Cluster Site Reliability Engineer
$15kThe Voleon Group
Voleon is a technology company that applies state-of-the-art AI and machine learning techniques to real-world problems in finance. For nearly two decades, we have led our industry and worked at the frontier of applying AI/ML to investment management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.Your colleagues will include internationally recognized experts in artificial intelligence and machine learning research as well as highly experienced finance and technology professionals. In addition to our enriching and collegial working environment, we offer highly competitive compensation and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant. Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale. You will support both on-prem and cloud infrastructure, and work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.The Cluster Operations team works on the frontline to triage and mitigate real-time operational issues. You will be an integral member of this team, solving day-to-day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs. ResponsibilitiesBe a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they ariseEnsure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliabilityDiagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teamsDevelop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't doHelp software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policiesAssist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usabilityRequirements5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech leadKnowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)Experience with cloud infrastructure (AWS or GCP)Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)Experience with distributed storage technologies (Lustre, Ceph, S3)Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automationBachelor degree in computer sciencePreferred QualificationsHands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark)Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed)Familiarity with hybrid/on-prem environmentsExperience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environmentsExperience with HPC networking (InfiniBand, RDMA)Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust)“Friends of Voleon” Candidate Referral ProgramIf you have a great candidate in mind for this role and would like to have the potential to earn $15,000 if your referred candidate is successfully hired and employed by The Voleon Group, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Voleon Referral Bonus Program.Equal Opportunity EmployerThe Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation TypeRemoteDepartmentSoftwareCompensationBase Salary $205K – $235K • Offers BonusThe listed base salary range for this position is based upon the location(s) of this posting. Individual salaries are determined through a variety of factors, including, but not limited to, education, experience, knowledge, skills, and geography. Base salary does not include other forms of total compensation such as bonus compensation and other benefits. Our benefits package includes medical, dental, and vision coverage, life and AD&D insurance, 20 days of paid time off, 9 sick days, and a 401(k) plan with a company match.
$250k
...in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments... ..., and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering...SeniorPermanent employmentRemote work$152.5k - $205k
...everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and... ...designing, operating, securing, and troubleshooting production clusters and containerized workloads at scale.Strong Terraform...SeniorFlexible hours$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$215k - $275k
...scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.Proud... ...raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide...SeniorWork at office- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools... .... Build, maintain, monitor and improve our Kubernetes clusters. Work with development teams on migrating applications...Senior
- ...grown 800% over the last 12 months. Engineering at Ivo Engineers at Ivo are... ...without sacrificing accuracy [2024] Clustering legal documents descended from the same... ...SLAs. What? We're looking for a Senior Site level Reliability Engineer as part of Infrastructure...SeniorContract workWork at officeRemote workVisa sponsorshipRelocation packageFlexible hours
$175k - $250k
...50,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco,... ...unavailable. Modality: On-Site only. Must live within commuting... ..., performance, and reliability across environments.... ...and automate GPU compute clusters using tools such as Python...SeniorFull timeRemote workRelocationRelocation package$155k - $222.6k
...cloud platform. As a team of six engineers distributed across the US,... ...strong focus on automation, reliability, and operational excellence.... ...to critical projects such as cluster build out by building automation... ...~2+ years of experience in Site Reliability Engineering, DevOps...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
- ...attract incredibly creative scientists and engineers from leading academic institutions and... ...Summary We are looking for a Site Reliability Engineer to own the digital infrastructure... ...into resource utilization and cluster health. Auto-Scaling: Enable automatic...Visa sponsorship
$139.76k - $287.75k
...their business.We are seeking a Senior Site ReliabilityEngineer to help... ...in advancing the reliability, scalability, automation, observability... ...is a highly hands-on engineer with strong production experience... ...in Kubernetes, including cluster operations, troubleshooting,...SeniorWork at officeLocal areaRelocationRelocation package$117k - $209.33k
Job Requisition ID #26WD99273Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable, secure, and scalable cloud services for Autodesk GovCloud products.As part of a new SRE team supporting...SeniorFull timeFor contractors$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...SeniorPermanent employmentLocal areaWorldwideFlexible hours$140k - $205k
Senior Technology Site Reliability EngineerCooley is seeking a Senior Site Reliability Engineer to join the Infrastructure & Development Operations team.Position summary: The Senior Technology Site Reliability Engineer (“SRE”) is responsible for ensuring the reliability...SeniorFull timeTemporary workWork at officeFlexible hoursWeekend work- ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ...ownership, working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal areaWorldwide
$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations...SeniorFull timeWorldwideWeekend work$165k - $241.4k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$220k - $235k
...are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...SeniorFull timeContract workWork at office- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
$227.2k - $324.5k
About the Role:Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies... ...automation.We are seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site...SeniorFull timeContract workTemporary workLocal areaFlexible hours- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves...Senior
- ...getting here.)About the RoleWe're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.You'...Senior
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$165k - $241.4k
...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$15k
...ambitious goals for the future.As a Senior Software Engineer in Strategy Research Analytics, you will... ...shape technical direction, establish reliability standards, and drive consolidation... ...cloud technologies and on-prem compute clusters (e.g., Slurm, SSH, Unix)Exposure to...SeniorLocal area$350k
...the Tinker community. About the Role We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work... ...at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads. Logistics Location...Full timeVisa sponsorshipWork visaRelocation package$174.92k - $209.91k
...access to data as simple and reliable as electricity. With... ...ready to query, with no engineering or maintenance required.... ..., systems, and career sites.We’re looking for a talented Senior Software Engineer with a... ...engineering, data security, and cluster orchestration. You don’t...SeniorFull timeWork at officeRemote work- ...Udaip Cloud-Based Data And Ai Platform Engineer At U.S. Bank, we're on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it...SeniorTemporary workWork experience placement
$81.1k - $187k
...Site Reliability Engineer 3 We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving...SeniorTemporary workImmediate startFlexible hoursShift work$166.9k - $225.9k
...Summary: Drata's SRE team operates as both a central engineering function and an embedded reliability practice. You'll be part of a close-knit SRE team... ...What you'll bring: ~6+ years of experience in Site Reliability Engineering, Cloud Engineering, or building...SeniorWork at officeImmediate startWorldwideMonday to FridayFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Cluster Site Reliability Engineer. Be the first to apply!
- senior manager tax Berkeley, CA
- senior international accountant Berkeley, CA
- senior vmware engineer Berkeley, CA
- senior resident engineer Berkeley, CA
- senior performance engineer Berkeley, CA
- senior storage engineer Berkeley, CA
- senior director diversity & inclusion Berkeley, CA
- senior manager Berkeley, CA
- remote senior project manager Berkeley, CA
- senior implementation project manager Berkeley, CA

