Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff — Compute Cluster

Linuxcareers

This AI research company is building large physics foundation models for causal intelligence and weather prediction. You will design, build, and operate large-scale GPU clusters that power training, evaluation, and serving infrastructure for the research team. What You\'ll Do Design, deploy, and operate large distributed GPU clusters end-to-end, including provisioning, imaging, upgrades, and capacity planning Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage Monitor and continuously improve reliability and error recovery; build observability to catch failures before researchers do Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs What You Need Experience operating large-scale GPU clusters and container orchestration frameworks such as Kubernetes, Slurm, and Docker Strong systems background in Linux, networking, storage, and infrastructure-as-code Knowledge of cloud platforms including GCP, AWS, or Azure and their ML/AI service offerings Understanding of monitoring, logging, observability, and version control best practices for ML systems Familiarity with CUDA and NCCL, and performance profiling for distributed workloads Ability to own deliverables end-to-end from requirements through autonomous execution #J-18808-Ljbffr Linuxcareers

Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff — Compute Cluster. Be the first to apply!