Member of Technical Staff, Cluster Administration
$200k - $400kInferact
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.
You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.
Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.
Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.
Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.
Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.
Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.
Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.
Preferred qualifications:
Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.
Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.
Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.
Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.
Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.
Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.
Bonus points if you have:
Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.
Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.
Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.
Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.
Carried real operational responsibility for infrastructure used by many engineers or researchers.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
- Member of Technical Staff (Production Services) at CockroachDB CockroachDB Observability Distributed Systems SQL Accessibility Category-defining... ...directly on the systems that protect customer data, handle cluster connectivity and disaster recovery, and power the operational...SuggestedLocal areaRemote workWorldwideFlexible hours
$240k - $280k
...Direct message the job poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $280,000 You know how... ...engineers, working on: AI systems that make healthcare administration actually work smoothly Tools that help clinics deliver cutting...SuggestedFull timeRemote workWorldwideRelocation$300k
...the ground up. About the Role We’re looking for a deeply technical Member of Technical Staff to own RL and post‑training for large‑scale omni models.... ...mixed training/inference workloads across large GPU clusters. Experience with adjacent areas such as distributed pretraining...SuggestedH1bWork at officeVisa sponsorshipShift work- ...work at scale. This role demands deep technical expertise in production environments... ...ensure you can do your best work. As a Member of Technical Staff, you will: ~Design and write high-... ..., and experiment with ideas on our cluster and data infrastructure. ~...SuggestedFull timeLocal areaRemote work
- Member Of Technical Staff - Extreme-Scale Sparse Linear Algebra, Domain Decomposition & GPU Solver Architecture Vinci | Full-Time | Remote / Hybrid... ...solvers Deterministic parallel reductions across GPU clusters AI-accelerated solver components grounded in numerical...SuggestedFull timeRemote work
$324k - $396k
...concisely and accurately share knowledge with their teammates. Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques... ...data processing systems Building large-scale Kubernetes clusters for data storage, processing, and analysis on on-prem...Remote work$324k - $396k
About the Role Member of Technical Staff (X.AI LLC; Palo Alto, CA): Introduce innovative techniques and analyses to the AI field to facilitate... ...Hadoop, BigQuery. Experience building large-scale Kubernetes clusters for data storage, processing, and analysis on on-prem...Remote work$150k - $300k
...systems into our RL training stack. Core Technical Responsibilities LLM Serving Multi‑... ...distribution and cold‑start times across clusters. Inference Optimization & Performance... ...in open development and encourage team members to contribute to the broader AI community...Work at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work$150k - $300k
...and full fine-tuning runs on managed GPU clusters with a single API call or a few clicks.... ...infrastructure that runs the jobs. Core Technical Responsibilities Hosted Training... ...in open development and encourage team members to contribute to the broader AI community...Work at officeLocal areaRemote workVisa sponsorshipRelocation packageFlexible hours- Member of the Technical Staff, Systems Location: North America Remote / San Francisco, CA · Full-Time About Andromeda Compute is the most sought-after... ...infrastructure once reserved for hyperscalers. The first cluster filled almost instantly. The years since went into the...Full timeRemote work
$180k - $280k
...drowning in repetitive tasks, paper-based bottlenecks, and fragmented data – and we're building the software to fix it. As a Member of the Technical Staff at Finch, you'll own critical features end-to-end, ship fast, and work directly alongside product, ops, and design in a...Work at officeRemote workFlexible hours1 day per week$250k
...the creation of maintainable, scalable systems and make sound technical decisions. You lead large projects from ideation to delivery, balancing... ...excellent, and you actively invest in the growth of your team members. You can hold a team to high standards while being comfortable...H1bWork at officeWork from homeHome officeRelocation packageFlexible hours3 days per week- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As a founding member of the engineering team, you will impact the design and direction of Pixeltable at a formative stage, contributing to some of our most foundational...Full timePart timeWork at officeWork from homeFlexible hours2 days per week
- ...Career Launch is hiring for this role and related opportunities. Career Launch is hiring candidates for Member of Technical Staff roles and similar opportunities with fast-moving teams. This is a remote-friendly opportunity for candidates who are analytical, practical,...Remote work
$200k - $300k
...Member of Technical Staff — ML Infra (Data) Seattle, Washington About Nuance Labs Nuance Labs is building photorealistic, real-time AI avatars with emotional intelligence: a full-duplex audiovisual system that can listen, speak, react, interrupt, and respond like a real...H1bWork at officeVisa sponsorship- ...Member of Technical Staff, Machine Learning Protingent Staffing has an exciting direct hire Member of Technical Staff, Machine Learning with our client that is fully remote. Job Description As a Member of Technical Staff, Machine Learning, you will build core ML components...Remote work
- Role Description As a Senior Member of Technical Staff specializing in web data for pre-training, you will play a pivotal role in developing the large scale web data pipeline that underpins Cohere’s advanced language models. In this role, you will work extensively with...Full timeWork at officeRemote workFlexible hours
$200k - $350k
...behalf of a rapidly scaling technology company building sophisticated infrastructure and software systems. We are seeking a Member of Technical Staff with experience in YC or Venture backed start ups to lead high-impact technical initiatives and solve complex engineering...Full time- ...node hot-swapping ~Design storage systems for fast model checkpointing ~Container migration, inference autoscaling and multi-cluster serving ~Preemption handling and sandboxes for training, RL rollouts, and evals ~Make the AI stack run great out of the box...Full time
$110k - $370k
...At Cohere, we care deeply about building technologies that are broadly accessible and useful, regardless of language. As a Member of Technical Staff on the Multilingual team, you'll be at the forefront of advancing language models that serve the world. You’ll push the...Full timeWork at officeLocal areaRemote work$200k - $300.09k
Role Description The OpenClaw Foundation is seeking exceptional Members of Technical Staff (MTS) to serve as full-time maintainers, builders, and stewards of the OpenClaw ecosystem. This is not a traditional software engineering role. ~Operate as both an open-source...Full time- Role Description We are looking for a Member of Technical Staff, Research to investigate, design, test and develop state of the art (SOTA) methods and applications, which can be integrated into the broader AI engine FirstPrinciples is developing. You will collaborate with...Full timeRemote work
$200k - $350k
...a rapidly scaling technology company building sophisticated infrastructure and software systems. We are seeking a Senior Member of Technical Staff to build core systems, solve challenging engineering problems, and contribute directly to the evolution of the company's technology...Full time- ...Participating in on-call rotations Qualifications ~Impressive technical work you can go deep on, with impact in the world. That can... ...AI infrastructure once reserved for hyperscalers. The first cluster filled almost instantly. The years since went into the platform...Full time
- ...on-prem — deciding where compute should live, keeping jobs and clusters consistent across unreliable infrastructure, and recovering... ...looking for an engineer to own this core end to end, set its technical direction and build new features to make it more robust, performant...Full time
$200k - $350k
...rapidly scaling technology company building sophisticated infrastructure and software systems. We are seeking a Principal Member of Technical Staff to provide technical leadership across the company's most important engineering challenges. ~Define architecture for complex...Full time- ...open-source core — a hosted API server that gives infrastructure teams unified scheduling, governance, and security across all their clusters and clouds, while compute stays in their own environment (BYOC). We're looking for an engineer to own this control plane end to...Full time
- ...model behavior. Our objective is to help users complete tasks daily enjoyable with over 90%* reduced time. Role As a Member of Technical Staff, Machine Learning, you will build core ML components. You will work on real production systems from day one, learning how...Full time
$220k - $250k
Our client, an AI-driven healthcare company focused on personalizing patient care, is hiring a Member of Technical Staff to join their team remotely. The successful candidate will help build and scale the AI and data systems powering next-generation primary care solutions...Full timeRemote work- ...Shape how advanced AI systems reason through complex legal work. As a Member of Technical Staff focused on Legal Research, you will develop rigorous evaluation methods, datasets, and benchmarks for legal AI, working at the intersection of large language models, agentic...Full timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff, Cluster Administration. Be the first to apply!
- life support technician Remote
- personal computer support technician Remote
- systems support technician Remote
- technical support analyst Remote
- user support analyst Remote
- help desk technical support Remote
- senior technical analyst Remote
- technical support specialist Remote
- IT assistant Remote
- help desk assistant Remote














