Member of Technical Staff - GPU Infrastructure
$150k - $300kPrime Intellect
Building Open Superintelligence Infrastructure Prime Intellect is building the open superintelligence stack - from frontier agentic models to the infra that enables anyone to create, train, and deploy them. We aggregate and orchestrate global compute into a single control plane and pair it with the full RL post-training stack: environments, secure sandboxes, verifiable evals, and our async RL trainer. We enable researchers, startups and enterprises to run end-to-end reinforcement learning at frontier scale, adapting models to real tools, workflows, and deployment contexts. As our Solutions Architect for GPU Infrastructure, you'll be the technical expert who transforms customer requirements into production‑ready systems capable of training the world’s most advanced AI models. We recently raised $15mm in funding (total of $20mm raised) led by Founders Fund, with participation from Menlo Ventures and prominent angels including Andrej Karpathy (Eureka AI, Tesla, OpenAI), Tri Dao (Chief Scientific Officer of Together AI), Dylan Patel (SemiAnalysis), Clem Delangue (Huggingface), Emad Mostaque (Stability AI) and many others. Core Technical Responsibilities Customer Architecture & Design Partner with clients to understand workload requirements and design optimal GPU cluster architectures Create technical proposals and capacity planning for clusters ranging from 100 to 10,000+ GPUs Develop deployment strategies for LLM training, inference, and HPC workloads Present architectural recommendations to technical and executive stakeholders Infrastructure Deployment & Optimization Deploy and configure orchestration systems including SLURM and Kubernetes for distributed workloads Implement high‑performance networking with InfiniBand, RoCE, and NVLink interconnects Optimize GPU utilization, memory management, and inter‑node communication Configure parallel filesystems (Lustre, BeeGFS, GPFS) for optimal I/O performance Tune system performance from kernel parameters to CUDA configurations Production Operations & Support Serve as primary technical escalation point for customer infrastructure issues Diagnose and resolve complex problems across the full stack - hardware, drivers, networking, and software Implement monitoring, alerting, and automated remediation systems Provide 24/7 on‑call support for critical customer deployments Create runbooks and documentation for customer operations teams Technical Requirements Required Experience 3+ years hands‑on experience with GPU clusters and HPC environments Deep expertise with SLURM and Kubernetes in production GPU settings Proven experience with InfiniBand configuration and troubleshooting Strong understanding of NVIDIA GPU architecture, CUDA ecosystem, and driver stack Experience with infrastructure automation tools (Ansible, Terraform) Proficiency in Python, Bash, and systems programming Track record of customer‑facing technical leadership Infrastructure Skills NVIDIA driver installation and troubleshooting (CUDA, Fabric Manager, DCGM) Container runtime configuration for GPUs (Docker, Containerd, Enroot) Linux kernel tuning and performance optimization Network topology design for AI workloads Power and cooling requirements for high‑density GPU deployments Nice to Have Experience with 1000+ GPU deployments NVIDIA DGX, HGX, or SuperPOD certification Distributed training frameworks (PyTorch FSDP, DeepSpeed, Megatron‑LM) ML framework optimization and profiling Experience with AMD MI300 or Intel Gaudi accelerators Contributions to open‑source HPC/AI infrastructure projects Growth Opportunity You’ll work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure. You’ll collaborate with our world‑class engineering team while having direct impact on systems powering the next generation of AI breakthroughs. We value expertise and customer obsession - if you’re passionate about building reliable, high‑performance GPU infrastructure and have a track record of successful large‑scale deployments, we want to talk to you. Apply now and join us in our mission to democratize access to planetary scale computing. Compensation Cash Compensation Range of $150-300k plus Equity Incentives #J-18808-Ljbffr Prime Intellect
- ...is 5 people with a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own the cloud systems that serve our compression... ...'d get to build global low-latency, high-throughput GPU ML inference infra that sits in the critical path of...SuggestedVisa sponsorship
$250k
...compute platform building the next generation of agentic infrastructure for GPU-intensive workloads. Operating across the full... ...scale. This opportunity offers the chance to join as a Member of Technical Staff at a pivotal stage in the company's growth. You'll help...SuggestedFull time- Member of Technical Staff - Infrastructure Security We're partnering with a frontier AI research company that is building next-generation open-weight foundation... ...strategy from day one Work across cloud, Kubernetes, GPU infrastructure, CI/CD, and platform engineering...Suggested
- ...beyond. Role Overview Reflection.AI is looking for a Member of Technical Staff - Infrastructure Security to secure our geographically diverse multi-cloud... ..., AWS, and/or Azure Experience working with neocloud GPU providers such as VoltagePark, GMI Cloud, Crusoe, Anyscale...SuggestedRelocation package
$150k - $350k
...cost with today’s homogeneous, vertically integrated infrastructure. Gimlet addresses this by decoupling AI workloads... ...AI datacenters. Mission Gimlet Labs is seeking a Member of Technical Staff focused on kernels and GPU performance. In this role, you will work close to...Suggested$200k
Member of Technical Staff, Supercomputing Platform & Infrastructure Magic’s mission is to build safe AGI that accelerates humanity’s progress on the world’s most important... ...will design, build, and operate the large-scale GPU infrastructure that powers Magic’s model training...RelocationVisa sponsorship- # Founding Member of Technical Staff, AI Infrastructure**Location:** San Francisco / Bay Area preferred. Remote exceptional for the right person.We look for... ...turning inference behavior, traces, workload replay, GPU signals, and task-path evidence into reusable optimization...Full timeRemote work
$200k - $400k
...to an algorithm. We're building the infrastructure to understand human behavior at scale... ...possible to run. About the Role As a Member of Technical Staff in Research Infrastructure, you will... ...serving path to find where the FLOPs and GPU memory are going, and others still...Live inFlexible hours- ...observe their code. We are responsible for designing, building, and scaling core infrastructure that powers a high-volume data platform for AI applications. We are looking for team members who love building enabling systems that empower our engineers and power our rapidly...Work at office
- ...DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems.... ...latency (e.g. TensorRT) Understanding of distributed compute, GPU parallelism, and hardware-aware optimization Deep familiarity...
- About Us Gimlet is building the next generation of AI infrastructure: large-scale AI datacenters and the orchestration platform that coordinates... ...will work on Design, deploy, and operate large‑scale CPU, GPU, and accelerator clusters powering production AI inference....
$350k
...model training, reinforcement learning, reasoning systems, and infrastructure for large-scale experiments. Our team includes researchers and... ...Kubernetes and multi-cluster compute - operate CPU and GPU clusters as one platform with scheduling, autoscaling, and multi...$250k - $300k
...career? Join one of the most interesting infrastructure companies in the AI space right now.... ...they have built a platform that deploys GPU clusters into third-party datacentres at... ...Must Have: A track record of impressive technical work you can speak to in depth; the...Full timeRemote work$250k - $300k
...development? Join one of the most exciting AI infrastructure companies in the market, building a... ...that deploys and operates large-scale GPU clusters for some of the world's leading... ...Have: ~ A track record of impressive technical work you can speak to in depth, the years...Full timeRemote work- ...DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. We look for infrastructure engineers who are excited to tackle unsolved problems.... ...large-scale training fast, efficient, and reliable, so that every GPU cycle accelerates research progress. Responsibilities Design...
$150k - $250k
...people take ownership, grow together, and share both the challenges and the wins. What you'll do Build the supercomputing infrastructure that runs our agents. Our agents tackle long-horizon, high-performance workloads, and you'll design the cloud compute,...Work at officeRemote workFlexible hours- ...the life sciences. About the role We're looking for an Infrastructure Engineer to build, operate, and scale the deployment... ...HPC-adjacent workflows, bioinformatics tools, batch execution, GPU/CPU compute, and large-scale file/data movement. Work with...Work at office
- ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical... ...teams to focus on innovation, not on infrastructure. We aim to simplify the AI development... ...heterogeneous compute resources (CPU and GPU) efficiently? What data model will...Full timePart timeWork at officeWork from homeFlexible hours2 days per week
$150k - $300k
...key areas are: Building the infrastructure to serve LLMs efficiently at... ...our RL training stack. Core Technical Responsibilities LLM Serving... ...that operates across our cloud GPU fleets. GPU‑Aware Scheduling... ...development and encourage team members to contribute to the broader...Work at officeRemote workVisa sponsorshipRelocation packageFlexible hoursShift work- Job Description - Member of Technical Staff (Hardware) Location: San Francisco (on-site at our offices... ..., MatX, Tenstorrent or similar), GPU and accelerator teams at larger players... ...Google, Amazon, Qualcomm), or inference infrastructure companies working close to the...
$150k - $265k
...human again. Mission We're building the platform for the future of voice technology. Our market edge is extensible, reliable infrastructure designed for the full complexity of voice interactions. 18 months, 150k developers, adding 1000 every day. Give it a try here...Full timeShift work- .... Successful candidates typically come from staff or principal-level roles and are recognized for establishing technical direction, leading large-scale initiatives,... ...teams use to right‑size space and budgets. This infrastructure already powers 16,000 workplaces and 9,000+...Work at officeLocal areaMonday to Thursday
- ...curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet...
- About Mandolin Nearly every disease will become treatable in our lifetimes. Mandolin is laying the clinical and financial infrastructure to get groundbreaking treatments to patients faster, powered by AI agents. Mandolin partners closely with the largest healthcare institutions...Local area
- Member of the Technical Staff, Product (Backend)Location: North America Remote / San Francisco, CA · Full... ...to give startups the scaled AI infrastructure once reserved for hyperscalers. The... ...GPUs under management and billions of GPU-hours supported, on everything from...Remote work
- ...next Transformer or next AlphaFold breakthrough. As a Member of Technical Staff, you will build this autonomous AI research system to usher... ...evolutionary search algorithms, agent loops, parallel GPU infrastructure, and fine-tuning. What You’ll Work On Autonomous AI R&...
- ...Horowitz, GIC, Goldman Sachs, KKR, Visa, and others. Technical Skills Develop and maintain infrastructure that powers digital asset custody, trading, staking,... ...to solve problems, and assist or teach other team members when possible. You may be a fit for this role if you...Worldwide
$175k - $240k
...biology, physics, chemistry, and AI. The Role As a Member of Technical Staff, Infrastructure Engineer, you'll play a key role in designing, scaling,... ..., and fault tolerance for heterogeneous workloads (CPU, GPU, memory-intensive). Establish and uphold best practices...Full timeWork at office- ...lifecycle Implement the platform-level quality and monitoring tooling that data and research teams build their checks on Scale infrastructure to improve engineering velocity and ensure reliability, with monitoring and alerting to match Work across the full data...Immediate start
- ...superintelligence stack: the infrastructure frontier AI labs build internally... ...they own. Core Technical Responsibilities This hybrid... ...heterogeneous hardware (CPU, GPU, TPU) Backend & Feature Development... ...development and encourage team members to contribute to the broader...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff - GPU Infrastructure. Be the first to apply!
- work from home technical support specialist San Francisco, CA
- product support technician San Francisco, CA
- helpdesk support technician San Francisco, CA
- help desk assistant San Francisco, CA
- senior technical associate San Francisco, CA
- IT help desk technician San Francisco, CA
- technical solutions specialist San Francisco, CA
- desktop support analyst San Francisco, CA
- trade support analyst San Francisco, CA
- senior IT support technician San Francisco, CA


