Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
$250kFull-time
Perplexity
Role Description
Your job is to take ownership of that infrastructure and hide its complexity behind a unified, self-serve platform for running training and inference workloads.
- Build a self-serve compute platform.
- Design and own the systems that let inference engineers and researchers launch training jobs and operate inference services without managing GPU provisioning, cluster configuration, or provider-specific infrastructure.
- Operate the GPU fleet:
- Own provisioning, lifecycle management, reliability, and capacity integration across providers, giving teams a consistent way to use compute regardless of where it runs.
- Solve for GPU scarcity:
- Build the scheduling and placement logic that finds available capacity across providers, packs it efficiently, and gets the right workload onto the right hardware under real constraints.
- Support two very different workloads:
- Keep long-running distributed training jobs healthy while simultaneously guaranteeing the availability and latency of production inference services on the same fleet.
- Own the Kubernetes for GPU orchestration:
- Write the operators and CRDs, and manage many clusters across providers so the platform behaves the same everywhere we run.
- Make failure boring:
- Build the fault tolerance, autoscaling, and observability that keep the fleet utilized and let workloads survive node loss, provider hiccups, and capacity shifts without human intervention.
- Set technical direction across teams:
- Partner with inference and cloud infrastructure engineers to turn operational constraints into a coherent platform architecture and roadmap.
Qualifications
- Deep Kubernetes experience — custom operators, CRDs, and multi-cluster federation, not just running kubectl apply.
- You've managed GPU clusters at scale: NVIDIA hardware, CUDA, and the networking that makes them fast (InfiniBand or RoCE).
- You've orchestrated compute across multiple clouds (CoreWeave, AWS, GCP, or similar) and understand how different each one really is.
- Strong distributed systems fundamentals: scheduling, resource allocation, and fault tolerance under load.
- You write infrastructure and systems-level code in Go, Rust or C++.
- You've supported both long-running training jobs and high-availability inference services, and you know why they pull infrastructure in opposite directions.
- You own problems end-to-end and do well when the path forward isn't laid out for you.
Requirements
- Inference serving stacks: vLLM, SGLang, or TensorRT-LLM.
- Slurm or other HPC schedulers.
- GPU kernel work in CUDA or Triton — not required, but notable.
- High-speed interconnects: InfiniBand, RoCE, or RDMA in production.
- Observability for ML workloads: Prometheus, Grafana, or Weights & Biases.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure) in Remote vacancy
- ...a team of researchers, engineers, designers, and more, who... ...? We are looking for Members of Technical Staff to join the Model Serving... ...running production infrastructure at a large scale ~ Experience... ...with Kubernetes, and GPU workloads on those clusters ~ Experience with...SuggestedFull timeWork experience placementWork at officeRemote workFlexible hours
- ...to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI... ...migration, inference autoscaling and multi-cluster serving ~Preemption handling and... ...and with ML training or inference infrastructure ~Strong Python, and comfort reaching...SuggestedFull time
- ...Infrastructure Platform Engineer We are looking for an Infrastructure platform Engineer to design, build, and operate the cluster infrastructure behind Gimlet's heterogeneous inference cloud. Unlike... ..., and operate large-scale CPU, GPU, and accelerator clusters powering...Suggested
- ...leader in AI cloud infrastructure serving tens of... ...One person, one GPU.If you'd like to... ...Infrastructure Engineering organization... ...-performance AI clusters by welding together... ...seeking a seasoned Staff Storage Software Engineer with... ...Leadership: Set technical direction for storage...SuggestedWork at officeLocal areaWork from homeFlexible hours
- ...best data and AI infrastructure platform so our customers... .... Founded by engineers — and customer... ...to solve technical challenges, from... ...thousands of Kubernetes clusters, and must deliver... ...efficiency. As a Staff Software Engineer and Tech... ...batch, stateful, GPU) with high...SuggestedFull time
$188k - $275k
Role Description As a Staff Engineer on Marimo's molab... ...specialized kubernetes-based clusters and integrate with... ...of experience in software engineering ~Strong... ...Distributed systems ~Cloud infrastructure ~Experience... ...~Experience with GPU resource allocation and...Full timeTemporary workCasual workWork at officeFlexible hours$198k - $326k
...opportunity for every member of the global... ...team. LinkedIn’s AI Infrastructure organization is responsible... ...looking for a Senior Staff Software Engineer with deep expertise at... ...systems, machine learning, GPU infrastructure, and... .... This is a highly technical, high-leverage role...For contractorsWork at officeFlexible hours$200k - $400k
Role Description We're looking for a TPU and AMD GPU performance engineer to make vLLM a first-class inference engine across non-NVIDIA accelerators... ..., compiler integrations, runtime paths, and benchmarking infrastructure. ~Work at the boundary of inference systems, kernels,...Full timeRemote workVisa sponsorship$300k
...looking for a deeply technical Member of Technical Staff to own RL and... ...and engineers who’ve operated... ...and build the infrastructure needed to run them... ...serving latency, GPU utilization, policy... .... Strong software engineering fundamentals... ...large GPU clusters. Experience with...H1bWork at officeVisa sponsorshipShift work$200k - $400k
Role Description We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem... ...GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling...Full time$109k - $160k
...enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate... ...more at . About the role A Software Engineer contributes to the design, implementation... ...hardware teams to evolve our GPU performance testing platform to...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours$146.6k - $215.1k
...the RoleWe're looking for a Staff Software Engineer to join our Cloud ML team —... ...both the cloud-side ML infrastructure and the applied ML research... ...across the team, and set the technical direction for our ML platform... ...low-latency, multi-tenant, GPU-aware, and unforgiving of...Work at office$207k - $300k
...design and end-to-end software delivery of... ...closely with hardware engineering and chip design teams... ...developing large-scale infrastructure, distributed systems... ...Science, or a related technical field.8 years of experience... ...and Google-internal cluster systems, alongside a...Remote workWorldwide$197.3k - $313.7k
...Salesforce.Slack is looking for a Staff Software Engineer to join our Desktop team... ...and frontend stakeholders.Technical Standards: You will be... ...frontend and desktop infrastructure to support new product features... ...between junior and senior members to surface blockers....Full timeRemote work$183.37k - $214.5k
...here.Role Overview:We are seeking a Staff Software Engineer for the ID.me Member Support Applications Team to lead... ...& Agent Design: Lead the technical vision for integrating LLMs (specifically... ...and member experience.Shared Infrastructure: Develop shared UI components and...Full timeTemporary workWork at officeRemote workWorldwideFlexible hours$198k - $299k
...query. Our in-house OLAP engine, Nova, processes... ...makes Nova the critical infrastructure in the loop, and as... ....We’re looking for a Staff Software Engineer who wants to... ...scale. You’ll influence technical direction through... ...compression, caching, and cluster-level resource...Work at officeImmediate startWorldwideHome office$250k - $300k
...one of the most exciting AI infrastructure companies in the market,... ...deploys and operates large-scale GPU clusters for some of the world's... ...platform evolves while solving engineering challenges that have a... ...track record of impressive technical work you can speak to in depth...Permanent employmentRemote work$250k - $300k
...the most interesting infrastructure companies in the AI space... ...that deploys GPU clusters into third-party datacentres... ...0. As part of the engineering team, you'll help... ...containers, and PCIe, writing software that has to hold up... ...record of impressive technical work you can speak to...Permanent employmentRemote work$238k - $302k
...techniques and build the infrastructure to store, process,... ...map data. As a member of the infrastructure... ...improvements of our software stack, own resource planning... ...a team of software engineers to help us scale our... ...: ~ Experienced technical leader (or as a people...Full timeRemote work$193.93k - $352.29k
...About the Role Our software team is growing, and we are looking for talented engineers to join us and be instrumental... ...Data Platform, Simulation, and Technical Infrastructure. Data Platform: The... ...different compute modalities (CPU, GPU, FPGA) etc. At Nuro,...Full time$170k - $240k
...our office. About the Role We’re looking for early members of our software engineering infrastructure team. You’ll work closely with the founding team and have ownership of a wide variety of technical and design decisions for Suno’s technical architecture and products...Full timeWork at officeFlexible hours$185k - $245k
Role Description Cribl Inc is seeking a Staff Software Engineer to join our mission to unlock the... ...ship Cribl products. As an active member of our team, you will: ~Contribute... ...solutions that improve our cloud service, infrastructure, and tools. ~Solve infrastructure...Full timeTemporary workRemote work$175k - $287k
...economic opportunity for every member of the global workforce. Our... ...part of our world-class software engineering team, you will help build the next-generation infrastructure and platforms that power LinkedIn... ...:You will own the technical strategy for broad or complex...For contractorsWork experience placementWork at officeFlexible hours$120k - $160k
...Application Security platform for the software development revolution. Modern software... ...servers, ensuring robust and scalable infrastructure. Maintain diverse scan environments... ...users. What We're Looking For Engineering Expertise : Bachelor’s in engineering...Full timeShift work- ...gaps in patient care, drive member enrollment, and patient... ...growth without hiring more staff.We are on a mission to improve... ...together.Role Summary:As an Infrastructure Engineer, you will have the... ...If you are experiencing a technical issue with your application...Full timeTemporary workWork at officeRemote work3 days per week
$140k - $165k
...electronic devices and IT infrastructure, enabling enhanced... ...a hands-on AI Engineer to design, deploy, and... ...infrastructure — including GPU clusters, model serving (e.g.,... ...you are motivated by technical challenges, we offer... ...our current team members to ensure fairness and...Full time$226k - $369k
...economic opportunity for every member of the global workforce.... ...are seeking a Principal Staff Software Engineer to join our organization.... ...you will serve as a senior technical leader and architect... ...and driving next-generation infrastructure that powers AI-first unified...For contractorsWork at officeRemote workWork from homeFlexible hours- Role Description The Staff Software Engineer – Cloud Infrastructure Engineer is responsible for all stages of the software... ...to define and prioritize technical requirements that meet client needs... ...them design, deploy, and operate clusters and workloads. ~Build custom tooling...Full timeWork experience placement
- ...AI Infrastructure Specialist As vCluster... ...from bare metal GPU nodes through... ...first team members a neocloud or... ...engages with at a technical depth, and the... ...Engineer, your role will... ...40M+ virtual clusters created since... ...don't just ship software; we define the...Remote workFlexible hours
- ...Description We are looking for engineers who can reason about... ...operationalizes our infrastructure ~Money paths:... ...Qualifications ~Impressive technical work you can go deep... .... The first cluster filled almost instantly... ...management ~Billions of GPU-hours supported, on everything...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure). Be the first to apply!
Related searches
- technical assistant Remote
- work from home technical support specialist Remote
- end user support technician Remote
- help desk technical support Remote
- technical support assistant Remote
- tech assistant Remote
- remote support technician Remote
- application support technician Remote
- technical associate Remote
- technical analyst Remote








