Cluster Engineer
STN Inc
Role Description
We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.
Responsibilities
- Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
- Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
- Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
- Build and support production AI infrastructure running hundreds to thousands of GPUs.
- Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
- Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
- Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
- Configure and tune distributed AI software stacks including:
- PyTorch
- NCCL
- CUDA
- UCX
- MPI
- Slurm
- Pyxis/Enroot
- Optimize GPU scheduling and resource allocation for both training and inference environments.
- Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
- Identify performance regressions and troubleshoot distributed training issues at scale.
- Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
- Work closely with ML engineers to improve training scalability and inference efficiency.
- Create automation to deploy, validate, benchmark, and monitor GPU clusters.
- Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.
Qualifications
- 7+ years designing or operating large-scale Linux infrastructure.
- 5+ years supporting production GPU clusters for AI or HPC workloads.
- Demonstrated experience building multi-node GPU training environments from the ground up.
- Deep expertise with distributed PyTorch training.
- Extensive experience troubleshooting and optimizing NCCL communications.
- Strong understanding of distributed AI communication patterns, including:
- AllReduce
- ReduceScatter
- AllGather
- Broadcast
- Point-to-point communications
- Experience benchmarking distributed training using tools such as:
- nccl-tests
- NVIDIA DCGM
- Nsight Systems
- MLPerf (preferred)
- Strong understanding of GPU memory management, including:
- KV Cache
- Activation checkpointing
- Tensor Parallelism
- Pipeline Parallelism
- Data Parallelism
- Experience optimizing LLM inference throughput, including:
- Tokens/sec optimization
- Batch sizing
- Continuous batching
- KV cache tuning
- Memory bandwidth optimization
- Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
- Expert-level Linux systems administration skills.
- Experience with Slurm workload manager.
- Experience using Pyxis and Enroot for containerized GPU workloads.
- Strong scripting skills using Python and Bash.
Technical Expertise
- AI Frameworks:
- PyTorch
- CUDA
- NCCL
- Triton (preferred)
- TensorRT-LLM (preferred)
- Cluster Scheduling:
- Slurm
- Pyxis
- Enroot
- GPU Networking:
- Strong understanding of:
- InfiniBand
- RoCE v2
- RDMA
- GPUDirect RDMA
- GPUDirect Storage
- UCX
- MPI
- Network topology optimization
- Congestion control
- QoS
- ECN/PFC
- High-speed Ethernet (200/400/800 GbE)
- Strong understanding of:
- Storage:
- Experience designing or tuning storage for AI workloads, including:
- Parallel file systems
- Distributed storage
- Object storage
- NVMe
- Checkpoint optimization
- Dataset staging
- GPUDirect Storage
- Storage bandwidth optimization
- Metadata performance
- Experience designing or tuning storage for AI workloads, including:
- Performance Engineering:
- Experience with:
- NCCL benchmarking
- Multi-node scaling analysis
- GPU utilization optimization
- Communication/computation overlap
- NUMA optimization
- CPU affinity
- PCIe topology
- GPU topology (NVLink/NVSwitch)
- Memory bandwidth analysis
- End-to-end performance profiling
- Experience with:
Preferred Qualifications
- Experience deploying AI workloads on Kubernetes.
- Experience with NVIDIA GPU Operator.
- Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).
- Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
- Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
- Familiarity with MLPerf benchmarking.
- Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
- Experience automating infrastructure using Ansible, Terraform, or similar tools.
- Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.
$96.6k - $160k
...You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers... ...Remediation Project Design Electrical Engineer in across the IAD Cluster. If you meet these qualifications, exude passion, and enjoy the...SuggestedWork at officeImmediate startFlexible hours$131k - $175k
...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-Life... ...platforms integrate cleanly into large-scale AI and cloud clusters.What You’ll Do Lead the design, validation, and deployment of...SuggestedRemote workFlexible hours$15k
...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering...SuggestedWork at officeLocal areaRemote work$145.92k - $209.24k
...remote work a few days per week.Travel: Up to 15%, domestic and international. Job ID: 1750The Role: We're looking for an HPC Cluster Engineer to join our Infrastructure Team. Our mission is to build and operate the internal high-performance computing platform that IonQ...SuggestedPermanent employmentContract workWork at officeRemote work$250k
...The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering... ..., observability, and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering teams to...SuggestedFull timeRemote work- ...-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.THE ROLEGlobal Cluster Engineering (GCE) at AMD has a unique opportunity for a Cluster Architecture Documentation & Standards Engineer who thrives in a fast-paced...Remote workShift work
- DescriptionYour team’s dynamic:As Field Engineer, you will provide on-site and off-site professional services to Genetec customers and... ...operating systems (Active Directory, SQL, file sharing, IIS, clustering, GPO, performance monitoring, etc.) Excellent knowledge of networking...Full timeRemote workFlexible hours
$107.5k - $204.5k
...Raytheon is seeking an experienced Principal level Electrical Engineer to verify FPGA based designs for control and/or signal processing... ...with SLURM workload manager for job scheduling and compute cluster resource managementPrior experience mentoring junior verification...Temporary workWork experience placementWork at officeRemote workRelocationFlexible hours- ...performance, availability and reliability Collaborating with team engineers to design and implement servers, storage, containers, network... ...with Dell, VMware, NetApp, RHEL, Cloud, Containers, Clustering, LAN, WAN, Data Center, life cycle, and remote access technologies...Full timeWork experience placementLocal areaRemote work
$139.2k - $232.6k
...standards, and deliver with purpose. As a **Senior Electrical Engineer and licensed Professional Engineer specializing in telecommunications... ..., NVIDIA NVLink, and high-density fiber management for GPU cluster environments.FAT/SAT commissioning experience for mission-...Full timeFor contractorsFor subcontractorWork at officeRemote work- ...company where you matter.Your Impact As a Senior Security Operations Engineer II, you will play a key role in building secure, reliable, and... ...of Kubernetes platforms and container registries, including cluster provisioning, upgrades, hardening, policy enforcement, image...Work at officeRemote work
- ...or employee).The Impact you will have in this role:The Systems Engineering family is responsible for the entire technical effort to... ...Primary Responsibilities:Install, configure, and manage Kafka clusters, including Confluent Kafka components.Maintain and optimize existing...Remote workFlexible hours
$142.2k - $213.2k
...offers an excellent opportunity for a Sr Principal Cyber Systems Engineer - X-Lab Platform Engineer (26-375) to join our team of skilled... ...primary administrator for the container platforms, covering cluster build, upgrade, workload onboarding, and decommissioningAny task...Full timeContract workWork experience placementRemote workRelocationFlexible hoursShift work$119k - $170k
...cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (... ...availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks.Lead full-cycle incident response by...Full timeWork at officeLocal areaRemote work3 days per week- ...Principal Systems Engineer Job Description: About Foundever™ Foundever™ is a global leader in the customer experience (CX) industry... ..., DNS, Group Policy, PowerShell scripting, IIS, Windows clustering). Mandatory, in-depth hands-on experience with VMware vSphere...Remote work
$193k - $234k
...seeking a high-energy, detail-oriented Staff Network Production Engineer to lead the physical and logical implementation of our global... ...-in" testing and site acceptance testing (SAT) for new network clusters, ensuring zero-defect handovers to the Operations team.Optimize...Temporary workRemote work$170k - $180k
...Interests in Space as a Space Domain Awareness (SDA) Radar Systems Engineer providing Engineering Support as a member of the Deep Space... ...engineering, signal processing, agile software development, GPGPU cluster development or optimization, space systems operations and...Full timeContract workTemporary workFor contractorsWork experience placementCasual workLive outWork at officeRemote workFlexible hours2 days per week1 day per week$150k - $300k
Hudson River Trading (HRT) is looking for Systems Engineers to join our growing Research & Development team. This team builds and maintains exceptionally large and growing distributed compute clusters, multi petabyte-scale storage layers, operating systems, automation...Full timeWork at officeLocal areaImmediate startRemote workWorldwide$94.4k - $198.2k
...The Opportunity:CACI is seeking a talented Operational Systems Engineer to join our team in Chantilly, VA to help lead the implementation... ...testingExperience in managing large-scale remote workstation clusters, focusing on high-availability and minimal latency for highly scalable...Contract workWork experience placementRemote workFlexible hoursShift work- ...are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States... ...private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production...Work at officeRemote workFlexible hours
- ...Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview ~ We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner...Full timeLocal areaShift work
$100k - $160k
...Service Mesh Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud... ...mesh platforms — primarily Istio and Linkerd — across multi-cluster Kubernetes environments. Implement and operate mTLS, certificate...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...(Onsite 1-3 days/week - Tues., Wed., Thurs.)Organization: Global Instrument Panel Cluster & HUD EngineeringJob Type: Full-TimeRole OverviewThe HUD Mechanical & Packaging Senior Engineer owns end-to-end build to print design of mechanical architecture and packaging execution...Full timeLocal areaWork from homeFlexible hours3 days per week1 day per week
$82.9k - $146.17k
...systems, and other new and dynamic programs. In this role as a GNC engineer, you may be involved in many of the following activities: •... ...Implement simulation software on a high performance computing cluster including Linux systems or VMs • Work in an Agile DevOps...Full timeTemporary workWork experience placementFor subcontractorWork at officeRemote workRelocationFlexible hoursShift work- The Sr Systems Engineer Platform - Messaging Platform will play a key role in designing, implementing, and maintaining enterprise messaging... ...of configuration artifacts, rolling upgrades of messaging clusters, and dynamic provisioning of topics and consumer policies.Implement...Contract workLocal areaRemote workFlexible hours
$120k - $140k
...our success! Job Summary The Professional Service Systems Engineer works alongside technicians, technical support, other HPS... ...Ware or Microsoft Virtual Server. · Experience with Microsoft Cluster technology · Experience with government technical publications...Work experience placementLocal areaRemote work$90k - $176k
...performance, memory, I/O, configuration, security, networking, clustering, and storageProven track record supporting and troubleshooting... ...specialist within MongoDB and will be helping your peer engineers in advance diagnostics. Also, you will be encouraged to handle...Local areaWorldwideFlexible hours- What will you do?This position will oversee a multi-disciplined engineering team that provides engineered to order solutions for customer... ...you report to?The position reports to the Project Engineering Cluster Leader US-CANWhat qualifications will make you successful for...Work at officeFlexible hours
- ...Area Preferred) Position: Senior Service Delivery & Support Engineer Department: Public Sector Job Type: Full-time Job... ...platform experience. Experience with HPE server platforms, VMware clusters, pfSense, Fortinet, Cisco, or comparable infrastructure...Full timePart timeImmediate startRemote work
$50k
...in a truly global environment.About the roleAs a Senior Project Engineer Manager, you’ll own high-impact capital projects from idea to... ...reverse osmosis, clean-in-place systems, sanitary piping, valve clusters and more.The typical base salary hiring range for this role is...Full timeFor contractorsWork experience placementRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Cluster Engineer. Be the first to apply!



