Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Cluster Engineer

Full-time

STN Inc

Role Description

We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance.

Responsibilities

  • Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads.
  • Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency.
  • Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization.
  • Build and support production AI infrastructure running hundreds to thousands of GPUs.
  • Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers.
  • Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance.
  • Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS.
  • Configure and tune distributed AI software stacks including:
    • PyTorch
    • NCCL
    • CUDA
    • UCX
    • MPI
    • Slurm
    • Pyxis/Enroot
  • Optimize GPU scheduling and resource allocation for both training and inference environments.
  • Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases.
  • Identify performance regressions and troubleshoot distributed training issues at scale.
  • Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O.
  • Work closely with ML engineers to improve training scalability and inference efficiency.
  • Create automation to deploy, validate, benchmark, and monitor GPU clusters.
  • Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture.

Qualifications

  • 7+ years designing or operating large-scale Linux infrastructure.
  • 5+ years supporting production GPU clusters for AI or HPC workloads.
  • Demonstrated experience building multi-node GPU training environments from the ground up.
  • Deep expertise with distributed PyTorch training.
  • Extensive experience troubleshooting and optimizing NCCL communications.
  • Strong understanding of distributed AI communication patterns, including:
    • AllReduce
    • ReduceScatter
    • AllGather
    • Broadcast
    • Point-to-point communications
  • Experience benchmarking distributed training using tools such as:
    • nccl-tests
    • NVIDIA DCGM
    • Nsight Systems
    • MLPerf (preferred)
  • Strong understanding of GPU memory management, including:
    • KV Cache
    • Activation checkpointing
    • Tensor Parallelism
    • Pipeline Parallelism
    • Data Parallelism
  • Experience optimizing LLM inference throughput, including:
    • Tokens/sec optimization
    • Batch sizing
    • Continuous batching
    • KV cache tuning
    • Memory bandwidth optimization
  • Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance.
  • Expert-level Linux systems administration skills.
  • Experience with Slurm workload manager.
  • Experience using Pyxis and Enroot for containerized GPU workloads.
  • Strong scripting skills using Python and Bash.

Technical Expertise

  • AI Frameworks:
    • PyTorch
    • CUDA
    • NCCL
    • Triton (preferred)
    • TensorRT-LLM (preferred)
  • Cluster Scheduling:
    • Slurm
    • Pyxis
    • Enroot
  • GPU Networking:
    • Strong understanding of:
      • InfiniBand
      • RoCE v2
      • RDMA
      • GPUDirect RDMA
      • GPUDirect Storage
      • UCX
      • MPI
      • Network topology optimization
      • Congestion control
      • QoS
      • ECN/PFC
      • High-speed Ethernet (200/400/800 GbE)
  • Storage:
    • Experience designing or tuning storage for AI workloads, including:
      • Parallel file systems
      • Distributed storage
      • Object storage
      • NVMe
      • Checkpoint optimization
      • Dataset staging
      • GPUDirect Storage
      • Storage bandwidth optimization
      • Metadata performance
  • Performance Engineering:
    • Experience with:
      • NCCL benchmarking
      • Multi-node scaling analysis
      • GPU utilization optimization
      • Communication/computation overlap
      • NUMA optimization
      • CPU affinity
      • PCIe topology
      • GPU topology (NVLink/NVSwitch)
      • Memory bandwidth analysis
      • End-to-end performance profiling

Preferred Qualifications

  • Experience deploying AI workloads on Kubernetes.
  • Experience with NVIDIA GPU Operator.
  • Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.).
  • Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang.
  • Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments.
  • Familiarity with MLPerf benchmarking.
  • Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter.
  • Experience automating infrastructure using Ansible, Terraform, or similar tools.
  • Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.
Vacancy posted 10 days ago
Similar jobs that could be interesting for youBased on the Cluster Engineer in Remote vacancy
  • $96.6k - $160k

     ...You’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers...  ...Remediation Project Design Electrical Engineer in across the IAD Cluster. If you meet these qualifications, exude passion, and enjoy the... 
    Suggested
    Work at office
    Immediate start
    Flexible hours

    Amazon Data Services, Inc.

    Herndon, VA
    5 hours ago
  • $131k - $175k

     ...excellence has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-Life...  ...platforms integrate cleanly into large-scale AI and cloud clusters.What You’ll Do Lead the design, validation, and deployment of... 
    Suggested
    Remote work
    Flexible hours

    Arista Networks

    Santa Clara, CA
    3 days ago
  • $15k

     ...benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering... 
    Suggested
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    4 days ago
  • $145.92k - $209.24k

     ...remote work a few days per week.Travel: Up to 15%, domestic and international. Job ID: 1750The Role: We're looking for an HPC Cluster Engineer to join our Infrastructure Team. Our mission is to build and operate the internal high-performance computing platform that IonQ... 
    Suggested
    Permanent employment
    Contract work
    Work at office
    Remote work

    IONQ

    College Park, MD
    2 days ago
  • $250k

     ...The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering...  ..., observability, and monitoring frameworks for GPU compute clusters Collaborate with ML, data, and platform engineering teams to... 
    Suggested
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.THE ROLEGlobal Cluster Engineering (GCE) at AMD has a unique opportunity for a Cluster Architecture Documentation & Standards Engineer who thrives in a fast-paced... 
    Remote work
    Shift work

    AMD

    Texas
    3 days ago
  • DescriptionYour team’s dynamic:As Field Engineer, you will provide on-site and off-site professional services to Genetec customers and...  ...operating systems (Active Directory, SQL, file sharing, IIS, clustering, GPO, performance monitoring, etc.) Excellent knowledge of networking... 
    Full time
    Remote work
    Flexible hours

    Genetec

    Washington DC
    1 day ago
  • $107.5k - $204.5k

     ...Raytheon is seeking an experienced Principal level Electrical Engineer to verify FPGA based designs for control and/or signal processing...  ...with SLURM workload manager for job scheduling and compute cluster resource managementPrior experience mentoring junior verification... 
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Relocation
    Flexible hours

    Raytheon

    Marlborough, MA
    2 days ago
  •  ...performance, availability and reliability Collaborating with team engineers to design and implement servers, storage, containers, network...  ...with Dell, VMware, NetApp, RHEL, Cloud, Containers, Clustering, LAN, WAN, Data Center, life cycle, and remote access technologies... 
    Full time
    Work experience placement
    Local area
    Remote work

    WolfSpeed

    Durham, NC
    1 day ago
  • $139.2k - $232.6k

     ...standards, and deliver with purpose. As a **Senior Electrical Engineer and licensed Professional Engineer specializing in telecommunications...  ..., NVIDIA NVLink, and high-density fiber management for GPU cluster environments.FAT/SAT commissioning experience for mission-... 
    Full time
    For contractors
    For subcontractor
    Work at office
    Remote work

    Jacobs

    Portland, OR
    7 hours ago
  •  ...company where you matter.Your Impact As a Senior Security Operations Engineer II, you will play a key role in building secure, reliable, and...  ...of Kubernetes platforms and container registries, including cluster provisioning, upgrades, hardening, policy enforcement, image... 
    Work at office
    Remote work

    Axon

    Scottsdale, AZ
    11 hours ago
  •  ...or employee).The Impact you will have in this role:The Systems Engineering family is responsible for the entire technical effort to...  ...Primary Responsibilities:Install, configure, and manage Kafka clusters, including Confluent Kafka components.Maintain and optimize existing... 
    Remote work
    Flexible hours

    Dtcc

    Coppell, TX
    1 day ago
  • $142.2k - $213.2k

     ...offers an excellent opportunity for a Sr Principal Cyber Systems Engineer - X-Lab Platform Engineer (26-375) to join our team of skilled...  ...primary administrator for the container platforms, covering cluster build, upgrade, workload onboarding, and decommissioningAny task... 
    Full time
    Contract work
    Work experience placement
    Remote work
    Relocation
    Flexible hours
    Shift work

    Northrop Grumman

    Colorado
    11 hours ago
  • $119k - $170k

     ...cybersecurity.RoleWe are looking for a Staff Site Reliability Engineer (Production Engineer) to join our team. This is a hybrid role (...  ...availability across large-scale bare-metal Linux/BSD fleets, Kubernetes clusters, and custom routing stacks.Lead full-cycle incident response by... 
    Full time
    Work at office
    Local area
    Remote work
    3 days per week

    Zscaler

    Dallas, TX
    1 day ago
  •  ...Principal Systems Engineer Job Description: About Foundever™ Foundever™ is a global leader in the customer experience (CX) industry...  ..., DNS, Group Policy, PowerShell scripting, IIS, Windows clustering). Mandatory, in-depth hands-on experience with VMware vSphere... 
    Remote work

    Foundever

    United States
    3 days ago
  • $193k - $234k

     ...seeking a high-energy, detail-oriented Staff Network Production Engineer to lead the physical and logical implementation of our global...  ...-in" testing and site acceptance testing (SAT) for new network clusters, ensuring zero-defect handovers to the Operations team.Optimize... 
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    1 day ago
  • $170k - $180k

     ...Interests in Space as a Space Domain Awareness (SDA) Radar Systems Engineer providing Engineering Support as a member of the Deep Space...  ...engineering, signal processing, agile software development, GPGPU cluster development or optimization, space systems operations and... 
    Full time
    Contract work
    Temporary work
    For contractors
    Work experience placement
    Casual work
    Live out
    Work at office
    Remote work
    Flexible hours
    2 days per week
    1 day per week

    Odyssey Systems Consulting Group

    Colorado Springs, CO
    11 hours ago
  • $150k - $300k

    Hudson River Trading (HRT) is looking for Systems Engineers to join our growing Research & Development team. This team builds and maintains exceptionally large and growing distributed compute clusters, multi petabyte-scale storage layers, operating systems, automation... 
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Worldwide

    Hudson River Trading

    Austin, TX
    4 days ago
  • $94.4k - $198.2k

     ...The Opportunity:CACI is seeking a talented Operational Systems Engineer to join our team in Chantilly, VA to help lead the implementation...  ...testingExperience in managing large-scale remote workstation clusters, focusing on high-availability and minimal latency for highly scalable... 
    Contract work
    Work experience placement
    Remote work
    Flexible hours
    Shift work

    CACI International

    Chantilly, Loudoun County, VA
    4 days ago
  •  ...are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States...  ...private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production... 
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Charlotte, NC
    3 days ago
  •  ...Norway, Bhutan, and Ethiopia. To learn more, visit Position Overview ~ We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class, productized scheduling layer across that fleet. This person is the single technical owner... 
    Full time
    Local area
    Shift work

    Bitdeer Technologies Group

    Remote
    4 days ago
  • $100k - $160k

     ...Service Mesh Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud...  ...mesh platforms — primarily Istio and Linkerd — across multi-cluster Kubernetes environments. Implement and operate mTLS, certificate... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    Framingham, MA
    3 days ago
  •  ...(Onsite 1-3 days/week - Tues., Wed., Thurs.)Organization: Global Instrument Panel Cluster & HUD EngineeringJob Type: Full-TimeRole OverviewThe HUD Mechanical & Packaging Senior Engineer owns end-to-end build to print design of mechanical architecture and packaging execution... 
    Full time
    Local area
    Work from home
    Flexible hours
    3 days per week
    1 day per week

    General Motors

    Warren, MI
    11 hours ago
  • $82.9k - $146.17k

     ...systems, and other new and dynamic programs. In this role as a GNC engineer, you may be involved in many of the following activities: •...  ...Implement simulation software on a high performance computing cluster including Linux systems or VMs • Work in an Agile DevOps... 
    Full time
    Temporary work
    Work experience placement
    For subcontractor
    Work at office
    Remote work
    Relocation
    Flexible hours
    Shift work

    Lockheed Martin Corporation

    Littleton, CO
    3 days ago
  • The Sr Systems Engineer Platform - Messaging Platform will play a key role in designing, implementing, and maintaining enterprise messaging...  ...of configuration artifacts, rolling upgrades of messaging clusters, and dynamic provisioning of topics and consumer policies.Implement... 
    Contract work
    Local area
    Remote work
    Flexible hours

    O'Reilly Auto Parts

    Springfield, MO
    2 days ago
  • $120k - $140k

     ...our success! Job Summary The Professional Service Systems Engineer works alongside technicians, technical support, other HPS...  ...Ware or Microsoft Virtual Server. · Experience with Microsoft Cluster technology · Experience with government technical publications... 
    Work experience placement
    Local area
    Remote work

    Hirsch

    Santa Ana, CA
    4 days ago
  • $90k - $176k

     ...performance, memory, I/O, configuration, security, networking, clustering, and storageProven track record supporting and troubleshooting...  ...specialist within MongoDB and will be helping your peer engineers in advance diagnostics. Also, you will be encouraged to handle... 
    Local area
    Worldwide
    Flexible hours

    MongoDB

    Palo Alto, CA
    3 days ago
  • What will you do?This position will oversee a multi-disciplined engineering team that provides engineered to order solutions for customer...  ...you report to?The position reports to the Project Engineering Cluster Leader US-CANWhat qualifications will make you successful for... 
    Work at office
    Flexible hours

    Schneider Electric

    Columbia, SC
    1 day ago
  •  ...Area Preferred) Position:  Senior Service Delivery & Support Engineer Department: Public Sector Job Type: Full-time Job...  ...platform experience. Experience with HPE server platforms, VMware clusters, pfSense, Fortinet, Cisco, or comparable infrastructure... 
    Full time
    Part time
    Immediate start
    Remote work

    Assured Data Protection

    Herndon, VA
    2 days ago
  • $50k

     ...in a truly global environment.About the roleAs a Senior Project Engineer Manager, you’ll own high-impact capital projects from idea to...  ...reverse osmosis, clean-in-place systems, sanitary piping, valve clusters and more.The typical base salary hiring range for this role is... 
    Full time
    For contractors
    Work experience placement
    Remote work

    Kerry Group

    Beloit, WI
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Cluster Engineer. Be the first to apply!