Senior HPC Cluster Engineer
$152k - $241.5kNVIDIA
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA (Electronic Design Automation) and high-performance computing workloads used across multiple teams and projects. Join our engineering team and collaborate with researchers and infrastructure teams to ensure our GPU clusters are highly performant, scalable and reliable.What you'll be doing:Develop and enhance our ecosystem around GPU-accelerated computing including developing scalable automation solutions.Continuously improve infrastructure provisioning, management, observability and day to day operation through automation.Provide technical leadership and strategic guidance for managing large-scale HPC systems, including the deployment of compute, networking, and storage.Foster strong customer and multi-functional partnerships to ensure consistent cluster support and rapidly adapt to evolving user needsSupport our researchers to run their EDA workloads including performance analysis and optimizations.Conduct root cause analysis and suggest corrective action. Proactively find and fix issues before they occur.Build innovative tooling to accelerate researchers' velocity, debugging and software performance at scale.What we need to see:Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.Minimum of 5 years of proven experience crafting and operating large scale compute infrastructure, including cluster configuration managements tools such as BCM or Ansible.Experience with AI/HPC job schedulers and orchestrators, such as Slurm, LSF, PBS or K8s. Applied experience with AI/HPC workflows that use MPI and NCCL.Proficient in using Linux including Rocky/Centos/RHEL and/or Ubuntu Linux distributions. A solid understanding of container technologies such Enroot and Docker.Proficiency in Python and BashExperience analyzing and tuning performance for a variety of EDA workloads. Excellent problem-solving to analyze complex systems, identify bottlenecks, and implement scalable solutions.Excellent communication and collaboration skills, with the ability to work effectively with various teams and individuals.Passion for continual learning and staying ahead of new technologies and effective approaches in the HPC infrastructure fields.Ways to stand out from the crowd:Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarking.Experience supporting EDA workloads and tools.Familiarity with High-Speed Networking pertaining to HPC including InfiniBand, RDMA and RoCE.Understanding of fast, distributed storage systems such as Lustre and GPFS for AI/HPC workload.Familiarity with metrics collection and visualization at scale with Prometheus, OpenSearch and Grafana.Our technology has no boundaries! NVIDIA is building the most groundbreaking and powerful compute platforms for the world to use. It’s because of our work that scientists, researchers and engineers can advance their ideas. At its core, our visual computing technology not only enables an amazing computing experience, but it is also energy efficient! We pioneered a supercharged form of computing loved by the most demanding computer users in the world - scientists, designers, artists, and gamers. It’s not just technology though! It is our people, some of the brightest in the world, and our diverse company culture make NVIDIA one of the most fun, innovative and dynamic places to work in the world! At the center of NVIDIA's culture are our core values like innovation, excellence and determination and team, that guide us to be the best we can be.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 19, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, WA, RedmondType: Full time
$145.92k - $209.24k
...days per week. Travel: Up to 15%, domestic and international. Job ID: 1750 The Role: We're looking for an HPC Cluster Engineer to join our Infrastructure Team. Our mission is to build and operate the internal high-performance computing platform that...SeniorPermanent employmentFull timeContract workWork at officeRemote work$184k - $287.5k
...CPUs, and a fully optimized NVIDIA AI and HPC software stack. We’re searching for a... ...emulation technology.Mentor architects and engineering teams to grow them into future leaders.Make... ...crowd:Knowledge of large-scale cloud and cluster level deployment and management systems....SeniorFull timeRemote workShift work$119.8k - $234.7k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft... ...(PSE) team is seeking a Senior AI Network Systems Engineer... ...scale AI training and inference clusters. Collaborate with silicon,... ...of experience supporting AI, HPC, cloud, or large-scale data center...SeniorOngoing contractWork at officeLocal areaWorldwide3 days per week- ...ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER — HPC & AUTOMATION (SILICON ENGINEERING)At SpaceX we’re leveraging... ...RESPONSIBILITIES:Deploy, upgrade, operate, maintain, and scale our suite of clusters and servicesCollaborate with engineers to develop automated,...SuggestedPermanent employmentTemporary workWork at officeWorldwideMonday to FridayWeekend work
$152k - $241.5k
...technologies to accelerate AI workloads, and we are looking for an engineer focused on performance validation, analysis, and tracking. In... ...systems to measure latency, throughput, and efficiency of AI and HPC workloadsAnalyze performance trends over time and identify...SeniorFull timeRemote work$108k - $172.5k
...lasting impact on the world.We are seeking a highly motivated Senior HPC Support Engineer focussing on InfiniBand and NVLink technology, passionate... ...the following:InfiniBand, RDMA/RoCEv2 and GPU Technology.Clustering or HPC Data-Center technologies including Upper Layer...SeniorFull timeWork experience placement- A global technology company is seeking a Senior Engineer to optimize core data infrastructure and ensure high-performance computing within their advertising platform. In this role, you will drive innovations and engage in all software development lifecycle aspects while...Senior
$184k - $287.5k
We are seeking for an expert Senior Compiler Engineer to join our Compute Compiler Team, with a focus on upstream engagement with the LLVM ecosystem and advancing NVIDIA’s compiler technology through trusted, sustained participation in open‑source communities. In this...SeniorFull timeRemote work$153k - $204k
...You'll Do: The Systems Engineering team owns the host software stack... ...that framework into HPC verification, Slurm-on-Kubernetes... ...About the role: As a Senior Software Engineer on the Systems... ...experience. HPC or large-cluster experience — InfiniBand/RoCE,...SeniorPermanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$159.2k - $215.3k
...latency, high-speed broadband connectivity to unserved and underserved communities around the world.A day in the life:As an Senior Hardware Engineer, you will have the opportunity to work on some of the most innovative, challenging, and exciting phased array antenna...SeniorPermanent employmentFlexible hours$187.36k - $245.3k
...already use IonQ quantum computers in tandem with high-performance computer (HPC) clusters in applications like quantum machine learning and image analysis. As a Senior Staff Software Engineer leading our HPC integration, you’ll help build and maintain our interfaces with...SeniorPermanent employmentFull timeContract workWork at office$165k - $242k
...'ll Do CoreWeave is seeking a highly skilled and motivated HPC Performance Engineer to join our HAVOCK Team, reporting into the Manager of Systems... ...bare‑metal systems from POST through joining a Kubernetes cluster. The team's primary responsibilities include maintaining a...Permanent employmentTemporary workCasual workWork at officeRemote workFlexible hours$140k - $200k
...shaping the future of global communications. For more information, visit kymetacorp.com What We Need Kymeta is looking for a Senior RF Antenna Engineer who will use metamaterial technology to provide next generation satellite communications services using flat, thinner,...SeniorWorldwideFlexible hours$188k - $275k
...more at What You'll Do: The Field Engineering organization at CoreWeave is dedicated to... ...entire customer lifecycle: leading new GPU cluster bring-up and acceptance, driving InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we operate...SeniorPermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$207k - $275k
...production environment is mission-critical. We’re seeking a Senior Manager of Production Engineering to lead and expand our SRE team and practices. This... ...(e.g., custom data centers, edge compute, HPC clusters). Working knowledge of DPUs, service mesh architectures...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$119.8k - $234.7k
...Hardware EngineeringDiscipline: Sourcing EngineeringCompany: MicrosoftOverviewJoin Microsoft in this amazing opportunity as a Senior Sourcing Engineer driving end‑to‑end sourcing strategies across our Global Tier1 Contract Manufacturing. This role enables impact through...SeniorOngoing contractContract workLocal area3 days per week$153k - $204k
...CRWV) in March 2025. Learn more at What You’ll Do As a Senior Engineer in Compute Services, you will be responsible for building... ...Perform day 2 lifecycle tasks and maintenance on running clusters Identify gaps and implement fault-tolerant architectures...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$119.8k - $234.7k
...MicrosoftOverviewMicrosoft Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s expanding Cloud... ...and optimize the Cloud infrastructure.We are looking for a Senior Silicon Engineer to join the team.#SCHIEResponsibilitiesLead the...SeniorOngoing contractPermanent employmentWork at officeLocal areaWorldwide3 days per week- NVIDIA Gruppe is seeking a Senior Software Engineer for Quantized Inference in Redmond, Washington. You will be responsible for implementing quantized and sparse recipes in inference engines and enhancing developer productivity across the team. Successful candidates will...Senior
- HCLTech in Redmond, WA is seeking a highly talented PCB Layout Engineer to design and develop complex multilayer PCBs for next-generation computing, networking, wireless, and embedded hardware platforms. You will use Cadence Allegro, handle high-speed interfaces, and drive...Senior
- Amazon Leo in Redmond, WA is seeking a senior software engineer to join the SDN team building the control plane for a global satellite-terrestrial network. You will design, implement, and operate scalable networking software that powers high throughput, low latency connectivity...Senior
$165k - $242k
...AI cloud platform provider in Bellevue, WA, is seeking an HPC Performance Engineer to optimize bare-metal systems. This role involves collaborating... ...Linux kernel, and ensuring performance across distributed clusters. Ideal candidates will have over five years of experience...$184k - $287.5k
...perceive and understand the world. Today, we are increasingly known as “the AI computing company”.We're looking for a Senior Performance Compiler Engineer to join our team and work on the open-source Triton compiler project. This opportunity involves working with new...SeniorFull timeRemote work$152k - $241.5k
...perceive and understand the world. Today, we are increasingly known as “the AI computing company”.We are looking for versatile software engineers for our XLA team. NVIDIA is at the center for the AI revolution that's transforming how people live, work, and interact with...SeniorFull timeRemote work$184k - $287.5k
...inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.We are seeking an AI Compiler Engineer with deep expertise in compiler technologies to join our team. The ideal candidate brings broad experience across machine learning...SeniorFull timeRemote work$119.8k - $234.7k
...EngineeringCompany: MicrosoftOverviewAre you looking for a product group/engineering role where you will impact the product roadmap while working... ...a high degree of operational and execution excellence.As a Senior Customer Experience Engineer in the Support-as-a-Feature Team,...SeniorOngoing contractWork at officeLocal area3 days per week$159.2k - $215.3k
...businesses, government agencies, and other organizations operating in places without reliable connectivity.As an Antenna Systems Engineer supporting the development, test, and design of antenna and RF Front-End, you will design, integrate, and characterize antennas and...SeniorPermanent employmentFlexible hours$137.3k - $185.7k
...locations without reliable connectivity.Amazon Leo Customer Terminal Engineering is looking for an individual who will lead our Aviation... ...certification regulations on a first-generation product. As a Senior Certification Aviation Engineer, you will engage with an experienced...SeniorPermanent employmentWork experience placementFlexible hours$152k - $241.5k
...their best work. Come join the team and see how you can make a lasting impact on the worldWe are looking for a Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning Compiler (DLC) team. Academic and commercial groups around the world...SeniorFull timeRemote work$119.8k - $234.7k
...leading progress in areas ranging from quantum hardware and error correction to comprehensive integration with Azure. As a Senior Quantum Engineer you will design and test the readout system for topological qubits at the component, subsystem and system levels. This opportunity...SeniorOngoing contractPermanent employmentLocal area3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior HPC Cluster Engineer. Be the first to apply!
- senior associate architect Redmond, WA
- senior dynamics crm developer Redmond, WA
- senior application security Redmond, WA
- senior supervisor Redmond, WA
- senior cloud data engineer Redmond, WA
- senior customer service manager Redmond, WA
- senior Redmond, WA
- senior customer success engineer Redmond, WA
- senior business development Redmond, WA
- senior operations technician Redmond, WA

