Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC AI Cluster Engineer

$176k - $333.5k
Full-time

NVIDIA

NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers and AI clusters based on groundbreaking technologies. We are looking for an outstanding engineer, be a key player to the most exciting computing hardware and software to contribute to the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs. You will work with the latest Accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.

What you will be doing:

  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting

  • Manage Linux job/workload schedules and orchestration tools

  • Develop and maintain continuous integration and delivery pipelines

  • Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources

  • Deploy monitoring solutions for the servers, network and storage

  • Perform troubleshooting bottom up from bare metal, operating system, software stack and application level

  • Being a technical resource, develop, re-define and document standard methodologies to share with internal teams

  • Support Research & Development activities and engage in POCs/POVs for future improvements

What we need to see:

  • A degree in Computer Science, Engineering, or a related field (or equivalent experience) and 8+ years of experience

  • Knowledge of HPC and AI solution technologies from CPU’s and GPU’s to high speed interconnects and supporting software

  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s

  • Excellent knowledge of Windows and Linux (Redhat/CentOS and Ubuntu) networking (sockets, firewalld, iptables, wireshark, etc.) and internals, ACLs and OS level security protection and common protocols e.g. TCP, DHCP, DNS, etc.

  • Experience with multiple storage solutions such as Lustre, GPFS, Weka.io. Familiarity with newer and emerging storage technologies.

  • Python programming and bash scripting experience.

  • Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/chef

  • Deep knowledge of Networking Protocols like InfiniBand, Ethernet

  • Deep understanding and experience with virtual systems (for example VMware, Hyper-V, KVM, or Citrix)

  • Familiarity with cloud computing platforms (e.g. AWS, Azure, Google Cloud)

Ways to stand out from the crowd:

  • Knowledge of CPU and/or GPU architecture

  • Knowledge of Kubernetes, container related microservice technologies

  • Experience with GPU-focused hardware/software (DGX, Cuda)

  • Experience with RDMA (InfiniBand or RoCE) fabrics

With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 24, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior HPC AI Cluster Engineer in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...technology powers everything from generative AI to autonomous systems, and we continue to...  ..., and tools that enable researchers and engineers to develop the next generation of AI/ML systems...  .... We are looking for a strong AI & HPC Observability Engineer to build and scale... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...high-performance computing, cloud, and AI. Whether you’re designing next-gen...  ...THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms...  ...administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE... 
    Suggested

    AMD

    San Jose, CA
    2 days ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have a pivotal...  ...storage systems like Lustre and GPFS for AI/HPC workloadsFamiliarity with deep learning frameworks... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU...  ...on the world.We are seeking a highly skilled and experienced HPC Cluster Engineer to design, deploy, and operate GPU Compute Clusters for EDA (... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...performance computing, cloud, and AI. Whether you’re designing next...  ...of large-scale AI/ML clustered infrastructure. You will be working...  ...growing team of multi-disciplined engineers that operates across industry...  ...patterns, Kubernetes for HPC/AI (GPU operators, device plugins... 
    Suggested
    Flexible hours

    AMD

    Santa Clara, CA
    2 days ago
  • $160k - $198k

     ...of our team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and...  ...systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute...  ...Kubernetes), paired with cloud-agnostic cluster abstractors like SkyPilot to manage... 
    Senior
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  •  ...A leading AI technology firm in California is seeking an experienced Senior Software Engineer to develop and optimize AI infrastructure software using state-of-the-art GPU systems. Candidates should have a Bachelor's degree in a technical field and a minimum of 5 years... 
    Senior

    Intelliswift - An LTTS Company

    Sunnyvale, CA
    17 hours ago
  • $152k - $241.5k

     ...into the unlimited potential of AI to define the next era of...  ...the world. We are looking for a Senior Software Engineer to join our mission to continue improving our HPC infrastructure. Our team builds...  ...infrastructure or control planes for HPC clusters, large‑scale AI/ML platforms,... 
    Senior

    NVIDIA Gruppe

    Santa Clara, CA
    17 hours ago
  •  ...ASML Germany GmbH is seeking an experienced engineer to develop and maintain HPC infrastructure supporting scalable workloads across products. You will design compute clusters, storage, and networking, delivering stable, production-ready platforms with strong observability... 
    Senior

    ASML Germany GmbH

    San Jose, CA
    17 hours ago
  • $176k - $333.5k

     ...deep learning ignited modern AI and enabled the next era of computing...  ...today. We are looking for a Senior Software Engineer to join our mission to continue improving our HPC infrastructure. Our team...  ...to meet the demands of our HPC clusters Evaluate new and innovative technologies... 
    Senior

    NVIDIA

    Santa Clara, CA
    17 hours ago
  •  ...for the testing and evaluation of current and next-generation HPE HPC products. Ensure development issues are resolved in a cost-...  ...remote labs. Drive appropriate automated test execution to test engineers at various global locations. Provide training and guidance to... 
    Local area
    Remote work

    Net2Source

    San Jose, CA
    2 days ago
  • KLA is seeking a Principal HPC Architect to design, build, optimize...  ...for scientific computing, AI/ML workloads, and data-intensive...  ...HPC infrastructure and lead engineering initiatives. The role emphasizes...  ..., and reliability across clusters and storage. #J-18808-Ljbffr... 
    Senior

    KLA

    Milpitas, CA
    12 hours ago
  • $152k - $241.5k

    We’re currently seeking a Senior AI Developer Technology Engineer, Financial Sector!Would you like to help shape the future of financial AI and data analytics...  ...in-depth analysis and optimization of complex AI and HPC workloads to ensure the best possible performance on... 
    Senior
    Full time
    Work experience placement
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195.2k - $275.58k

    Job Details:Job Description: The Software and AI (SAI) organization is seeking a highly skilled Software Development Engineer to contribute to the development and optimization...  ...with top experts in deep learning, HPC, compilers, and systems optimization.Contribute... 
    Senior
    Full time
    Local area
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Intel

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are looking for a software engineer with a strong background in parallel processing and GPU...  ...of performance at the intersection of AI, high-performance computing, and financial...  ...analyze, optimize, and scale complex AI and HPC workloads for modern CPU and GPU architectures... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $227k - $300k

     ...Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles....  ...the Edge. We are looking for a great Senior Staff AI Engineer to join our seasoned AI team and lead the...  ....Develop algorithms to automatically cluster log patterns and detect software regressions... 
    Senior
    Work at office
    Worldwide
    Flexible hours
    Shift work
    3 days per week

    Sonatus

    Sunnyvale, CA
    2 days ago
  • $203.45k - $344.3k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...transmission, data processing, data clustering, data mining, data evaluation, data...  ...costs, data architecture and closed-loop engineering system.Job ResponsbilitiesResponsible for... 
    Senior
    Full time
    Temporary work
    Work experience placement

    XPENG Motors

    Santa Clara, CA
    3 days ago
  • HPE Labs - Senior Software Engineer - Integrated HPC & Quantum SolutionsThis role has been designed as 'Hybrid' with...  ...GPU/CPU function and applicability, cluster design, job scheduling, workload...  ...computing platforms, GPU computing, AI/ML frameworks (e.g., TensorFlow, PyTorch... 
    Senior
    Full time
    Work experience placement
    Work at office
    Local area
    Immediate start
    2 days per week

    Hewlett Packard Enterprise

    Milpitas, CA
    12 hours ago
  • $184k - $356.5k

     ...product that powers innovative AI research and developers. We...  ...an AI infrastructure software engineer to join our team. You'll be instrumental...  ...AI in production. As a senior DGX Cloud AI Infrastructure...  ...working with the large scale AI cluster and cloud-native... 
    Senior
    Full time

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $262k - $364k

     ...Infrastructure layer of the GDC product line.Drive and influence engineering excellence across the organization, establish Service Level...  ...full-stack as we continue to push technology forward.The Infra Cluster team delivers the core Kubernetes-based compute infrastructure... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

     ...a Solutions Architect with a performance engineering background who can help our most sophisticated...  ...Robotics customers accelerate Physical AI workloads using NVIDIA's full-stack...  ...within at least one of these areas: LLM and HPC. Having expertise ranging from operator-level... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...NVIDIA is seeking a Senior Software Engineer to lead AI resiliency for the world’s most powerful AI supercomputers. You will drive the development of...  ...scalable fault tolerance, while enabling robust production deployments in cloud and HPC environments. #J-18808-Ljbffr
    Senior

    NVIDIA

    Santa Clara, CA
    17 hours ago
  • $176k - $276k

     ...intelligence.Join our team of innovative engineers who develop and maintain software facilitating...  ...while managing and maintaining large GPU clusters interconnected via NVLink and InfiniBand....  ...is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...that accelerate next‑generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems....  ...'s most important challenges.As the Director of Cloud, HPC & Sovereign AI Customer Engineering within the Compute & Enterprise AI SolutionsCustomer Engineeringorganization... 
    Remote work

    Advanced Micro Devices

    Santa Clara, CA
    3 days ago
  • $193.3k - $261.5k

     ...are seeking an experienced engineer to work on distributed AI/ML systems. This role...  ...high-speed networking or HPC interconnects is valued highly...  ...features for the largest clusters, with the largest customers...  ...seriously, you can both expect senior mentorship and will be... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    12 hours ago
  • $139k - $204k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave...  .... About the role As part of the Cluster Orchestration team, you will play a key...  ...possible with AI. What You'll Do As a Senior Software Engineer I (IC3), you will own multiple services... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    24 days ago
  •  ...ASML US, LLC is seeking an engineer to develop and maintain HPC infrastructure, including compute clusters, storage, and networking to support scalable workloads across products. You will work on containerization, hybrid deployments, observability, and performance tuning... 
    Senior

    ASML US, LLC

    San Jose, CA
    17 hours ago
  •  ...NVIDIA is seeking a senior software engineer for its communication libraries and network software team. The role centers on designing...  ...runtimes for deep learning frameworks and HPC interfaces on GPU clusters. You will contribute to MPI/OpenSHMEM specs, develop... 
    Senior

    NVIDIA AI

    Santa Clara, CA
    17 hours ago
  • $255k - $340k

     ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ...home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ...lead for integrating OEM and white-label HPC AI/ML, general purpose compute, storage,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    12 hours ago
  • $200k - $400k

     ...by those who field intelligent machines at scale. At Scout AI, we’re developing Fury, the first robotic foundation model...  ...precision, and relentless work. The Role We're looking for a Senior or Staff AI Engineer to join the Fury Orchestration Team with a deep passion for... 
    Senior
    Full time
    Relocation package

    Scout Ai

    Sunnyvale, CA
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC AI Cluster Engineer. Be the first to apply!