Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC AI Cluster Engineer

$176k - $276k
Full-time

NVIDIA

NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers and AI clusters based on groundbreaking technologies. We are looking for an outstanding engineer, be a key player to the most exciting computing hardware and software to contribute to the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs. You will work with the latest Accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.What you will be doing:Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alertingManage Linux job/workload schedules and orchestration toolsDevelop and maintain continuous integration and delivery pipelinesDevelop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resourcesDeploy monitoring solutions for the servers, network and storagePerform troubleshooting bottom up from bare metal, operating system, software stack and application levelBeing a technical resource, develop, re-define and document standard methodologies to share with internal teamsSupport Research & Development activities and engage in POCs/POVs for future improvementsWhat we need to see:A degree in Computer Science, Engineering, or a related field (or equivalent experience) and 8+ years of experienceKnowledge of HPC and AI solution technologies from CPU’s and GPU’s to high speed interconnects and supporting softwareExperience with job scheduling workloads and orchestration tools such as Slurm, K8sExcellent knowledge of Windows and Linux (Redhat/CentOS and Ubuntu) networking (sockets, firewalld, iptables, wireshark, etc.) and internals, ACLs and OS level security protection and common protocols e.g. TCP, DHCP, DNS, etc.Experience with multiple storage solutions such as Lustre, GPFS, Weka.io. Familiarity with newer and emerging storage technologies.Python programming and bash scripting experience.Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/chefDeep knowledge of Networking Protocols like InfiniBand, EthernetDeep understanding and experience with virtual systems (for example VMware, Hyper-V, KVM, or Citrix)Familiarity with cloud computing platforms (e.g. AWS, Azure, Google Cloud)Ways to stand out from the crowd:Knowledge of CPU and/or GPU architectureKnowledge of Kubernetes, container related microservice technologiesExperience with GPU-focused hardware/software (DGX, Cuda)Experience with RDMA (InfiniBand or RoCE) fabricsWith competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior HPC AI Cluster Engineer in Santa Clara, CA vacancy
  • $176k - $276k

    NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers and AI clusters based on groundbreaking technologies. We are looking for an outstanding engineer, be a key player... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...technology powers everything from generative AI to autonomous systems, and we continue to...  ..., and tools that enable researchers and engineers to develop the next generation of AI/ML systems...  .... We are looking for a strong AI & HPC Observability Engineer to build and scale... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...scientific discovery to powering AI and the technologies people...  ...we’ll advance your career.HPC & AI Researcher Engineer - PhD preferredSOFTWARE...  ...team is seeking a senior engineer to lead the enablement...  ...deployment at production level clusters. You will collaborate with... 
    Suggested

    AMD

    Santa Clara, CA
    14 hours ago
  •  ...high-performance computing, cloud, and AI. Whether you’re designing next-gen...  ...THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms...  ...administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE... 
    Suggested

    AMD

    San Jose, CA
    14 hours ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency efforts. As an Engineer, you will have a pivotal...  ...storage systems like Lustre and GPFS for AI/HPC workloadsFamiliarity with deep learning frameworks... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $160k - $198k

     ...of our team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and...  ...systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute...  ...Kubernetes), paired with cloud-agnostic cluster abstractors like SkyPilot to manage... 
    Senior
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...everything from generative AI to autonomous systems,...  ...enable researchers and engineers to develop the next...  ...globally.We are looking for a senior networking engineer to...  ...are scaling it toward clusters of ten thousand nodes...  ...fabrics in AI or HPC environments - InfiniBand... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $76.92 - $95.19 per hour

     ...across air taxis, UAS, AI and powertrain development...  ...re seeking exceptional engineers, operators and builders...  ....What You’ll DoAs a Senior AI Systems Engineer, you...  ...performance computing (HPC), or ML infrastructure....  ...with cloud-agnostic cluster abstractors like SkyPilot... 
    Senior
    Full time
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    1 day ago
  •  ...for the testing and evaluation of current and next-generation HPE HPC products. Ensure development issues are resolved in a cost-...  ...remote labs. Drive appropriate automated test execution to test engineers at various global locations. Provide training and guidance to... 
    Local area
    Remote work

    Net2Source

    San Jose, CA
    4 days ago
  • $100k

     ...leading the industry on cutting-edge AI technology, revolutionizing...  ...and looking for contributors of all seniorities.Tenstorrent is building next-generation...  ...autonomous AI systems. As the RISC-V AI / HPC & Agentic Software Engineering Lead, you will operate at the... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our Compute...  ...harness the power of AI to deliver...  ...integrate cleanly with HPC schedulers, storage, and...  ...supporting large‑scale HPC clusters using Slurm, LSF or Kubernetes...  ...or Ruby.Mentored other engineers and influenced... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    We are looking for a software engineer with a strong background in parallel processing and GPU...  ...of performance at the intersection of AI, high-performance computing, and financial...  ...analyze, optimize, and scale complex AI and HPC workloads for modern CPU and GPU architectures... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $203.45k - $344.3k

     ...forefront of innovation, integrating advanced AI and autonomous driving technologies into...  ...transmission, data processing, data clustering, data mining, data evaluation, data...  ...costs, data architecture and closed-loop engineering system.Job ResponsbilitiesResponsible for... 
    Senior
    Full time
    Temporary work
    Work experience placement

    XPENG Motors

    Santa Clara, CA
    14 hours ago
  • $184k - $287.5k

     ...a Solutions Architect with a performance engineering background who can help our most sophisticated...  ...Robotics customers accelerate Physical AI workloads using NVIDIA's full-stack...  ...within at least one of these areas: LLM and HPC. Having expertise ranging from operator-level... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $144.7k - $261.3k

     ...environments, cloud infrastructure, and ML/AI GPU platforms for AV research and...  ...The Role : GM is looking for a Senior Performance Engineer to join the AV Capacity and Performance...  ...design and high-performance computing (HPC). Containerization: Hands-on experience... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours
    3 days per week

    General Motors

    Sunnyvale, CA
    3 days ago
  • $139k - $204k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave...  ....  About the role As part of the Cluster Orchestration team, you will play a key...  ...possible with AI. What You'll Do As a Senior Software Engineer I (IC3), you will own multiple services... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    14 days ago
  • $323k

     ...Summary:Are you passionate about optimizing AI workloads and delivering real-world...  ...devices?We’re looking for an experienced engineer to help customers achieve best-in-class inference...  ...to audiences ranging from engineers to senior leadership.Responsibilities:Develop... 
    Senior
    Work at office
    Local area

    ARM

    San Jose, CA
    6 hours ago
  • $152k - $241.5k

     ...computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain...  ...known as “the AI computing company”.NVIDIA is hiring a Senior AI Compiler Engineer. GPUs are driving rapid progress in deep learning—from LLMs... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    We are looking for outstanding Senior High Performance AI Engineers to build the next generation of agentic AI systems for the CUDA ecosystem. Our team works across the full agentic AI stack—from training and improving models, to designing agent architectures and multi... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...consulting, and customer experience with agile engineering and problem-solving creativity. United by...  ...company that delivers enterprise AI platforms and services, helping organizations...  ...and measurable business impact. As a Senior Forward Deployed Engineer, you will work... 
    Senior

    Publicis Media

    San Jose, CA
    5 days ago
  • $255k - $340k

     ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands...  ...home day is currently Tuesday.Hardware Engineering at Lambda is responsible for building and...  ...lead for integrating OEM and white-label HPC AI/ML, general purpose compute, storage,... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $152k - $241.5k

    We are now looking for a Senior AI Frameworks Engineer (C++/Python)! NVIDIA's high-performance computing platforms are powering the AI revolution across many applications and industries. Within our software stack, CUTLASS stands out as a popular open-source ecosystem dedicated... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU...  ...on the world.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $255.65k - $299k

    Immigration sponsorship is not available for this positionWhat you can expect:As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI algorithms and software applications. You will ensure AI training, inference, deployment, and operation... 
    Senior
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    3 days ago
  • $184k - $287.5k

    Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts...  ...we can make a lasting impact on the world.NVIDIA is hiring senior software engineers in its Infrastructure, Planning and Process Team (IPP), to... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago
  • $190k - $237k

     ...team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage...  ...ML systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute...  ...), paired with cloud-agnostic cluster abstractors like SkyPilot to manage multi... 
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  • $184k - $287.5k

     ...is the industry leader in high performance computing, gaming and AI. Our GPUs and SOCs give outstanding performance and efficiency,...  ...optimized during the Blackwell generation alone! Now we're hiring the engineer who will lead the rebuild of that toolchain around AI.We focus... 
    Senior
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $264.5k

    Nvidia's SOC Design (SOCD) team is looking for an Applied AI Engineer who is passionate about eliminating bottlenecks in SOC integration workflows through intelligent automation. If you are driven to build AI-powered tools, agents, and automation solutions to dramatically... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    NVIDIA is looking for a Senior Applied AI Engineer to help build intelligent software systems that improve engineering productivity and software quality at scale. In this role, you will develop AI-powered workflows and services that help teams analyze complex code changes... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $187k - $215k

     ...the intelligence layer for global trade and logistics. We turn fragmented, messy data into instant, reliable intelligence, enabling AI to reason and act in high-stakes, real-world environments. We are a small, sharp team operating at the frontier, and we are looking for... 
    Senior
    Work at office
    Flexible hours

    KlearNow.AI

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC AI Cluster Engineer. Be the first to apply!