Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior HPC and AI Networking Performance Research and Analysis Engineer

NVIDIA

Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer

Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines to learn, perceive, reason and solve problems. Today, visual computing is a crucial tool in helping people get along with technology, and NVIDIA has extended its technology into datacenters, mobile devices and cars. There has never been a more exciting time to join our team - if this role sounds like a fit for you, we'd love to hear from you!

NVIDIA is seeking a Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.

What you'll be doing:

  • Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on NVIDIA supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).
  • Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.
  • Implementing performance analysis tools.
  • Collaborating with many teams from hardware to software to provide performance analysis insights.
  • Defining performance test planning, setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

What we need to see:

  • B.Sc in Computer Science or Software Engineering or equivalent experience
  • 5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)
  • Demonstrated Performance Analysis skills and methodologies.
  • Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).
  • Fast and self-learning capabilities with strong analytical and problem-solving skills.
  • Programming Languages: Python, Bash and C languages
  • Experience with Linux OS distros.
  • Great teammate with good communication and interpersonal skills

Ways to stand out from the crowd:

  • In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.
  • Knowledge in CUDA, and NCCL libraries.
  • Knowledge in Congestion Control algorithms.
  • In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).
  • Strong Performance Analysis skills and methodologies using modern tools.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. We have a unique legacy of innovation that's fueled by great technology—and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world!

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior HPC and AI Networking Performance Research and Analysis Engineer in Santa Clara, CA vacancy
  • $184k - $287.5k

     ...unlimited potential of AI to define the...  ...in the research, design and implementation...  ...demanding high performance computing, and...  ...services for HPC workloads,...  ...performance analysis and optimizations...  ...Science, Electrical Engineering or related...  ...appliances like Network Appliance.... 
    Senior
    Performance
    Network

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $152k - $241.5k

     ...unlimited potential of AI to define the...  ...and experienced HPC Cluster Engineer to design,...  ...Automation) and high-performance computing...  ...collaborate with researchers and infrastructure...  ...deployment of compute, networking, and storage....  ...performance analysis and optimizations... 
    Senior
    Performance
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the...  ...for rack, networking, and datacenter provisioning...  ...management. As a Senior Software Engineer - Datacenter...  ...today's fastest HPC and AI workloads....  ...systems in high-performance or distributed environments...  ...monitoring and analysis. Experience... 
    Senior
    Performance
    Network
    Remote work

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...currently seeking a Senior Developer Technology Engineer for High-Performance Databases! Would you enjoy researching new algorithms and...  ...perform in-depth analysis and optimization...  ...storage systems, networking, and distributed computer...  .... NVIDIA uses AI tools in its... 
    Senior
    Performance
    Network

    NVIDIA

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...currently seeking a Senior Developer Technology Engineer! NVIDIA's...  ...team is a global network of world-class experts...  ...this role, you will research and develop techniques...  ...key customers to perform in-depth analysis and optimization...  .... NVIDIA uses AI tools in its recruiting... 
    Senior
    Performance
    Network
    Work experience placement

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...learning ignited modern AI — the next era of...  ...Learning Compiler Engineer. NVIDIA is hiring...  ...leading inference performance, fast build time,...  ...deep learning networks and developing compiler...  ...optimizations and analysis, crafting and...  ...relevant work or research experience in performance... 
    Senior
    Performance
    Network

    NVIDIA

    Santa Clara, CA
    2 days ago
  • $140k - $160k

     ...Federal is looking for a Senior HPC Engineer, as ASRC Federal InuTeq provides High Performance Computing services across...  ...(hardware, software, and network). Understands research use cases, researches and...  ...systems System performance analysis and tuning Building, installing... 
    Senior
    Performance
    Network
    Contract work
    Weekend work

    ASRC Federal Holding Company

    Mountain View, CA
    1 day ago
  •  ...Senior Software Engineer, Systems/Solutions Test This role has...  ...executing tests for networking products, including routers...  ..., scalability, and performance across complex...  ..., perform root-cause analysis, and drive corrective...  ...technologies, including AI-assisted testing workflows... 
    Senior
    Performance
    Network
    Work at office
    2 days per week

    Hewlett Packard Enterprise

    Sunnyvale, CA
    3 days ago
  • $184k - $287.5k

     ...seeking an outstanding Senior Systems Engineer to join our dynamic...  ...firmware engineers, networking architects, and experts...  ..., data science, and AI to develop seamless and...  ..., scalable, and performant hardware-accelerated...  ...in system performance analysis and low-level programming... 
    Senior
    Performance
    Network
    Remote work
    Shift work

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $120k - $243k

     ...Senior Engineer, Solution Engineering This role has been...  ...requirements and performance expectations. The...  ...strong expertise in networking technologies, virtualization...  .... Leverage AI tools, technologies,...  ..., perform root cause analysis, and support corrective... 
    Senior
    Performance
    Network
    Work experience placement
    Work at office

    Hewlett Packard Enterprise Development LP

    Sunnyvale, CA
    9 hours ago
  • $184k - $287.5k

     ...boundaries of innovation and engineering? At NVIDIA, we lead...  ...—driving progress in AI, graphics, and high‑performance systems. As a Senior Hardware Systems...  ...issues, driving root‑cause analysis, and ensuring timely...  ...designs, such as new power networks, liquid cooling, or... 
    Senior
    Performance
    Network

    NVIDIA

    Santa Clara, CA
    6 hours ago
  • $70 - $78 per hour

     ...the production engine designed for today...  ...'s always on, AI driven...  .... What does a Senior QA Engineer do...  ...software defects, research root causes, debug...  ...coverage and defect analysis. Proven...  ...Experience with performance and/or security...  ...WPP’s unmatched network, this is a place... 
    Senior
    Performance
    Network
    Contract work
    Work at office
    3 days per week

    WPP Production

    Sunnyvale, CA
    2 days ago
  •  ...Senior Systems Engineer Graphcore is one of the...  ...generation of AI breakthroughs...  ...melting pot of AI research specialists,...  ...reliability and performance of next-...  ...power anomalies, network configuration,...  ...perform root cause analysis and propose...  ...experience with HPC systems, AI... 
    Senior
    Performance
    Network

    Graphcore

    Milpitas, CA
    2 days ago
  • $190.58k - $200k

     ...Research Computing GPU Systems Engineer Business Affairs: University IT (UIT...  ...infrastructure, driving system performance and reliability...  ...research in AI/ML, computational biology...  ...and InfiniBand RDMA networking. Develop...  ...experience. ~5+ years in HPC systems administration... 
    Performance
    Network
    Hourly pay
    Full time
    Flexible hours
    Weekend work
    Afternoon shift

    Stanford University

    Stanford, CA
    3 days ago
  • $184k - $287.5k

     ...potential of AI to define the...  ...a dedicated engineer for the Senior Systems Software...  ...on GPU Performance at Scale. At...  ...Collaborate with researchers, developers,...  ...Engage with HPC, OS, CPU, GPU...  ...GPUs, CPUs, and networking hardware. Engage...  ...to systems analysis. Linux... 
    Senior
    Performance
    Network

    NVIDIA

    Santa Clara, CA
    5 days ago
  • $200k - $400k

     ...Foundation Models Engineer The Institute of...  .... We believe performance, fault tolerance,...  ...This is not a network operations role. This...  ...concepts · Congestion analysis and routing optimization...  ...to translate research ideas into...  ...distributed systems, HPC, or large-scale training... 
    Senior
    Performance
    Network
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    5 days ago
  •  ...Description PlusAI is a Physical AI company pioneering AI-based...  ...-growing teams. As a Research Engineer, you will deliver mission-...  ...participating in vehicle performance analysis, tuning, and...  ...control and modern neural network architectures, with expertise... 
    Senior
    Performance
    Network
    Full time

    PlusAI

    Santa Clara, CA
    11 days ago
  • $82 - $85 per hour

     ...Metal Design-Expert, Tolerance Analysis-Expert, Manufacturing-Metal-...  ...Mechanical and Thermal team, the Senior Mechanical Engineer is tasked with leading the...  ...devices designed for the 5G era and AI networks, ensuring adaptability, high performance, and efficiency. Key... 
    Senior
    Performance
    Network
    Hourly pay
    Contract work

    Akraya

    San Jose, CA
    5 days ago
  • $75 - $80 per hour

     ...Description Job Title : Senior RF Test & PCB Design Engineer Position Description...  .... Simulation & Analysis: Perform Signal Integrity (SI) and...  ...transmission line theory, matching networks, and de-embedding...  ...RF: Experience applying AI/ML models for predictive... 
    Senior
    Performance
    Network
    Contract work

    Protingent

    San Jose, CA
    6 hours ago
  • $186k - $279k

     ...The Storage Benchmarking Engineer will design, execute, and analyze performance benchmarks using...  ...entire stack (compute, network, and storage), along with...  ...LL DO Configure HPC lab environment so all systems...  ...~ Proficiency in data analysis and visualization tools... 
    Senior
    Performance
    Network
    Work at office
    Flexible hours

    Everpure LLC

    Santa Clara, CA
    5 days ago
  • $130k - $225k

     ...Description Arista Networks is an industry leader...  ...awards, such as Best Engineering Team, Best Company for...  ...standards of quality and performance in everything we do....  ...and bleeding edge AI clusters. With this comes...  ...delivering root cause analysis and fixes to the design... 
    Senior
    Performance
    Network

    Arista Networks, Inc.

    Santa Clara, CA
    5 days ago
  •  ...hardware security,AI drivenresilience...  ...more than great engineering—it takes a team...  ...We are hiring a Senior QA Engineer – Performance & Reliability to...  ..., and deep-dive analysis for performance...  ...DDR memory, and networking protocols. Performance...  ...world's leading research, technology and... 
    Senior
    Performance
    Network

    Axiado

    San Jose, CA
    1 day ago
  • $150k - $230k

     ...Description Arista Networks is an industry leader...  ...awards, such as Best Engineering Team, Best Company for...  ...standards of quality and performance in everything we do....  ...Networks is seeking a Senior Optical Signal...  .... Drive root cause analysis of SI/PI-related issues... 
    Senior
    Performance
    Network
    Full time
    Contract work

    Arista Networks

    Santa Clara, CA
    a month ago
  • $130k - $200k

     ...Description Arista Networks is an industry leader...  ...awards, such as Best Engineering Team, Best Company for...  ...standards of quality and performance in everything we do....  ...Arista seeks to add a Senior Thermal Engineer to...  ...perform complete thermal analysis and design using... 
    Senior
    Performance
    Network
    Full time

    Arista Networks

    Santa Clara, CA
    6 days ago
  • $150k - $230k

     ...Description Arista Networks is an industry leader...  ...awards, such as Best Engineering Team, Best Company for...  ...standards of quality and performance in everything we do....  ...an exceptional Senior Optical Transceiver Design...  ...qualification, and root-cause analysis during development and... 
    Senior
    Performance
    Network

    Arista Networks, Inc.

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...NVIDIA Math Libraries team is looking for a senior engineer to join our development efforts in the area of kernel generation for AI and HPC, specifically targeting matrix...  ...designing, and implementing high quality and performance numerical dense linear algebra software... 
    Senior
    Performance

    NVIDIA

    Santa Clara, CA
    3 days ago
  • $140k - $175k

     ...Description Arista Networks is an industry leader...  ...awards, such as Best Engineering Team, Best Company for...  ...standards of quality and performance in everything we do....  ...and strategic Senior Electrical Manufacturing...  ...traffic generation and analysis, full line-rate testing... 
    Senior
    Performance
    Network
    Full time
    Contract work

    Arista Networks

    Santa Clara, CA
    6 days ago
  • $100k - $200k

     ...quality of life. As a Senior Wireless and Algorithm Design Engineer, you will be responsible...  ...satellite communication networks. You will leverage your...  ...that meet rigorous performance standards. This position...  ...artificial intelligence (AI) tools to support parts of... 
    Senior
    Performance
    Network
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    eSpace

    Saratoga, CA
    4 days ago
  • $50 - $70 per hour

     ...enabling enhanced performance and user experiences...  ...within Systems Engineering group is looking for a motivated Senior QA Automation Engineer...  ...like generative AI and advanced...  ...controllers, and networking setup. Analyze...  ...drive root-cause analysis for complex system... 
    Senior
    Performance
    Network
    Hourly pay
    Contract work

    SK Hynix Memory Solutions America Inc.

    San Jose, CA
    4 days ago
  • $300 per month

     ...vertically integrated AI infrastructure...  ...be part of a high-performing team that believes...  ...— and Production Engineering sits at the heart...  ...demanding AI and HPC workloads. You’...  ...reviews and root cause analysis Build, operate,...  ...with compute, networking, storage, and platform... 
    Senior
    Performance
    Network
    Temporary work

    Crusoe

    Sunnyvale, CA
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior HPC and AI Networking Performance Research and Analysis Engineer. Be the first to apply!