Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer, AI Cluster Networking Performance Engineer

$167.7k - $245.2k

Jobleads-US

Meet the Team Cisco Validated Infrastructure Services (CVIS) establishes and proves the performance of large AI clusters before they are handed over to customers. We deploy each cluster, validate it against performance baselines, and deliver the evidence that it is ready for service. The platform behind that work is full stack. It combines workflow and API services, network and infrastructure automation, and the UI used to run validation across labs and customer environments.

Your Impact

As a Senior Software Engineer, you will be responsible for things like debugging, profiling and optimizing the full AI cluster networking and GPU performance stack, from boot-time and PCIe behavior through Linux kernel, SmartNIC/DPU, RDMA, NCCL, and application-level collective communication.

  • Investigate kernel-level failures, boot-time anomalies, PCIe link issues, driver behavior, and network latency.
  • Diagnose and tune RDMA/RoCEv2, NCCL, SmartNIC/DPU, NIC queues, congestion, throughput, and collective communications.
  • Read and correlate unified hardware traces across host CPUs, PCIe switches, SmartNICs, and GPU compute streams.
  • Capture and analyze Wireshark packets, system logs, NVIDIA support logs, traces, profiles, and observability data.
  • Partner with hardware and software teams to eliminate end-to-end bottlenecks.

Minimum Qualifications

  • Bachelors + 7 years of related experience, or Masters + 4 years of related experience, or PHD +1 year of related experience or equivalent related work experience.
  • 4+ years of experience working with Linux kernel, networking, GPU, and distributed-systems debugging skills.
  • Experience writing code with C/C++/Python and experience with user-mode and kernel-mode drivers.
  • Experience with Linux network stack, RDMA/RoCEv2, NCCL, PCIe, and performance profiling.
  • Prior experience working with reason from packet captures, traces, counters, logs, and reproducible benchmarks.

Preferred Qualifications

  • SmartNIC/DPU optimization and DOCA SDK experience.
  • Cisco switching, AI cluster network fabrics, multi-plane/ToR designs, and high-speed Ethernet or InfiniBand environments.
  • Familiarity with NVIDIA diagnostic tooling, Nsight/profiling tools, kernel tracing, and observability platforms.

Why Cisco?

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you.

Message to applicants applying to work in the U.S. and/or Canada

The starting salary range posted for this position is $167,700.00 to $245,200.00 and reflects the projected salary range for new hires in this position in U.S. and/or Canada locations, not including incentive compensation*, equity, or benefits. Individual pay is determined by the candidate's hiring location, market conditions, job‑related skillset, experience, qualifications, education, certifications, and/or training. The full salary range for certain locations is listed below. For locations not listed below, the recruiter can share more details about compensation for the role in your location during the hiring process.

  • U.S. employees are offered benefits, subject to Cisco’s plan eligibility rules, which include medical, dental and vision insurance, a 401(k) plan with a Cisco matching contribution, paid parental leave, short and long‑term disability coverage, and basic life insurance.
  • Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to receive grants of Cisco restricted stock units, which vest following continued employment with Cisco for defined periods of time.
  • U.S. employees are eligible for paid time away as described below, subject to Cisco’s policies:

10 paid holidays per full calendar year, plus 1 floating holiday for non‑exempt employees 1 paid day off for employee’s birthday, paid year‑end holiday shutdown, and 4 paid days off for personal wellness determined by Cisco Non‑exempt employees** receive 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full‑time employees Exempt employees participate in Cisco’s flexible vacation time off program, which has no defined limit on how much vacation time eligible employees may use (subject to availability and some business limitations) 80 hours of sick time off provided on hire date and each January 1st thereafter, and up to 80 hours of unused sick time carried forward from one calendar year to the next Additional paid time away may be requested to deal with critical or emergency issues for family members Optional 10 paid days per full calendar year to volunteer For non‑sales roles, employees are also eligible to earn annual bonuses subject to Cisco’s policies. Employees on sales plans earn performance‑based incentive pay on top of their base salary, which is split between quota and non‑quota components, subject to the applicable Cisco plan. For quota‑based incentive pay, Cisco typically pays as follows: .75% of incentive target for each1% of revenue attainment up to 50% of quota; 1.5% of incentive target for each1% of attainment between 50% and 75%; 1% of incentive target for each1% of attainment between 75% and 100%; and Once performance exceeds100% attainment, incentive rates are at or above 1% for each1% of attainment with no cap on incentive compensation. For non‑quota‑based sales performance elements such as strategic sales objectives, Cisco may pay0% up to125%of target. Cisco sales plans do not have a minimum threshold of performance for sales incentive compensation to be paid.

The applicable full salary ranges for this position, by specific state, are listed below:

  • New York City Metro Area: $167,700.00 - $282,000.00
  • Non‑Metro New York state & Washington state: $149,100.00 - $250,900.00

* For quota‑based sales roles on Cisco’s sales plan, the ranges provided in this posting include base pay and sales target incentive compensation combined.

** Employees in Illinois, whether exempt or non‑exempt, will participate in a unique time off program to meet local requirements.

Cisconians power the future. We make impact as a team, innovating fast and fearlessly to create meaningful solutions on a large scale. The depth and breadth of our technology doesn't just benefit our customers – it also means limitless opportunities for us to experiment and learn. We understand the power each of our unique backgrounds bring when we work together. Because of that, we have a global network of thinkers, doers, experts, and curious creators who help one another do their life’s best work.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer, AI Cluster Networking Performance Engineer in Milpitas, CA vacancy
  •  ...scientific discovery to powering AI and the technologies...  ...:Join AMD's IT Systems Engineering team and help build the networking foundation powering some...  ...advanced AI and high-performance computing environments....  ...supporting AMD Instinct GPU clusters used for AI training,... 
    Senior
    Performance

    AMD

    San Jose, CA
    2 days ago
  •  ...NVIDIA Corporation seeks a Senior Solution Engineer, Networking to lead expertise in high‑performance network tech for AI clusters. You will collaborate with Enterprise Experience teams, R&D, and field engineers to diagnose, reproduce, and root cause complex customer issues... 
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • $136.3k - $231.7k

     ...of physicists, engineers, data scientists...  ...scientist, software engineers, application...  ...engineers, and senior product...  ...including Linux networking, file systems,...  ...technologies, and performance tuning.Hands-on...  ...with Ansible, cluster provisioning solutions...  ...drivers, CUDA, AI/ML... 
    Performance
    Minimum wage
    Full time
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    2 days ago
  •  ...Cisco Systems, Inc. is seeking a Senior Software Engineer to debug, profile, and optimize the AI cluster networking and GPU performance stack across boot-time to kernel level, SmartNIC/DPU, and NCCL-based communications. You will investigate kernel-level failures,... 
    Senior
    Performance

    Jobleads-US

    Milpitas, CA
    12 hours ago
  •  ...Inc. in Milpitas, CA is seeking a Senior Software Engineer to own CVIS/NVIS AI cluster validation, end-to-end solution...  ..., GPU debugging, and performance profiling. You will model workloads...  ...reports while validating Cisco network-switch integrations and coordinating... 
    Senior
    Performance

    Jobleads-US

    Milpitas, CA
    12 hours ago
  •  ...Cisco Systems, Inc. seeks a Senior Software Engineer to own CVIS/NVIS and AI cluster validation end-to-end. You will run GPU debugging, performance profiling, bottleneck analysis, and craft...  ...release-ready evidence across Cisco network-switch solutions. You will develop... 
    Senior
    Performance

    Jobleads-US

    Milpitas, CA
    12 hours ago
  • $139k - $204k

     ...is The Essential Cloud for AI™. Built for pioneers by pioneers...  ...superior infrastructure performance with deep technical...  ...About the role As part of the Cluster Orchestration team, you will...  ...AI. What You'll Do As a Senior Software Engineer I (IC3), you will own multiple... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    a month ago
  • $184k - $287.5k

    NVIDIA is seeking a Senior Software Engineer to help us develop distributed storage services for AI/ML. In this role you will work closely...  ...to ensure scalable, high-performance, and reliable solutionsHistory...  ...that runs on large-scale clusters, multi-petabyte to exabyte in... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...multi-rack, multi-tenant AI/ML datacenters with...  ...GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud...  ...What you’ll be doing:Perform deep-dive debugging of...  ...multi-rack, multi-tenant clusters: scheduler behavior,...  ...native stacks across networking (RDMA/RoCE), storage,... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...United StatesJob Title: Senior Kubernetes Platform Engineer / ArchitectLocation:...  ...Kubernetes clusters.Develop container platform...  ...Kubernetes clusters.Perform cluster lifecycle...  ...Configure namespaces, RBAC, network policies, storage...  ....Experience with AI/ML platform deployments... 
    Senior
    Performance

    Apptad

    Milpitas, CA
    2 days ago
  • $200k - $322k

     ...the infrastructure and software platform that enables...  ..., train, and deploy AI at scale. As demand...  ...We are looking for a Senior Software Engineer to design and build the...  ...providers, regions, clusters, and products.Automate...  ...the reliability, performance, and scalability of capacity... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...of the generative AI revolution, building the software and systems that...  ...are looking for a Senior Software Engineer to lead the bring...  ...will lead deep performance and reliability investigations...  ...that keep large clusters productive. This...  ...compute, memory, networking, and... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...unlimited potential of AI to define the next era...  ...team is building the software stack that makes large...  ...We own the platform — performance, CI/CD pipelines, validated...  ...efficiency on edge cluster configurationsProduce...  ...Science, Computer Engineering, Electrical Engineering... 
    Senior
    Performance
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...and powers the largest AI workloads worldwide...  ...for a hands-on Storage Software Engineer to join the storage...  ...keep our largest GPU clusters fast, reliable, and durable...  ...) — I/O and metadata performance, data corruption, and...  ...and operations, networking, and security, and... 
    Senior
    Performance
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $168k - $270.25k

    As a Senior Software Engineer on NVIDIA’s Global Network Visibility (GNV) team within NVIDIA's Global...  ...triage, and bring AI, storage, backbone, and...  ...configuration errors during cluster bring-upDefine a telemetry...  ...and clear tradeoffs among performance, resiliency, usability,... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $193.93k - $291.15k

     ...and profound opportunity for AI to drive positive change in the...  ...accelerators, and multi-cluster scheduling and orchestration....  ...Computer Science, Electrical Engineering, or a closely related field,...  ...the ability to reason about performance, failure modes, and reliability... 
    Senior
    Performance
    Work experience placement
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    4 days ago
  • $152k - $241.5k

     ...is hiring experienced software engineers with kubernetes experience...  ...to help scale up its AI Infrastructure. We...  ...kubernetes including cluster operations, operator development...  ...to cluster and network telemetry.Working with...  ...with maximum performance. Evaluating system failures... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $160k - $320k

     ...looking for a Senior Member of Technical...  ...of Kubernetes cluster management,...  ...backend engineering skills with deep...  .... It supports AI/ML workloads,...  ...with compute, networking, and storage systems...  ...professional software engineering...  ...GPU, or other performance-sensitive workloads... 
    Senior
    Performance
    Work at office
    Local area
    Remote work
    Relocation package
    3 days per week

    Nutanix

    San Jose, CA
    22 hours ago
  • $171k - $231.5k

     ...seeking a Staff Software Engineer to join the Core...  ...Kubernetes Services (IKS) clusters hosting Intuit...  ...design patterns, networking, troubleshooting,...  ...practices, and AI to develop...  ...mentoring junior and senior engineers, promoting...  ...a strong pay for performance rewards approach.... 
    Senior
    Performance
    Worldwide

    Intuit

    Mountain View, CA
    3 days ago
  • $152k - $241.5k

     ...GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production...  ...upgrades, repair, and cluster lifecycle management.Develop...  ...systems, CPU systems, networking, Linux, and Kubernetes;...  ...-X, or GPU cluster performance validation.Experience building... 
    Senior
    Performance
    Permanent employment
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Job Title: Senior Kubernetes Platform Engineer / Architect Location: Milpitas...  ...Kubernetes clusters. Develop container...  ...Kubernetes clusters. Perform cluster lifecycle...  ...namespaces, RBAC, network policies, storage...  ...Experience with AI/ML platform deployments... 
    Senior
    Performance

    Apptad Inc

    Milpitas, CA
    2 days ago
  •  ...NVIDIA Corporation seeks a Senior AI/ML Performance and Efficiency Engineer for GPU Clusters to advance AI efficiency across research workloads. You will partner with researchers to detect and fix infrastructure and application bottlenecks, delivering scalable improvements... 
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    1 day ago
  • $176k - $276k

     ...is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we...  ...most exciting computing hardware and software to contribute to the latest...  ...develop and bring up large scale performance platforms.What you will be doing:Design... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    22 hours ago
  • $184k - $287.5k

     ...Joining NVIDIA's DGX Cloud AI Efficiency Team means...  ...an AI infrastructure software engineer to join our team. You'...  ...of AI systems. As a senior DGX Cloud AI...  ...with the large scale clusters Experience in defining...  ...Artificial Intelligence, High-Performance Computing, and... 
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...tapping into the unlimited potential of AI to define the next era of...  ...world.We are seeking an outstanding Software Engineer to join our US-based networking software team. As a technical leader...  ...innovative, scalable, and high-performance hardware-accelerated software solutions... 
    Senior
    Performance
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $174k - $252k

     ...maintain and improve switch software.Manage individual project...  ....Manage scalability and performance tuning of large-scale networks.Triage product or system...  ....Google's software engineers develop the next-generation...  ...Contribute to projects enabling AI networking and high-... 
    Senior
    Performance

    Google

    Sunnyvale, CA
    4 days ago
  • $184k - $287.5k

    We are seeking a Senior Software Engineer to help build and improve AI Developer Tools connected through web APIs, SDKs, CLIs, and Agents. Apply your expertise...  ...CUDA development, from coding to profiling and performance fine-tuning.In this role, you will architect cloud... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA is looking for a Senior Software Engineer in Object Storage to design, implement, and extend the...  ...service that is critical to NVIDIA AI/ML research teams creating best of breed...  ...of dataAnalyzing and improving system performance at all levelsAutomating storage... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $264.5k

    NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure...  ...network to support NVIDIA software development workflows and...  ...-on experience with high performance network and network...  ...existing vacancy. NVIDIA uses AI tools in its recruiting... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    22 hours ago
  • $152k - $241.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An...  ...on the world. NVIDIA seeks a driven Software Engineer to advance our Kubernetes and AI Observability...  ...and development.Understanding of performance, security, and reliability in complex distributed... 
    Senior
    Performance
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer, AI Cluster Networking Performance Engineer. Be the first to apply!