Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Quality Engineer - GPU & Machine Learning

Advanced Micro Devices Inc

ADVANCE YOUR CAREER. ADVANCE THE WORLD.At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.THE ROLE: We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.THE PERSON:You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production. KEY RESPONSIBILITIES:Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms. Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them. Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage. Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.PREFERRED EXPERIENCE:Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster softwareExperience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.Familiarity with AI training, inference, and HPC benchmark methodologies.Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.ACADEMIC CREDENTIALS: BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience). LOCATION: San Jose, California#LI-DR1#LI-HYBRID Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Principal Software Quality Engineer - GPU & Machine Learning in San Jose, CA vacancy
  • $272k - $431.25k

    We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for GPU firmware and GPU system software, working directly with engineering...  ...influencing engineering teams to improve quality and fleet manageabilityWays to stand out... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $140k - $224.25k

     ..., building data driven tools to improve software quality, and ensuring customers have the best experience...  ...a creative, and hands-on software engineer with a test to failure approach who is a...  ...and optimize the testing workflows in GPU domain.Write maintainable, reliable, and... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $166.7k - $283.4k

     ...expert teams of physicists, engineers, data scientists and problem...  ...including several traditional machine learning techniques and deep learning...  ...Computing - HPC (including GPU), Machine Learning, Deep...  ...QualificationsSr. AI Infrastructure Software Engineer - C++ FocusLove C++... 
    Suggested
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    4 days ago
  • $139k - $208.4k

     ...ResponsibilitiesAs a Senior Engineer GPU Graphics Core Performance Verification...  ..., RTL, Verification, and Software teams to define performance...  ...innovation, and continuous learning while staying ahead of...  ...and research in graphics and machine learning, to shaping the GPU... 
    Suggested
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    1 day ago
  • $180k - $300k

     ...We are at the forefront of software and hardware innovation,...  ...RoleWe are looking for a Principal Software Engineer in QA to join our Software...  ...rigor and scalability to quality across the full stack — from...  ...Qualifications• Experience with machine learning frameworks and ML... 
    Suggested

    d-Matrix

    Santa Clara, CA
    3 days ago
  •  ...optimizing and developing deep learning frameworks for AMD...  ...critical in enhancing GPU kernels, deep learning...  ...and advanced engineering principles to drive continuous...  ...graph compilers. Software Engineering Best Practices...  ...from source to machine code. ACADEMIC CREDENTIALS... 

    AMD

    Santa Clara, CA
    4 days ago
  • $190k - $300k

     ...We are at the forefront of software and hardware innovation,...  ...within the US/Canada.The role: Principal Software Engineer, SDK & Lowering StackWhat...  ..., system software, and machine learning fundamentals.Demonstrated experience designing high-quality, scalable APIs and SDKs, with... 
    Work experience placement

    d-Matrix

    Santa Clara, CA
    4 days ago
  •  ...technology. We are at the forefront of software and hardware innovation, pushing the...  ...3+ days per week.The Role: Principal Software Engineer, KernelsWhat you will do:The role requires...  ...data structures, system software, and machine learning fundamentals.Proficient in C/C++ and... 
    Work experience placement
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $249k

     ...and build for travelers everywhere.Principal Software Engineer, Observability Introduction to the Team...  ...services, and tools to deliver high-quality experiences for travelers, partners,...  ...platform powered by data and machine learning provides secure, differentiated, and... 
    Full time

    Expedia

    San Jose, CA
    2 days ago
  •  ...the next generation of compiler and software infrastructure to accelerate Large...  ....THE PERSON:We are looking for a Principal Software Development Engineer to lead technical development in...  ...components and optimization pipelines for machine learning• Design and implement MLIR-based... 

    AMD

    San Jose, CA
    3 days ago
  • $190k - $210k

     ...California, United StatesProducts - Engineering /Fulltime /HybridOver 50,000...  ...a global networking leader, learn why there is no better time...  ...detailsTitle of position: Principal Software EngineerPosition type: Full...  ..., Electrical Engineering, Machine Learning, or a related... 
    Full time
    H1b
    Local area
    Work from home
    Work visa
    Shift work

    Extreme Networks, Inc.

    San Jose, CA
    3 days ago
  • $272k - $431.25k

     ...ship an Always-On, low-overhead GPU profiling service that runs...  ...-on delivery across system software, drivers, and CUDA to make profiling...  ...technical direction for an engineering team; mentor engineers, drive...  ...and shipping production quality system software or drivers with... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...performance and Python for extensibility, Dynamo orchestrates GPU shards, routes requests, and manages shared KV cache across...  ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap for memory management of... 
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

    We're looking for a Principal Engineer to join our CSP Engagements team as the technical focal point...  ...and stress tools (e.g., STREAM, GPU Burn, GPU BLAST) are updated and validated...  ...pattern analysis — identify configuration, software, or workload differences that explain... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...OPPORTUNITYWe're looking for a senior software engineer who combines deep systems performance work...  ...AI—someone who can shape software from GPU kernels through distributed training...  ...Extensive HIP/CUDA experience optimizing deep learning and OSS LLM inference/training kernels... 
    Shift work

    AMD

    Santa Clara, CA
    3 days ago
  •  ...ROLE:AMD is looking for an influential software engineer who is passionate about improving the performance...  ...performance from the lowest-level GPU kernels to large-scale distributed...  ...Supervised Fine-Tuning (SFT) and Reinforcement Learning (e.g., RLHF, GRPO). Candidates must... 

    AMD

    Santa Clara, CA
    1 day ago
  • $145.6k - $276.8k

     ...Senior Principal Software Engineer At RTX, the world's largest aerospace and defense company, 185,000 great minds are united by purpose...  ...clearance Qualifications We Prefer: Experience applying Machine Learning models for signal classification, modulation recognition... 
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Relocation
    Flexible hours

    Socket.dev

    San Jose, CA
    13 hours ago
  • Bachelor's or Master's degree in Computer science, Software Engineering, or related field. 10+ years of experience in Android mobile app...  ...sandboxing, code obfuscation, penetration testing). Exposure to data science or machine learning frameworks (TensorFlow, Pandas, NumPy, R).
    Full time

    Boston Scientific

    San Jose, CA
    13 hours ago
  • $170k - $277k

     ...SummaryIn the Layer-7 Security Software team, we are responsible for...  ...and Content Inspection Engine runs on Hardware, Virtualized...  ...such as the industry first Machine Learning powered NGFW, Credential Phishing...  ..., and interact with quality assurance and field support... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    9 hours ago
  • $258k - $387k

     ...investors.About the RoleAs a Principal Software Engineer, you will help define and...  ...building production-quality systems software.5+ years...  ...collection systems.Experience with GPU programming, NVIDIA...  ..., CUDA, image processing, machine learning infrastructure, or accelerated... 
    Immediate start
    Flexible hours

    Nuro

    Mountain View, CA
    2 days ago
  •  ...In the Layer-7 Security Software team, we are responsible for...  ...Identification and Content Inspection Engine runs on hardware,...  ...such as the industry first machine‑learning powered NGFW, credential phishing...  ...critical components in high quality and performance As an expert... 

    Palo Alto Networks, Inc.

    Santa Clara, CA
    3 days ago
  •  ...OneROCm — driving a unified ROCm software stack across AMD’s broad...  ...adaptability, and technical breadth to learn new domains, grow their...  .... Workload Performance Engineering: Lead the profiling, analysis...  ...PREFERRED EXPERIENCE: Knowledge in GPU architectures, basic... 

    AMD

    San Jose, CA
    2 days ago
  • $165.22k - $283.23k

     ...intelligence into the physical world. As a Principal Engineer, you will set the technical direction...  ...that blend cloud infrastructure, machine learning, and hardware interaction. You will...  ...Qualifications ~10+ years of professional software development experience ~... 
    Local area
    Immediate start

    Siemens

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...computer graphics, with our invention of the GPU. The GPU has also shown to be...  ...simulates human intelligence, running deep learning algorithms and acting as the brain of computers...  ...group is looking for Architects, Software Engineers, and AI application developers to join... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...s fueled by great technology—and amazing people.We are looking for a Principal Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA’s high-performance GPU infrastructure. You will play a meaningful role in crafting scalable... 
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $224k - $356.5k

     ...next era of computing. An era in which our GPU acts as the brains of computers, robots,...  ....NVIDIA's Local AI team is building the software stack that makes large language models...  ..., or PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    3 days ago
  • $249k

     ...and build for travelers everywhere.Principal Software Development Engineer — Platform &...  ...services, and tools to deliver high-quality experiences for travelers, partners...  ...technology platform powered by data and machine learning provides secure, differentiated, and... 
    Full time

    Expedia

    San Jose, CA
    3 days ago
  • $200k - $220k

     ...seek talented, passionate, and committed engineers, technologists, and business leaders to join...  ...is seeking an experienced AI Network Software Solution Architect to lead the design and...  ...workloads. This role requires deep expertise in GPU fabric design, high-speed switching,... 
    Worldwide

    Supermicro

    San Jose, CA
    4 days ago
  • $272k - $431.25k

    NVIDIA’s invention of the GPU in 1999 sparked the growth of the...  .... More recently, GPU deep learning ignited modern deep learning...  ...superchip. We are looking for expert engineers to come and help design rack...  ...of teamwork, love to produce quality work and commitment to finish... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $170k - $220k

     ...Description Job Description Staff Software Engineer, GPU AlgorithmsWe are looking for a full-...  ...software engineering role is to improve the quality, accuracy, and interpretation of...  ....Exploring the application of machine learning and artificial intelligence (AI) techniques... 
    Full time

    DeepSight Technology

    Santa Clara, CA
    11 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Quality Engineer - GPU & Machine Learning. Be the first to apply!