Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Developer - AI/ML Performance Validation & Systems Testing

Full-time

Advanced Micro Devices Inc

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE: We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.THE PERSON:You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production. KEY RESPONSIBILITIES:Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms. Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them. Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage. Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.PREFERRED EXPERIENCE:Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster softwareExperience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.Familiarity with AI training, inference, and HPC benchmark methodologies.Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.ACADEMIC CREDENTIALS: BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience). LOCATION: San Jose, California#LI-DR1#LI-HYBRID Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Principal Software Developer - AI/ML Performance Validation & Systems Testing in San Jose, CA vacancy
  •  ...to powering AI and the technologies...  ...seeking a Principal Software Engineer to...  ...software validation across...  ...server-class systems. In this individual...  ...component testing, through...  ..., workload, performance, stress, stability...  ...-service developer pre-submit...  ...oneAPI, SYCL)AI/ML frameworks (... 
    Performance
    Contract work
    Shift work

    AMD

    San Jose, CA
    4 days ago
  •  ...supercomputing, high-performance computing, cloud, and AI. Whether you’...  ...PMTS AI/ML Compiler...  ...join AMD's AI Software organization....  ...productsDesign, develop, and optimize...  ....Implement, validate, and maintain...  ...test suites.Lead performance...  ...critical software systems.Proficiency... 
    Performance

    AMD

    San Jose, CA
    2 days ago
  • $174k - $252k

    Write and test production software for agentic validation platforms, closed-loop execution...  ...complex product or system issues across web...  ...in specialized ML areas, leveraging ML...  ...to client systems, developer infrastructure, or...  ...Espresso), or closed-loop AI agent toolchains.... 
    Suggested
    Shift work

    Google

    San Jose, CA
    21 hours ago
  • $224k - $356.5k

     ...introductions (NPIs), distributed systems, familiarity with software testing and deployment, and...  ...to join the EDA Team. Principal Software Engineer...  ...stand out from the crowd:Developing ML/AI infrastructure....  ...Artificial Intelligence, High-Performance Computing, and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...potential of generative AI to power the...  ...the forefront of software and hardware innovation...  ...week.The role: Principal System Software Engineer,...  ...other software (ML and compilers) and...  ...distributed, high-performance software design and...  ...OpenMPIExperience with software testing... 
    Performance
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  •  ...experiences—from AI and data centers,...  ...gaming and embedded systems. Grounded in a culture...  ...a unified ROCm software stack across AMD’s...  ..., and help develop next-generation products...  ..., frameworks, and performance optimization layers...  ...software, AI/ML frameworks, libraries... 
    Performance

    AMD

    San Jose, CA
    2 days ago
  • $124k - $195.5k

     ...Intelligence, High-Performance Computing, and...  ...generative AI to autonomous vehicles...  ...looking for a Software Engineer to...  ...most advanced ML models on some...  ...most powerful GPU systems.What You'll Be...  ...pain points of validating, monitoring and...  ...will design, develop and maintain engineering... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $163k - $253k

     ...licensing and R&D software. In this role,...  ...at scale.AI/PI Group is an...  ...learning and system engineering to develop and operate our...  ...generation AI/ML solutions.What...  ...Optimize inference performance on accelerated...  ...(A/B testing, canary releases...  ...a current and valid agreement with... 
    Performance
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    1 day ago
  •  ...experiences—from AI and data...  ...and embedded systems. Grounded in...  ...AI systems, performance engineering,...  ...and agentic software development....  ...generation, automated validation, measurable...  ...workflows.Develop systems that...  ...generate, compile, test, benchmark,...  ...building AI, ML, agentic,... 
    Performance

    AMD

    Santa Clara, CA
    4 days ago
  • $203k - $258.6k

     ...part of the Cisco AI and Automation...  ...optimized for performance and cost-...  ...Engineer, you will develop software consistent with...  ...Computer Science, AI/ML, or a related...  ...on building, testing, and deploying...  ...Code) to build, validate, and deploy...  ...specifically applying system-level design to... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours
    3 days per week

    CISCO Systems

    San Jose, CA
    1 day ago
  • $150k - $210k

     ...Position - Embodied AI EngineerWe are...  ...pipelines, and training systems that let research...  ...startup teams to co-develop hardware-software systems and validate them in real-world...  ...source robotics or ML infrastructure tooling...  ...and vacation time. Performance based Short-Term... 
    Performance
    Full time
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    Santa Clara, CA
    21 hours ago
  •  ...experiences—from AI and data...  ...and embedded systems. Grounded in...  ...hardware and software engineering teams...  ...candidates, validate correctness,...  ...simulation, firmware, performance debugging,...  ...accuracy.Develop tools that...  ...validation using tests, benchmarks,...  ...applied AI, ML, agentic,... 
    Performance

    AMD

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...is seeking an Applied AI Engineer to innovate, develop, and integrate...  ...be doing:LLM-Powered Validation Pipelines: Design and deploy AI systems that make post-silicon...  ...of AI impact, close performance gaps, and drive iteration...  ...building and deploying ML/AI systems or data-intensive... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $165.2k - $223.6k

     ...professional software development experience...  ...in system design or architecture...  ...of system performance, memory...  ...servers Design, develop, and optimize...  ...on custom ML accelerators...  ..., testing, and production...  ...to-end model validation, with continuous...  ...Technologies: AI AWS Hardware... 
    Performance
    Full time
    Internship

    Annapurna Labs Inc.

    Cupertino, CA
    6 days ago
  • $165.2k - $223.6k

     ...about designing, developing, and...  ...the future of AI-driven cloud...  ...are seeking a Software Development Engineer...  ...security, AI/ML, distributed systems, emerging AWS...  ...and validate solutions - Architect...  ..., performance, and customer...  ...through load testing, and collaborate... 
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...unlimited potential of AI to define the next...  ...is building the software stack that makes large...  ...own the platform — performance, CI/CD pipelines, validated recipes, and model...  ...— that lets developers run groundbreaking...  ...in GPU computing, ML systems, or high-performance... 
    Performance
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...discovery to powering AI and the...  ...opportunity to develop deep expertise...  ...design, embedded systems, and AI infrastructure...  ...hardware and software domains. You...  ...in AI/ML, programming, and...  ...the creation, validation, and optimization...  ...GPUs, balancing performance, latency, power... 
    Performance

    AMD

    San Jose, CA
    14 hours ago
  • $171.6k - $257.4k

     ...can thrive. AI Engineer - Customer...  ...the core ML/AI systems that power self-...  ...evaluation, fine‑tuning, validation, deployment and...  ...SLAs); write performant, well‑tested code (primarily...  ...Development: Design, develop, and maintain...  ...~10+ years software engineering experience... 
    Performance
    Local area

    Worky Ltd

    San Jose, CA
    21 hours ago
  • $272k - $431.25k

     ...looking for a Principal Engineer...  ...-to-end performance, working directly...  ..., and validate that...  ...and drive systemic improvements...  ...configuration, software, or...  ...milestonesDefine test strategies...  ...GPU/HPC/ML...  ...groundbreaking developments in Artificial...  ...NVIDIA uses AI tools in... 
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...serving generative AI and reasoning...  ...Built in Rust for performance and Python for extensibility...  ...like a single system at datacenter...  ....We are seeking a Principal Systems Engineer to...  ...performance storage, or ML systems...  ...architectural decisions and validate improvements in... 
    Performance
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $132.4k - $217.6k

     ...600.00Bio-Techne develops innovative software and instrumentation...  ...we are embedding AI-driven...  ...production-grade systems used in regulated...  ...observability, security, and performance of distributed...  ...Computer Science, AI/ML, or related...  ...training/tuning, validation, deployment, and... 
    Performance
    Full time
    Temporary work
    Internship
    Worldwide
    Flexible hours

    Bio-Techne

    San Jose, CA
    4 days ago
  • $270k - $340k

     ....What You’ll Do:As a Principal AI and ML fundamentalist who is an expert at developing cutting-edge AI solutions...  ...to prototype and validate complex solutions from...  ...production-ready systems.What You Need:M.S or...  .... We drive a pay-for-performance culture and reward performance... 
    Performance
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...ignited modern AI—the next era...  ...we need a principal-level, hands...  ...production systems and the architectural...  ..., performance, observability...  ...for testing, debugging,...  ...like mature software, not prototypes...  ...integration.Help validate and operationalize...  ..., and developer tooling.Experience... 
    Performance
    Full time
    Live in

    Nvidia

    Santa Clara, CA
    3 days ago
  • $203k - $258.6k

     ...26Meet the TeamCX AI Incubation team is...  ...you will design and develop transformative AI...  ..., network test automation, infrastructure...  ...expertise in AI/ML, software development, and a...  ...Analyze data, develop, validate, and deploy...  ...to improve model performance, scalability, and... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $223.1k - $301.1k

     ...applications that validate customer networks...  ...the integrity and performance of enterprise infrastructure...  ...of our global, AI-driven team, you...  ...the frontier of software innovation,...  ...automate network testing, validation, and lifecycle...  ...backend systems that guarantee network... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    1 day ago
  • $231.4k - $331.8k

     ...central to the AI era, powering...  ...and develop BIOS, BSP, and...  ...develop, and test device drivers...  ...and execute software test plans. Collaborate...  ...and validate software. Innovate...  ...for embedded systems. Experience in...  ...systems. AI/ML experience and...  ...plans earn performance-based incentive... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Milpitas, CA
    1 day ago
  • $142.8k - $274.8k

     ...25%Profession: Software EngineeringDiscipline...  ...the end-to-end AI stack and is...  ...Service, Azure ML, Cognitive...  ...looking for a Principal Software Engineer...  ...and with high performance, low latency, and...  ...ResponsibilitiesDesign, and develop large-scale...  ...distributed systems.4+ years of... 
    Performance
    Ongoing contract
    Work at office
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    21 hours ago
  •  ...applying Generative AI and traditional...  ...science and software engineering,...  ...Development: Design, develop, and fine-tune...  ...(RAG) systems to enhance model...  ...wide range of ML models (classification...  ...training, validation, and deployment...  ...improved model performance.... 
    Performance

    Omni Inclusive

    San Jose, CA
    21 hours ago
  • $184k - $287.5k

     ...the unlimited potential of AI to define the next era of...  ...at the intersection of ML infrastructure and large-scale systems, this is your opportunity...  ...architecture, build, methodology, validation, and applied AI teams....  ...security, reliability, performance, and evolution.Leading... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Lead AI/ML Engineer We are seeking...  ...of enterprise software engineering, distributed...  ...) for robust testing. · Construct local developer sandbox...  ...Pipeline Automation & System Architecture...  ...& Trajectory Validation · Author eval...  ...in LLM performance telemetry, prompt... 
    Performance
    Local area

    E-Solutions

    San Jose, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Developer - AI/ML Performance Validation & Systems Testing. Be the first to apply!