Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Software Developer - AI/ML Performance Validation & Systems Testing

Advanced Micro Devices Inc

ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE: We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.THE PERSON:You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production. KEY RESPONSIBILITIES:Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms. Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them. Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage. Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.PREFERRED EXPERIENCE:Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster softwareExperience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.Familiarity with AI training, inference, and HPC benchmark methodologies.Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.ACADEMIC CREDENTIALS: BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience). LOCATION: San Jose, California#LI-DR1#LI-HYBRID Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 17 days ago
Similar jobs that could be interesting for youBased on the Principal Software Developer - AI/ML Performance Validation & Systems Testing in San Jose, CA vacancy
  •  ...ServiceNow is the AI control tower...  ...cloud-native systems. The IC5...  ...for novel AI/ML challenges, establishing...  ...mentor and develop senior...  ...refactoring, testing...  ...Design high-performance data platforms...  ...organization Software Development Excellence...  ...Prototype and validate new... 
    Performance
    Full time
    Temporary work
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    4 days ago
  •  ...supercomputing, high-performance computing, cloud, and AI. Whether you’...  ...PMTS AI/ML Compiler...  ...join AMD's AI Software organization....  ...productsDesign, develop, and optimize...  ....Implement, validate, and maintain...  ...test suites.Lead performance...  ...critical software systems.Proficiency... 
    Performance

    AMD

    San Jose, CA
    13 hours ago
  •  ...potential of generative AI to power the...  ...the forefront of software and hardware innovation...  ...week.The role: Principal System Software Engineer,...  ...other software (ML and compilers) and...  ...distributed, high-performance software design and...  ...OpenMPIExperience with software testing... 
    Performance
    3 days per week

    d-Matrix

    Santa Clara, CA
    a month ago
  •  ...Nexxa is building the best AI systems for heavy industries —...  ...role is a blend of backend software engineering, ML infrastructure, and systems...  .... Own the reliability, performance, and observability of backend...  ...— logging, monitoring, testing, and CI/CD for ML services... 
    Performance
    Full time

    Nexxa.ai

    Sunnyvale, CA
    13 hours ago
  •  ...for physical AI — a unified...  ...and deploy as software. Today,...  ...platform for developers, researchers...  ...to-end: from system design through...  ...output validation, sandboxed tool...  ...implementation, testing, rollout,...  ...testable code, performance awareness,...  ...for AI/ML integration... 
    Performance
    Full time

    Dexmate

    Santa Clara, CA
    13 hours ago
  • $150k - $210k

     ...Position – Embodied AI EngineerWe are...  ...pipelines, and training systems that let research...  ...startup teams to co-develop hardware–software systems and validate them in real-world...  ...source robotics or ML infrastructure tooling...  ...and vacation time. Performance based Short-Term... 
    Performance
    Full time
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    Santa Clara, CA
    20 hours ago
  •  ...AI Applications EngineerAt AMD,...  ...opportunity to develop deep expertise...  ...design, embedded systems, and AI infrastructure...  ...hardware and software domains. You...  ...in AI/ML, programming, and...  ...the creation, validation, and optimization...  ...GPUs, balancing performance, latency, power... 
    Performance

    Advanced Micro Devices , Inc.

    San Jose, CA
    3 days ago
  • $120.75k - $161k

     ...generation of Agentic AI products that...  ...looking for a Software Engineer to...  ...and backend systems. In this role...  ...designers, and AI/ML engineers to...  ...based workflows Develop backend...  ...maintainable, and well-tested code...  ...improve system performance, reliability, and... 
    Performance
    Full time
    Work at office
    Remote work
    Flexible hours

    Eightfold

    Santa Clara, CA
    13 hours ago
  •  ...experiences—from AI and data...  ...and embedded systems. Grounded in...  ...hardware and software engineering teams...  ...candidates, validate correctness,...  ...simulation, firmware, performance debugging,...  ...accuracy.Develop tools that...  ...validation using tests, benchmarks,...  ...applied AI, ML, agentic,... 
    Performance

    AMD

    Santa Clara, CA
    a month ago
  • $140k - $165k

     ...semiconductor innovation, developing advanced memory...  ...enhanced performance and user experiences...  ...foundational AI infrastructure that...  ...next-gen enterprise systems. Work on...  ...engineering and applied ML — you'll bridge...  ...AI to automate testing or validation. Familiarity with... 
    Performance

    SK hynix memory solutions America Inc.

    San Jose, CA
    28 days ago
  • Distributed Systems Software Engineer, Python / Go...  ...for building and validating resilient distributed...  ...approach to test automation, reporting...  ...opportunity to develop CI pipelines which...  ...and developing AI/ML pipelines for automatic...  ...reliability, performance, and resilience... 
    Performance
    Full time
    Local area
    Remote work
    Worldwide

    Canonical

    San Jose, CA
    2 days ago
  • $165.2k - $223.6k

     ...about designing, developing, and...  ...the future of AI-driven cloud...  ...are seeking a Software Development Engineer...  ...security, AI/ML, distributed systems, emerging AWS...  ...and validate solutions - Architect...  ..., performance, and customer...  ...through load testing, and collaborate... 
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    a month ago
  • $2,500 per month

     ...chips, racks, software, and manufacturing...  ...Silicon Validation Firmware Engineer to develop the firmware and...  ...validates our AI/ML accelerator ASICs...  ...validation firmware, test content, and...  ...functional correctness, performance, and...  ...protocol, and system levels; root-cause... 
    Performance
    Work at office
    Relocation package

    Etched

    San Jose, CA
    2 days ago
  • $224k - $356.5k

     ...introductions (NPIs), distributed systems, familiarity with software testing and deployment, and...  ...to join the EDA Team. Principal Software Engineer...  ...stand out from the crowd:Developing ML/AI infrastructure....  ...Artificial Intelligence, High-Performance Computing, and... 
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $123.24k - $200k

     ...Role As a Sr./Principal AI Engineer...  ...foundational systems from the ground...  ...and Validate: Lead rapid...  ..., automated testing, and observability...  ...hunting copilots, developer productivity...  ...in software engineering,...  ...fields in high‑performance environments...  ...SageMaker, Azure ML) and... 
    Performance
    Work at office

    TSMC

    San Jose, CA
    5 days ago
  • $203k - $258.6k

     ...26Meet the TeamCX AI Incubation team is...  ...you will design and develop transformative AI...  ..., network test automation, infrastructure...  ...expertise in AI/ML, software development, and a...  ...Analyze data, develop, validate, and deploy...  ...to improve model performance, scalability, and... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    19 days ago
  •  ...applying Generative AI and traditional...  ...science and software engineering,...  ...Development: Design, develop, and fine-tune...  ...(RAG) systems to enhance model...  ...wide range of ML models (classification...  ...training, validation, and deployment...  ...improved model performance. Collaboration... 
    Performance

    Omni Inclusive

    San Jose, CA
    5 days ago
  • Lead AI/ML Engineer We are seeking a...  ...of enterprise software engineering, distributed...  ...) for robust testing. · Construct local developer sandbox...  ...Pipeline Automation & System Architecture ·...  ...& Trajectory Validation · Author eval...  ...experience in LLM performance telemetry,... 
    Performance
    Local area

    E-Solutions

    San Jose, CA
    5 days ago
  • $272k - $431.25k

     ...serving generative AI and reasoning...  ...Built in Rust for performance and Python for extensibility...  ...like a single system at datacenter...  ....We are seeking a Principal Systems Engineer to...  ...performance storage, or ML systems...  ...architectural decisions and validate improvements in... 
    Performance
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    a month ago
  • $223.1k - $301.1k

     ...applications that validate customer networks...  ...the integrity and performance of enterprise infrastructure...  ...of our global, AI-driven team, you...  ...the frontier of software innovation,...  ...automate network testing, validation, and lifecycle...  ...backend systems that guarantee network... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    13 days ago
  • $142.8k - $274.8k

     ...25%Profession: Software EngineeringDiscipline...  ...the end-to-end AI stack and is...  ...Service, Azure ML, Cognitive...  ...looking for a Principal Software Engineer...  ...and with high performance, low latency, and...  ...ResponsibilitiesDesign, and develop large-scale...  ...distributed systems.4+ years of... 
    Performance
    Ongoing contract
    Work at office
    Local area
    3 days per week

    Microsoft

    Mountain View, CA
    a month ago
  • $231.4k - $331.8k

     ...central to the AI era, powering...  ...and develop BIOS, BSP, and...  ...develop, and test device drivers...  ...and execute software test plans. Collaborate...  ...and validate software. Innovate...  ...for embedded systems. Experience in...  ...systems. AI/ML experience and...  ...plans earn performance-based incentive... 
    Performance
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Milpitas, CA
    a month ago
  • $249k

     ...travelers everywhere. Principal Data & AI Engineer, Reporting...  ..., portfolio, and performance data into trusted...  ...hallucination mitigation, validation, accuracy monitoring,...  ...Science, Information Systems, Data Science,...  ...development, analytics, and AI/ML integration. ~3+... 
    Performance
    Work at office

    Expedia Group

    San Jose, CA
    19 days ago
  • $180k - $300k

     ...generative AI to power the...  ...of software and hardware...  ...looking for a Principal Software Engineer...  ...for test strategy and...  ...product, and systems teams to design...  ...and ML workload execution...  ...GitLab.• Develop and...  ..., reliable validation of workloads...  ...and company performance. This is in... 
    Performance

    d-Matrix

    Santa Clara, CA
    a month ago
  • $250.8k - $286.2k

    Senior Lead AI Engineer (MLX Emerging AI...  ...responsible and reliable AI systems, changing banking...  ...of AI & ML are bringing humanity...  ...and scalable, high-performance AI infrastructure....  ...One. Design, develop, test, deploy, and support AI software components including... 
    Performance
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Jose, CA
    13 hours ago
  •  ...: At Viven, we're building AI-powered Digital Twins for businesses...  ...:  We are looking for an AI/ML Engineer with hands-on...  ...Viven’s products, with a focus on performance, scalability, and reliability....  ...rapidly into real-world systems. Key Responsibilities:  Design... 
    Performance
    Full time

    Careers

    Santa Clara, CA
    13 hours ago
  • $207k - $340k

     ...meaning it will be performed both from home...  ...team.LinkedIn's AI and Machine Learning...  ...scientists and software engineers, who develop and implement machine...  .... As a Principal Staff Software Engineer...  ...algorithms, models, and systems that power our...  ...who will use AI/ML to push the... 
    Performance
    For contractors
    Work experience placement
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    a month ago
  •  ...servers powering AI/ML training and inference...  ..., and drive validation from PCBA bring-up...  ...that enable high-performance AI training and inference...  ..., firmware, test, qualification, and...  ...telemetry to identify systemic issues and drive...  ...Collaborate with firmware, software, and operations... 
    Performance

    Amazon

    Cupertino, CA
    6 days ago
  •  ...and build solutions — combining AI, data, and security expertise...  ...Expertise ~3+ years in software engineering, security engineering...  ...~ Familiarity with AI/ML tooling, data pipelines, or data...  ...thrive in a fast-paced, high-performance startup environment Passion... 
    Performance

    TenEx

    San Jose, CA
    2 days ago
  •  ...experiences—from AI and data centers,...  ...gaming and embedded systems. Grounded in a culture...  ...compiler for high-performance GPU kernels,...  ...critical to AMD’s AI software roadmap.AMD GPUs...  ...Gluon kernels for ML kernels powering the...  ...engineers to develop and maintain the Triton... 
    Performance

    AMD

    San Jose, CA
    28 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Software Developer - AI/ML Performance Validation & Systems Testing. Be the first to apply!