Principal Software Developer - AI/ML Performance Validation & Systems Testing
Advanced Micro Devices Inc
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.THE ROLE: We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.THE PERSON:You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production. KEY RESPONSIBILITIES:Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms. Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them. Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage. Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.PREFERRED EXPERIENCE:Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster softwareExperience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.Familiarity with AI training, inference, and HPC benchmark methodologies.Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.ACADEMIC CREDENTIALS: BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience). LOCATION: San Jose, California#LI-DR1#LI-HYBRID Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
- ...supercomputing, high-performance computing, cloud, and AI. Whether you’... ...PMTS AI/ML Compiler... ...join AMD's AI Software organization.... ...productsDesign, develop, and optimize... ....Implement, validate, and maintain... ...test suites.Lead performance... ...critical software systems.Proficiency...Performance
$174k - $252k
Write and test production software for agentic validation platforms, closed-loop execution... ...complex product or system issues across web... ...in specialized ML areas, leveraging ML... ...to client systems, developer infrastructure, or... ...Espresso), or closed-loop AI agent toolchains....SuggestedShift work$224k - $356.5k
...introductions (NPIs), distributed systems, familiarity with software testing and deployment, and... ...to join the EDA Team. Principal Software Engineer... ...stand out from the crowd:Developing ML/AI infrastructure.... ...Artificial Intelligence, High-Performance Computing, and...PerformanceFull time- ...potential of generative AI to power the... ...the forefront of software and hardware innovation... ...week.The role: Principal System Software Engineer,... ...other software (ML and compilers) and... ...distributed, high-performance software design and... ...OpenMPIExperience with software testing...Performance3 days per week
- ...experiences—from AI and data centers,... ...gaming and embedded systems. Grounded in a culture... ...a unified ROCm software stack across AMD’s... ..., and help develop next-generation products... ..., frameworks, and performance optimization layers... ...software, AI/ML frameworks, libraries...Performance
$189k - $301k
...solving the complex system-level challenges... ...demands of future AI/ML workloads. Our team... ...to designing and developing scalable platforms... ...consumption and maximizing performance. To achieve this... ...both hardware and software engineers to... ...a current and valid agreement with Samsung...PerformanceWork at officeFlexible hours$163k - $253k
...licensing and R&D software. In this role,... ...at scale.AI/PI Group is an... ...learning and system engineering to develop and operate our... ...generation AI/ML solutions.What... ...Optimize inference performance on accelerated... ...(A/B testing, canary releases... ...a current and valid agreement with...PerformanceFlexible hours- ...experiences—from AI and data... ...and embedded systems. Grounded in... ...AI systems, performance engineering,... ...and agentic software development.... ...generation, automated validation, measurable... ...workflows.Develop systems that... ...generate, compile, test, benchmark,... ...building AI, ML, agentic,...Performance
$150k - $210k
...Position - Embodied AI EngineerWe are... ...pipelines, and training systems that let research... ...startup teams to co-develop hardware-software systems and validate them in real-world... ...source robotics or ML infrastructure tooling... ...and vacation time. Performance based Short-Term...PerformanceFull timeTemporary workFor contractorsLocal areaImmediate start- ...experiences—from AI and data... ...and embedded systems. Grounded in... ...hardware and software engineering teams... ...candidates, validate correctness,... ...simulation, firmware, performance debugging,... ...accuracy.Develop tools that... ...validation using tests, benchmarks,... ...applied AI, ML, agentic,...Performance
$152k - $241.5k
...is seeking an Applied AI Engineer to innovate, develop, and integrate... ...be doing:LLM-Powered Validation Pipelines: Design and deploy AI systems that make post-silicon... ...of AI impact, close performance gaps, and drive iteration... ...building and deploying ML/AI systems or data-intensive...PerformanceFull timeRemote work$165.2k - $223.6k
...about designing, developing, and... ...the future of AI-driven cloud... ...are seeking a Software Development Engineer... ...security, AI/ML, distributed systems, emerging AWS... ...and validate solutions - Architect... ..., performance, and customer... ...through load testing, and collaborate...PerformanceInternshipLocal areaFlexible hours$224k - $356.5k
...unlimited potential of AI to define the next... ...is building the software stack that makes large... ...own the platform — performance, CI/CD pipelines, validated recipes, and model... ...— that lets developers run groundbreaking... ...in GPU computing, ML systems, or high-performance...PerformanceFull timeLocal area- ...discovery to powering AI and the... ...opportunity to develop deep expertise... ...design, embedded systems, and AI infrastructure... ...hardware and software domains. You... ...in AI/ML, programming, and... ...the creation, validation, and optimization... ...GPUs, balancing performance, latency, power...Performance
$270k - $340k
....What You’ll Do:As a Principal AI and ML fundamentalist who is an expert at developing cutting-edge AI solutions... ...to prototype and validate complex solutions from... ...production-ready systems.What You Need:M.S or... .... We drive a pay-for-performance culture and reward performance...PerformanceLocal area$203k - $258.6k
...26Meet the TeamCX AI Incubation team is... ...you will design and develop transformative AI... ..., network test automation, infrastructure... ...expertise in AI/ML, software development, and a... ...Analyze data, develop, validate, and deploy... ...to improve model performance, scalability, and...PerformanceFull timeTemporary workLocal areaFlexible hours$132.4k - $217.6k
...600.00Bio-Techne develops innovative software and instrumentation... ...we are embedding AI-driven... ...production-grade systems used in regulated... ...observability, security, and performance of distributed... ...Computer Science, AI/ML, or related... ...training/tuning, validation, deployment, and...PerformanceFull timeTemporary workInternshipWorldwideFlexible hours$148.7k - $297.3k
...Vascular is seeking a Principal AI/ML Engineer to develop advanced machine... ...intravascular imaging systems. This role will... ...into scalable, high-performance solutions integrated... ...clinical, imaging, software, and systems... ...development, training, validation, performance optimization...Performance$272k - $431.25k
...looking for a Principal Engineer... ...-to-end performance, working directly... ..., and validate that... ...and drive systemic improvements... ...configuration, software, or... ...milestonesDefine test strategies... ...GPU/HPC/ML... ...groundbreaking developments in Artificial... ...NVIDIA uses AI tools in...PerformanceFull timeRemote work$272k - $431.25k
...serving generative AI and reasoning... ...Built in Rust for performance and Python for extensibility... ...like a single system at datacenter... ....We are seeking a Principal Systems Engineer to... ...performance storage, or ML systems... ...architectural decisions and validate improvements in...PerformanceFull timeLocal areaRemote work$272k - $431.25k
...ignited modern AI—the next era... ...we need a principal-level, hands... ...production systems and the architectural... ..., performance, observability... ...for testing, debugging,... ...like mature software, not prototypes... ...integration.Help validate and operationalize... ..., and developer tooling.Experience...PerformanceFull timeLive in$142.8k - $274.8k
...25%Profession: Software EngineeringDiscipline... ...the end-to-end AI stack and is... ...Service, Azure ML, Cognitive... ...looking for a Principal Software Engineer... ...and with high performance, low latency, and... ...ResponsibilitiesDesign, and develop large-scale... ...distributed systems.4+ years of...PerformanceOngoing contractWork at officeLocal area3 days per week$231.4k - $331.8k
...central to the AI era, powering... ...and develop BIOS, BSP, and... ...develop, and test device drivers... ...and execute software test plans. Collaborate... ...and validate software. Innovate... ...for embedded systems. Experience in... ...systems. AI/ML experience and... ...plans earn performance-based incentive...PerformanceFull timeTemporary workLocal areaFlexible hours$184k - $287.5k
...the unlimited potential of AI to define the next era of... ...at the intersection of ML infrastructure and large-scale systems, this is your opportunity... ...architecture, build, methodology, validation, and applied AI teams.... ...security, reliability, performance, and evolution.Leading...PerformanceFull time$123.24k - $200k
...Senior / Principal AI Engineer for Business... ...systems from the ground... ...Prototype and Validate: Lead rapid... ..., automated testing, and observability... ...copilots, developer productivity... ...in software engineering,... ...fields in high-performance environments... ...SageMaker, Azure ML) and...PerformanceWork at office$184k - $287.5k
...Architect with a performance engineering... ...accelerate Physical AI workloads using... ...and test of Autonomous Vehicles... ...NVIDIA hardware and software. If you are... ...solving. You will develop and improve... ...libraries, tools, and system software teams... ...of hands-on validated ML/DL performance...PerformanceFull timeRemote work- ...supercomputing, high-performance computing, cloud, and AI. Whether you’re... ...is seeking an AI Systems Engineer to help develop and optimize machine... ...of hardware and software, designing high-performance ML operator kernels,... ...through hardware validation and silicon bring‑...PerformanceWorldwide
$180k - $300k
...generative AI to power the... ...of software and hardware... ...looking for a Principal Software Engineer... ...for test strategy and... ...product, and systems teams to design... ...and ML workload execution... ...GitLab.• Develop and... ..., reliable validation of workloads... ...and company performance. This is in...Performance$249k
...travelers everywhere.Principal Data & AI Engineer, Reporting and... ..., portfolio, and performance data into trusted executive... ...mitigation, validation, accuracy monitoring,... ...Science, Information Systems, Data Science, Business... ...development, analytics, and AI/ML integration. 3+ years...PerformanceFull timeWork at office$207k - $340k
...meaning it will be performed both from home... ...team.LinkedIn's AI and Machine Learning... ...scientists and software engineers, who develop and implement machine... .... As a Principal Staff Software Engineer... ...algorithms, models, and systems that power our... ...who will use AI/ML to push the...PerformanceFor contractorsWork experience placementWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Software Developer - AI/ML Performance Validation & Systems Testing. Be the first to apply!
- principal software engineer San Jose, CA
- senior principal cloud computing engineer San Jose, CA
- principal architect San Jose, CA
- principal San Jose, CA
- principal cloud computing engineer San Jose, CA
- senior principal scientist San Jose, CA
- software implementation project manager San Jose, CA
- remote software sales San Jose, CA
- bank software San Jose, CA
- healthcare software sales San Jose, CA

