AI Systems Engineer - HPC
Advanced Micro Devices Inc
ADVANCE YOUR CAREER. ADVANCE THE WORLD.At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms engineering team. The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON:You have a passion for learning. You are passionate about the field of large-scale distributed computing in AI and HPC workloads. You take responsibility for end-to-end outcomes of your efforts. You want to build scalable and highly performant HPC/AI/Data services with AMD hardware, software, people and processes. You have a curiosity to learn and improve scalable HPC systems. You have significant experience in working across a globally distributed organization. KEY RESPONSIBILITIES: Develop, implement, and maintain GPU-based clusters, ensuring optimal performance Administer ML/AI platforms – Distributed ML services, LLMs and AI inferencing, by managing deployments, resource allocation, monitoring, and security. Automate system provisioning and Cluster management end to end Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technical expertise. Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards. Use AI/ML to continuously improve internal processes and tools that are used in end-to-end delivery of your services in this team PREFERRED EXPERIENCE:Experience in developing Python based AI apps and UI HPC infrastructure engineering for AI/HPC domain SLURM and Kubernetes management Managing GPU clusters optimizing GPU-based services/tools/software Experience in creating web services with HPC backend (like AI) Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and Cluster interconnect with 400G networking. Demonstrated experience with AI workload schedulers and allocation optimization. Automation/monitoring tool - Ansible / Saltstack, Terraform, Prometheus, Grafana Strong organizational, problem-solving, and troubleshooting skills, with the ability to manage multiple projects simultaneously. Excellent verbal and written communication skills, with the ability to collaborate effectively with team members and stakeholders at all levels of the organization. EDUCATION: Bachelor's or master's degree in computer science or computer engineering preferred. LOCATION:San Jose, CAThis role is not eligible for visa sponsorship.#LI-MR1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
- ...Ubuntu, etc.), Windows Server operating systems, Windows Client operating systems, and VMWare... ...of current and next-generation HPE HPC products. Ensure development issues are... ...appropriate automated test execution to test engineers at various global locations. Provide training...SuggestedLocal areaRemote work
- ..., high-performance computing, cloud, and AI. Whether you’re designing next-gen processors... ...forward.THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms... ...administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI...Suggested
$152k - $241.5k
...technology powers everything from generative AI to autonomous systems, and we continue to shape the future... ...and tools that enable researchers and engineers to develop the next generation of AI/... .... We are looking for a strong AI & HPC Observability Engineer to build and...SuggestedFull time$176k - $276k
NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building... ...intelligence and GPU computing. Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs....SuggestedFull time$100k
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance... ...-performance computing, and emerging autonomous AI systems. As the RISC-V AI / HPC & Agentic Software Engineering Lead, you will operate at the hardware-software boundary...SuggestedPermanent employment$160k - $198k
...celebrates all of our team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage the critical infrastructure... ...focus on AI/ML systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute Management: Familiarity...Local area- ...advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day,... .... Join us and, together, we’ll advance your career.HPC & AI Researcher Engineer - PhD preferredSOFTWARE SYSTEMS DESIGN ENGINEER The Role: AMD’s server software and...
$190k - $237k
...celebrates all of our team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage the critical infrastructure... ...focus on AI/ML systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute Management: Familiarity...Local area$76.92 - $95.19 per hour
...company works across air taxis, UAS, AI and powertrain development — building real-world systems that combine software... ...aviation. We’re seeking exceptional engineers, operators and builders to join... ...systems, high-performance computing (HPC), or ML infrastructure.Multi-...Full timeLocal areaVisa sponsorshipNight shift$136.3k - $231.7k
...hands without us. KLA invents systems and solutions for the manufacturing... ...expert teams of physicists, engineers, data scientists and problem-... ...as part of their algorithms. AI, including several traditional... ...class team of physicists, HPC system designers, machine learning...Minimum wageFull timeWork experience placementFlexible hours$190.2k - $360.5k
The Opportunity We are looking for a Principal AI Systems Engineer with deep C++ expertise to help build the next generation of AI-enabled product and platform capabilities. This role sits at the intersection of large-scale systems engineering, applied AI, and production...Full timeTemporary workLocal areaRemote workWorldwide$296k - $395k
...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens... ...day is currently Tuesday.Hardware Engineering at Lambda is responsible for... ...this is that team.What You’ll DoOwn system integration validation for new HPC AI/ML, general purpose compute, storage...Work at officeLocal areaWork from homeFlexible hours$148.7k - $201.2k
...performance across our product line. The Nitro Team is looking for engineers with systems knowledge and experience in area such as Linux OS boot... ...performance computing workloads.The Nitro High Memory and HPC team owns the purpose built platform development for the High...InternshipLocal areaFlexible hours$152k - $241.5k
...’ll keep critically important systems running while working on the technologies... .... You’ll harness the power of AI to deliver groundbreaking... ...they integrate cleanly with HPC schedulers, storage, and... ...Perl, or Ruby.Mentored other engineers and influenced technical direction...Full time- ...technologies: next-generation communications, artificial intelligence (AI), and semiconductors. Together, these initiatives position the... ...production methods. Job Summary: The AI Server/Rack System Engineer leads the end-to-end integration, configuration, and...Local area
$136.3k - $231.7k
...into your hands without us. KLA invents systems and solutions for the manufacturing of wafers... ...R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work... ...real-time and high-performance computing (HPC) systems, including performance optimization...Minimum wageFull timeWork experience placementFlexible hours$149.75k - $275.58k
Job Details:Job Description: Join Intel's AI Frameworks PyTorch team to shape the future of Artificial Intelligence (AI) and High-Performance Computing (HPC). As an AI Frameworks Engineer, you will design, build, and optimize AI software frameworks that enable breakthrough...Full timeLocal areaImmediate startShift work$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency... ...optimization experience with NSight Systems and NSight ComputeExperience with... ...systems like Lustre and GPFS for AI/HPC workloadsFamiliarity with deep learning...Full timeRemote work$152k - $241.5k
We are looking for a software engineer with a strong background in parallel... ...at the intersection of AI, high-performance computing, and... ..., GPUs, and sophisticated systems, identifying and eliminating bottlenecks... ..., and scale complex AI and HPC workloads for modern CPU and...Full time- ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...are seeking a DevOps / Platform Engineer to join our team building and operating... ...like Terraform.Background in HPC, Slurm, or GPU-based compute systems...
$139k - $257.55k
...management, brand consistency, reusable design systems, and collaboration workflows that empower... ...team is exploring the next generation of AI-native creative systems that redefine how... ....We are looking for forward-thinking engineers who are excited to explore ambiguous...Full timeTemporary workLocal areaWorldwide- ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the... ...acceleration, designing scalable infrastructure, and monitoring system performance. Key ResponsibilitiesHigh-Performance AI...Full timeLocal areaImmediate start
$170k - $234k
...the global leader in materials science and engineering solutions that are at the foundation of... ...create and service is essential to advancing AI and accelerating the commercialization of... ...materials designExperience using cloud/HPC environments for large-scale model...Full time$184k - $287.5k
...Architect with a performance engineering background who can help our most... ...customers accelerate Physical AI workloads using NVIDIA's full-... ..., libraries, tools, and system software teams at NVIDIA to influence... ...one of these areas: LLM and HPC. Having expertise ranging from...Full timeRemote work- ...leading technology company for AI and Bitcoin mining... ...AI Scheduling & Orchestration Engineer to lead the workload placement... ...the intersection of distributed systems and AI, driving the architectural... ...deadlock scenarios in large-scale HPC environments. Mentor junior...Full timeLocal area
$184k - $287.5k
...the convergence of architecture, silicon, systems, and manufacturing. The System-... ...Manufacturing & System Co-Design Workflow Engineer to lead the methodology and infrastructure... ...workflows. This is the infrastructure that makes AI genuinely usable in a rigorous...Full timeImmediate start- ...communications, artificial intelligence (AI), and semiconductors. Together, these... ...of the AI Server / Rack Architecture Engineer (focused on L10/L11 system-level execution) is to own the... ...years of High-Performance Computing (HPC) or Enterprise Server Architecture...Work at officeLocal area
$169k - $338k
...WalmartBusiness Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering... ...lead the technical development of next-generation agentic AI systems and intelligent automation solutions that ensure mission-...Full timeTemporary workPart time- Job DescriptionWe are seeking a System Application Engineer (SAE) to join our System Application Engineering team, supporting next‑generation... ...semiconductor devices for customers in High Performance Computing (HPC) and AI‑driven applications.In this role, you will work closely...
- ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of... ...Datacenter GPU Platform Application Engineering team as a System Application... ...Center GPU customers across Cloud, HPC, and OEM segments. In this...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Systems Engineer - HPC. Be the first to apply!
- machine learning ai engineer San Jose, CA
- ai ml engineer San Jose, CA
- ai prompt engineer San Jose, CA
- ai engineer San Jose, CA
- ai engineer remote San Jose, CA
- ai developer San Jose, CA
- senior ai engineer San Jose, CA
- application system engineer San Jose, CA
- sr systems engineer San Jose, CA
- software system engineer San Jose, CA



