Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Systems Engineer - HPC

Full-time

Advanced Micro Devices Inc

ADVANCE YOUR CAREER. ADVANCE THE WORLD.At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms engineering team. The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON:You have a passion for learning. You are passionate about the field of large-scale distributed computing in AI and HPC workloads. You take responsibility for end-to-end outcomes of your efforts. You want to build scalable and highly performant HPC/AI/Data services with AMD hardware, software, people and processes. You have a curiosity to learn and improve scalable HPC systems. You have significant experience in working across a globally distributed organization. KEY RESPONSIBILITIES: Develop, implement, and maintain GPU-based clusters, ensuring optimal performance Administer ML/AI platforms – Distributed ML services, LLMs and AI inferencing, by managing deployments, resource allocation, monitoring, and security. Automate system provisioning and Cluster management end to end Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technical expertise. Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards. Use AI/ML to continuously improve internal processes and tools that are used in end-to-end delivery of your services in this team PREFERRED EXPERIENCE:Experience in developing Python based AI apps and UI HPC infrastructure engineering for AI/HPC domain SLURM and Kubernetes management Managing GPU clusters optimizing GPU-based services/tools/software Experience in creating web services with HPC backend (like AI) Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and Cluster interconnect with 400G networking. Demonstrated experience with AI workload schedulers and allocation optimization. Automation/monitoring tool - Ansible / Saltstack, Terraform, Prometheus, Grafana Strong organizational, problem-solving, and troubleshooting skills, with the ability to manage multiple projects simultaneously. Excellent verbal and written communication skills, with the ability to collaborate effectively with team members and stakeholders at all levels of the organization. EDUCATION: Bachelor's or master's degree in computer science or computer engineering preferred. LOCATION:San Jose, CAThis role is not eligible for visa sponsorship.#LI-MR1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 21 hours ago
Similar jobs that could be interesting for youBased on the AI Systems Engineer - HPC in San Jose, CA vacancy
  •  ...Ubuntu, etc.), Windows Server operating systems, Windows Client operating systems, and VMWare...  ...of current and next-generation HPE HPC products. Ensure development issues are...  ...appropriate automated test execution to test engineers at various global locations. Provide training... 
    Suggested
    Local area
    Remote work

    Net2Source

    San Jose, CA
    4 days ago
  •  ..., high-performance computing, cloud, and AI. Whether you’re designing next-gen processors...  ...forward.THE ROLE:We are seeking an AI Systems Engineer to join our AMD IT compute platforms...  ...administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI... 
    Suggested

    AMD

    San Jose, CA
    21 hours ago
  • $152k - $241.5k

     ...technology powers everything from generative AI to autonomous systems, and we continue to shape the future...  ...and tools that enable researchers and engineers to develop the next generation of AI/...  .... We are looking for a strong AI & HPC Observability Engineer to build and... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $176k - $276k

    NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building...  ...intelligence and GPU computing. Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs.... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $100k

    Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance...  ...-performance computing, and emerging autonomous AI systems. As the RISC-V AI / HPC & Agentic Software Engineering Lead, you will operate at the hardware-software boundary... 
    Suggested
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    4 days ago
  • $160k - $198k

     ...celebrates all of our team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and manage the critical infrastructure...  ...focus on AI/ML systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute Management: Familiarity... 
    Local area

    Archer Aviation

    San Jose, CA
    1 day ago
  •  ...advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day,...  .... Join us and, together, we’ll advance your career.HPC & AI Researcher Engineer - PhD preferredSOFTWARE SYSTEMS DESIGN ENGINEER The Role: AMD’s server software and... 

    AMD

    Santa Clara, CA
    21 hours ago
  • $190k - $237k

     ...celebrates all of our team members.What You’ll DoAs a Staff AI Systems Engineer, you will architect, deploy, and manage the critical infrastructure...  ...focus on AI/ML systems, high-performance computing (HPC), or ML infrastructure.Multi-Cloud & Compute Management: Familiarity... 
    Local area

    Archer Aviation

    San Jose, CA
    4 days ago
  • $76.92 - $95.19 per hour

     ...company works across air taxis, UAS, AI and powertrain development — building real-world systems that combine software...  ...aviation. We’re seeking exceptional engineers, operators and builders to join...  ...systems, high-performance computing (HPC), or ML infrastructure.Multi-... 
    Full time
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    1 day ago
  • $136.3k - $231.7k

     ...hands without us. KLA invents systems and solutions for the manufacturing...  ...expert teams of physicists, engineers, data scientists and problem-...  ...as part of their algorithms. AI, including several traditional...  ...class team of physicists, HPC system designers, machine learning... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    4 days ago
  • $190.2k - $360.5k

    The Opportunity We are looking for a Principal AI Systems Engineer with deep C++ expertise to help build the next generation of AI-enabled product and platform capabilities. This role sits at the intersection of large-scale systems engineering, applied AI, and production... 
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    San Jose, CA
    21 hours ago
  • $296k - $395k

     ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens...  ...day is currently Tuesday.Hardware Engineering at Lambda is responsible for...  ...this is that team.What You’ll DoOwn system integration validation for new HPC AI/ML, general purpose compute, storage... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    2 days ago
  • $148.7k - $201.2k

     ...performance across our product line. The Nitro Team is looking for engineers with systems knowledge and experience in area such as Linux OS boot...  ...performance computing workloads.The Nitro High Memory and HPC team owns the purpose built platform development for the High... 
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...’ll keep critically important systems running while working on the technologies...  .... You’ll harness the power of AI to deliver groundbreaking...  ...they integrate cleanly with HPC schedulers, storage, and...  ...Perl, or Ruby.Mentored other engineers and influenced technical direction... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...technologies: next-generation communications, artificial intelligence (AI), and semiconductors. Together, these initiatives position the...  ...production methods.  Job Summary: The AI Server/Rack System Engineer leads the end-to-end integration, configuration, and... 
    Local area

    Foxconn-PCE Technology

    Santa Clara, CA
    28 days ago
  • $136.3k - $231.7k

     ...into your hands without us. KLA invents systems and solutions for the manufacturing of wafers...  ...R&D. Our expert teams of physicists, engineers, data scientists and problem-solvers work...  ...real-time and high-performance computing (HPC) systems, including performance optimization... 
    Minimum wage
    Full time
    Work experience placement
    Flexible hours

    KLA-Tencor

    Milpitas, CA
    1 day ago
  • $149.75k - $275.58k

    Job Details:Job Description: Join Intel's AI Frameworks PyTorch team to shape the future of Artificial Intelligence (AI) and High-Performance Computing (HPC). As an AI Frameworks Engineer, you will design, build, and optimize AI software frameworks that enable breakthrough... 
    Full time
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency...  ...optimization experience with NSight Systems and NSight ComputeExperience with...  ...systems like Lustre and GPFS for AI/HPC workloadsFamiliarity with deep learning... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    We are looking for a software engineer with a strong background in parallel...  ...at the intersection of AI, high-performance computing, and...  ..., GPUs, and sophisticated systems, identifying and eliminating bottlenecks...  ..., and scale complex AI and HPC workloads for modern CPU and... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of...  ...are seeking a DevOps / Platform Engineer to join our team building and operating...  ...like Terraform.Background in HPC, Slurm, or GPU-based compute systems... 

    AMD

    San Jose, CA
    1 day ago
  • $139k - $257.55k

     ...management, brand consistency, reusable design systems, and collaboration workflows that empower...  ...team is exploring the next generation of AI-native creative systems that redefine how...  ....We are looking for forward-thinking engineers who are excited to explore ambiguous... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    3 days ago
  •  ...community where each individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the...  ...acceleration, designing scalable infrastructure, and monitoring system performance. Key ResponsibilitiesHigh-Performance AI... 
    Full time
    Local area
    Immediate start

    F5 Networks

    San Jose, CA
    21 hours ago
  • $170k - $234k

     ...the global leader in materials science and engineering solutions that are at the foundation of...  ...create and service is essential to advancing AI and accelerating the commercialization of...  ...materials designExperience using cloud/HPC environments for large-scale model... 
    Full time

    Applied Materials

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...Architect with a performance engineering background who can help our most...  ...customers accelerate Physical AI workloads using NVIDIA's full-...  ..., libraries, tools, and system software teams at NVIDIA to influence...  ...one of these areas: LLM and HPC. Having expertise ranging from... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...leading technology company for AI and Bitcoin mining...  ...AI Scheduling & Orchestration Engineer to lead the workload placement...  ...the intersection of distributed systems and AI, driving the architectural...  ...deadlock scenarios in large-scale HPC environments. Mentor junior... 
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    22 days ago
  • $184k - $287.5k

     ...the convergence of architecture, silicon, systems, and manufacturing. The System-...  ...Manufacturing & System Co-Design Workflow Engineer to lead the methodology and infrastructure...  ...workflows. This is the infrastructure that makes AI genuinely usable in a rigorous... 
    Full time
    Immediate start

    Nvidia

    Santa Clara, CA
    6 days ago
  •  ...communications, artificial intelligence (AI), and semiconductors. Together, these...  ...of the AI Server / Rack Architecture Engineer (focused on L10/L11 system-level execution) is to own the...  ...years of High-Performance Computing (HPC) or Enterprise Server Architecture... 
    Work at office
    Local area

    Foxconn-PCE Technology

    Santa Clara, CA
    14 days ago
  • $169k - $338k

     ...WalmartBusiness Segment: Home OfficePosition Summary...As a Distinguished AI/ML Engineer within Walmart Global Tech's Site Reliability Engineering...  ...lead the technical development of next-generation agentic AI systems and intelligent automation solutions that ensure mission-... 
    Full time
    Temporary work
    Part time

    Walmart

    Sunnyvale, CA
    1 day ago
  • Job DescriptionWe are seeking a System Application Engineer (SAE) to join our System Application Engineering team, supporting next‑generation...  ...semiconductor devices for customers in High Performance Computing (HPC) and AI‑driven applications.In this role, you will work closely... 

    Advantest

    San Jose, CA
    1 day ago
  •  ...generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of...  ...Datacenter GPU Platform Application Engineering team as a System Application...  ...Center GPU customers across Cloud, HPC, and OEM segments. In this... 

    AMD

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Systems Engineer - HPC. Be the first to apply!