AI Cluster & Data Center Design Engr
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a highly skilled systems engineer to architect and design scalable AI/HPC clusters. This role involves evaluating and selecting compute, storage, networking, and power delivery components and solutions to optimize performance and reliability across global deployments. You will collaborate with cross-functional teams to deliver cutting-edge infrastructure for AI and high-performance computing workloads.THE PERSONAn experienced systems architect with a strong background in HPC, AI infrastructure, cluster, and data center engineering. You bring deep technical knowledge of compute, power, and networking components, a strategic mindset for system-level design, and the ability to collaborate across diverse technical domains. You thrive in fast-paced environments and are passionate about building efficient, scalable, and reliable compute platforms.Key Responsibilities:System Architecture & DesignDesign scalable AI/HPC clusters including compute, storage, and networking with specific focus on , power delivery Evaluate and select CPUs, GPUs, accelerators, interconnects, and memory configurations for optimal cluster performance.PowerDesign leading-edge power delivery solutions for high-density AI/GPU deployments.Define power budgets, redundancy schemes, and fault tolerance mechanisms.NetworkDesign network topologies to maximize overall cluster performanceUnderstand the network performance needs of different types of workloadsUnderstand advantages and performance trade-offs of network topologies for AI/HPC clustersStorageDesign and optimize storage solutions to maximize AI/HPC cluster performanceUnderstand advantages and performance trade-offs of cluster storage solutions, e.g. Lustre, Ceph, etc.CollaborationWork across multiple organizations with subject matter experts from hardware, software, network, data center, and operations teams to deliver scalable, efficient, and reliable compute infrastructure.Experience in HPC, AI infrastructure, or data center systems engineering.Strong understanding of rack and data center power deliveryKnowledge of GPU/CPU architectures, PCIe, UALink, InfiniBand, and Ethernet networking.Familiarity with AI/ML frameworks and workload characteristics.Excellent problem-solving, communication, and documentation skills.Preferred Qualifications:Experience in HPC, AI infrastructure, or data center systems engineering.Experience designing power delivery solutions for racks and data centersContributions to open-source HPC or AI infrastructure projects.Academic CredentialsBachelor's or Master's degree in Electrical Engineering, COmputer Engineering, Computer Science or related field.This role is not eligible for Visa Sponsorship#LI-KW1Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
- ...high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI... ...engineer to design scalable AI/HPC clusters with specific focus on compliance with... ...experts from hardware, software, network, data center, and operations teams to deliver...Suggested
$89.2k - $209.5k
...challenges with broad technical impact.Our AI Infrastructure Engineering team is... ...behind the world’s largest AI mega-cluster. These systems support every phase of next-generation data center lifecycle management — including planning, design, deployment, provisioning,...SuggestedTemporary workFlexible hours- We Are:The Global AI Infrastructure team is at the center of enabling infrastructure reinvention... ...hybrid environments to design, build, and operate... ...configure, and manage XPU-based clusters (GPU, DPU, LPU, CPU) across... ...with existing IT systems, data pipelines, security...SuggestedFull timeWork experience placementLive inWork at officeLocal area
- ...that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...seeking a Technical Program Manager to lead execution of AI cluster engineering programs with deep focus on GPU platforms, rack-...Suggested
$184k - $287.5k
...Senior Solutions Architect specializing in Data Center Systems & Performance to join our elite... ...optimizing the performance of world-class AI, deep learning, and HPC ecosystems. Come... ...benchmarking suites to stress-test high-performance clusters and establish performance baselines.Apply...SuggestedFull time- ...that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...the development and deployment of next-generation packaging design automation and AI-Driven workflows. You will focus on accelerating...
- ...that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...opportunity to collaborate with world-class engineers across silicon design, board engineering, validation, manufacturing, and product...Internship
- Advanced Micro Devices in Austin seeks an experienced AI systems engineer to design scalable AI/HPC clusters, selecting compute, storage and networking... ...deployments. You will collaborate with hardware, software, data center, and operations teams to deliver a robust compute...
$148.7k - $201.2k
AWS Infrastructure Services (AIS) owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re... ...the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling...Flexible hours- ...they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built... ..., rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and...Local area
$200k - $275k
Fluidstack is seeking a Network Design Engineer in Austin, Texas, to lead the design of advanced data center network fabrics essential for AI training. You will own designs from requirements to deployment, ensuring operationally sound solutions. The role requires expertise...- Grow with us AI Model Architect — Silicon-Software Co-Design Austin, Texas The Voice of the Model. The Architect of the Machine. The Mission Most AI architects... ...L2/L3. You understand why 5G inference isn’t just a data‑center problem in a smaller box. Compiler Curiosity - You...Contract workTemporary work
- ...they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built... ..., rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and...
- Fluidstack is seeking a network professional to own and design the facility networks across sites, focusing on OT networks, security, and... ...engineers to ensure reliable, scalable network infrastructure for AI compute deployments. The role requires experience with OT/network...
$162.6k - $244k
...Engineering Group > Hardware Engineering General Summary The Qualcomm Data Center AI System Hardware and Validation Engineering team develops rack... ...Lead execution for next‑generation Qualcomm rack‑scale designs within a matrixed organization spanning multiple teams and...Work experience placementWork at officeWork from home- ...the world's most important technology and AI companies on their most transformational... ...product at scale? How do you build and operate data center infrastructure fast enough to meet demand... ...| inclusive of helping companies design, build, and operate AI infrastructure efficiently...Full timeLive inWork at officeLocal areaShift work
- ...the world's most important technology and AI companies on their most transformational... ...product at scale? How do you build and operate data center infrastructure fast enough to meet demand... ...| inclusive of helping companies design, build, and operate AI infrastructure efficiently...Full timeLive inWork at officeLocal areaShift work
- ...the world's most important technology and AI companies on their most transformational... ...product at scale? How do you build and operate data center infrastructure fast enough to meet demand... ...| inclusive of helping companies design, build, and operate AI infrastructure efficiently...Full timeLive inWork at officeLocal areaShift work
$171k - $231.4k
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS... ...keep the cloud running. We support all AWS data centers and all of the servers, storage,... ...server system NPIs such as compute, gpu (AI/ML), or storage servers- Extensive experience...Local areaFlexible hours$126k - $232k
...trusted insights in electronic design, simulation, prototyping, test,... ...Responsibilities About the Role AI infrastructure is being... ...challenges. Keysight sits at the center of that shift, and this Industry... ...customers need in the areas of AI data centers, accelerated computing,...Flexible hoursShift work$183k - $247.6k
Do you want to shape the future of AI? Join the team building the foundation of the world... ...come to life at scale. Here, you’ll design, deliver, and operate next-generation infrastructure... ...and bring it to market- 5+ years of data center engineering or operations experience-...Local areaFlexible hours$173.9k - $235.2k
...you want to shape the future of Generative AI at AWS? Join the team building the... ...models come to life at scale. Here, you’ll design, deliver, and operate next-generation infrastructure... ...years of systems development in an IT or data center environment experience- 3+ years of...InternshipLocal areaFlexible hours- ...company for Bitcoin mining and AI cloud. Bitdeer is committed to... ...for its customers. Apart from designing industry-leading ASIC chips... ...build distributed systems for cluster management, scheduling, and resource... ...~ Solid understanding of 1) Data structures and algorithms 2)...Work experience placementLocal area
$150k - $250k
...instead of things they had to do. Powerful AI will be the biggest lever for human... ...layer of the stack. We acquire power, design and build data centers, and operate them – with teams... ...teams, shaping the interfaces that GPU cluster operators, AI platform engineers, and...Local area- ...high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI... ...moves the world forward.THE TEAM:AMD's Data Center GPU organization is transforming the... ...spanning platform development, cluster deployment, performance optimization...Remote workFlexible hours
- ...choice that helps market-leading brands design, build and deliver innovative products... ...power and thermal management solutions for AI data centers and other mission-critical applications... ...supporting hardware that connect rack, cluster, and facility-level infrastructure into...Full timeVisa sponsorshipFlexible hours
$114.6k - $234.6k
...global leader in the RDMA cluster networking domain and... ...driving the development and design of state-of-the-art RDMA... ...specifically for AI, ML, HPC workloads.We strive... ...Oracle brings together the data, infrastructure,... ...validation in the AI data center builds to the OCI standards...Temporary workFlexible hours$152k - $241.5k
...supports our Electronic Design Automation (EDA)... ...software, operating systems, cluster schedulers, networking,... ...system signals, scheduler data, and network and... ...infrastructure across multiple data centers or heterogeneous... ...vacancy. NVIDIA uses AI tools in its recruiting...Full timeRemote work- ...a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer... ...procurement, transport logistics, data center design and construction, equipment management,... ...Overview ~ We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm...Remote jobFull timeLocal areaShift work
$133.1k - $306.4k
Leads the design, architecture, engineering, and operational... ...fabrics supporting AI, HPC, and cloud infrastructure... ...brings together the data, infrastructure,... ...capacity models for multi-cluster deployments.Drive technology... ...routing protocols, data center networking, and network...Temporary workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Cluster & Data Center Design Engr. Be the first to apply!

