Post-Training Platform Infrastructure Engineer
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a systems-minded engineer who lives at the intersection of large-scale model inference, distributed systems, and performance optimization. This role focuses on post-training and inference infrastructure, with particular emphasis on P/D disaggregation, KV cache lifecycle management, and efficient offloading mechanisms across both inference and reinforcement learning (RL) systems. THE PERSON:You enjoy reverse-engineering modern ML infrastructure, reasoning about memory and compute tradeoffs, and turning research insights into production-grade features. You are comfortable diving into unfamiliar frameworks, understanding their architectural choices, identifying bottlenecks, and improving them through principled engineering. KEY RESPONSIBILITIES:Research and deeply understand modern LLM inference frameworks, including: Architecture and design tradeoffs of P/D (prefill / decode) disaggregation KV cache lifecycle, memory layout, eviction strategies, and reuse KV cache offloading mechanisms across GPU, CPU, and storage backends Analyze and compare inference execution paths to identify: Performance bottlenecks (latency, throughput, memory pressure) Inefficiencies in scheduling, cache management, and resource utilization Develop and implement infrastructure-level features to: Improve inference latency, throughput, and memory efficiency Optimize KV cache management and offloading strategies Enhance scalability across multi-GPU and multi-node deployments Apply the same research-driven approach to RL frameworks: Study post-training and RL systems (e.g., policy rollout, inference-heavy loops) Debug performance and correctness issues in distributed RL pipelines Optimize inference, rollout efficiency, and memory usage during training Collaborate with research and applied ML teams to: Translate model-level requirements into infrastructure capabilities Validate performance gains with benchmarks and real workloads Document findings, architectural insights, and best practices to guide future system design PREFERRED EXPERIENCE:Strong background in systems engineering, distributed systems, or ML infrastructure Hands-on experience with GPU-accelerated workloads and memory-constrained systems Solid understanding of: LLM inference workflows (prefill vs decode) Attention mechanisms and KV cache behavior Multi-process / multi-GPU execution models Proficiency in Python and C++ (or similar systems languages) Experience debugging performance issues using profiling tools (GPU, CPU, memory) Ability to read, understand, and modify complex open-source codebases Strong analytical skills and comfort working in research-heavy, ambiguous problem spaces Direct experience with LLM inference frameworks or serving stacks Familiarity with: GPU memory hierarchies (HBM, pinned memory, NUMA considerations) KV cache compression, paging, or eviction strategies Storage-backed offloading (NVMe, object stores, distributed file system) Experience with distributed RL or post-training pipelines Knowledge of scheduling systems, async execution, or actor-based runtimes Contributions to open-source ML or systems projects Experience designing benchmarking suites or performance evaluation frameworks ACADEMIC CREDENTIALS: Bachelor’s or master's degree in computer science, computer engineering, electrical engineering, or equivalent LOCATION:San Jose, CA (Hybrid). May consider other US locations.#LI-CJ3#HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$147k - $211k
...with developing large-scale infrastructure, distributed systems or networks... ..., and software test engineering.Google's software engineers... ...and providing the essential platforms that enable developers to build... ..., and relevant education or training. US: $147000 - $211000 (USD)...TrainingWorldwide$174k - $253k
...technologies. Google's software engineers develop the next-generation... ...solutions.The AI and Infrastructure team is redefining what’s possible... ...and providing the essential platforms that enable developers to... ..., and relevant education or training. US: $174000 - $253000 (USD)...TrainingWorldwide$207k - $301k
...coach a distributed team of engineers.Facilitate alignment and clarity... ...and developing large-scale infrastructure, distributed systems or... ...the next generation of Google platforms, we make Google's product portfolio... ..., and relevant education or training. US: $207000 - $301000 (USD)...Training$126k - $181k
...scripts to improve product and engineering health.Develop test plans... ...functionality. Improve existing test infrastructure or create new test... ...Capacity Assurance team in Platforms Infrastructure Engineering is... ..., and relevant education or training. US: $126000 - $181000 (USD)...Training$147k - $211k
...Review code developed by other engineers and provide feedback to... ...software solutions.The AI and Infrastructure team is redefining what’s possible... ...and providing the essential platforms that enable developers to... ..., and relevant education or training. US: $147000 - $211000 (USD)...TrainingWorldwide$159k - $230k
...network, power and cooling infrastructure for end-to-end system testing... ...hardware/software designers, test engineers, on project planning within... ...working environment.Our Platforms Infrastructure Engineering... ..., and relevant education or training. US: $159000 - $230000 (USD)...Training$174k - $253k
...Review code developed by other engineers and provide feedback to... ...software solutions.The AI and Infrastructure team is redefining what’s possible... ...and providing the essential platforms that enable developers to... ..., and relevant education or training. US: $174000 - $253000 (USD)...TrainingWorldwide$240k - $250k
...Saviynt's AI-powered identity platform manages and governs human... ...DOING Design and operate the infrastructure powering Cloud and Edge... ...and security. AI & Agentic Engineering Apply AI-assisted engineering... ...skill sets; experience and training; licensure and...Training$176k - $276k
...of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy... ...operate the Kubernetes-based platform and shared services used to... ...looking for a hands-on senior engineer to own the lifecycle and... ...least until August 3, 2026.This posting is for an existing vacancy....Full timeRemote workWeekend work$144k - $175k
...highly motivated and versatile Senior Platform Engineer to join our platform engineering team.... ...bridging the gap between development, infrastructure, and our burgeoning AI/ML initiatives,... ...the underlying platform used for model training, experimentation, and serving.Focus on...TrainingLocal area- ...Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...home day is currently Tuesday.Engineering at Lambda is responsible for building... ..., and RBAC across the platform.Lead incident response, root-cause analysis, and post-mortems for platform issues.Mentor...Work at officeLocal areaWork from homeFlexible hours
$174k - $253k
...with developing large-scale infrastructure, distributed systems or networks... ...projects.Google's software engineers develop the next-generation... ...). We focus on building a platform that allows enterprises to run... ..., and relevant education or training. US: $174000 - $253000 (USD)...Training$207k - $301k
...coach a distributed team of Engineers.Facilitate alignment and clarity... ...forward.The AI and Infrastructure team is redefining what’s possible... ...and providing the essential platforms that enable developers to... ..., and relevant education or training. US: $207000 - $301000 (USD)...TrainingWorldwide$178k - $321k
...working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data... ...setting, and the governed data and AI infrastructure everything else depends on. We hire on... ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix...$132k - $189k
...qualifications:Bachelor's degree in Electrical Engineering, Optics, Physics, a related field, or... ..., efficiency, and integration.Our Platforms Infrastructure Engineering team designs and builds... ..., and relevant education or training. US: $132000 - $189000 (USD) + 15% bonus...TrainingWorldwide$159k - $231k
...qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a... ...As a Power Integrity Engineer within Platforms Infrastructure Engineering, you will play a pivotal... ..., and relevant education or training. US: $159000 - $231000 (USD) + 15% bonus...TrainingWorldwide$100k
...in software models, compilers, platforms, networking, and... ...seniorities.Tenstorrent’s AI Software Infrastructure team builds the platforms that... ...may differ from the one in this posting.Who You AreStrong backend or infrastructure engineer with experience building and operating...Permanent employment- ...TeamPlatform Systems is the engineering foundation of AIMS (... ...productive and the platform reliable — spanning... ...across AIMS and infrastructure teams — representing... ...online and offline model training pipelines and the platform... ...military service.Job Posting Date:06-30-2026Job...TrainingHourly payFull timeImmediate startFlexible hours
$132k - $190k
...qualifications:Bachelor’s degree in Electrical Engineering, Computer Engineering, Physics, a... ...are deployed in the data center.Our Platforms Infrastructure Engineering team designs and builds... ..., and relevant education or training. US: $132000 - $190000 (USD) + 15% bonus...TrainingWorldwide- ...ROLEWe are hiring AI / ML Platform Engineers to build the platform layer... .... This role focuses on the infrastructure and platform systems that support... ...execution, distributed training and inference, experiment... ...Policy” is available here.This posting is for an existing vacancy.Training
$150k - $218k
...the test fixture for better engineering efficiency and data... ...designs powerful computing infrastructures with custom-built machines.... ...and providing the essential platforms that enable developers to build... ...and relevant education or training. US: $150000 - $218000 (USD...TrainingWorldwide$132k - $190k
...experience in mechatronics engineering and robotics product development... ...to join our Physical infrastructure Robotics team. This team is... ...and providing the essential platforms that enable developers to build... ..., and relevant education or training. US: $132000 - $190000 (USD)...TrainingContract workWorldwide- ...moving fast, and the infrastructure that got us here needs... ...-generation AI/ML platform is one of the highest... ...Platform Systems is the engineering foundation of AIMS,... ...path looks like across training pipelines, AI... ...military service.Job Posting Date:07-01-2026Job Requisition...TrainingHourly payFull timeImmediate startFlexible hours
$207k - $301k
...direction to meet anticipated infrastructure needs.Oversee the planning,... ..., the work of a Software Engineer goes beyond just Search. Software... ...to lead the growth of our Platforms Infrastructure team. In this... ..., and relevant education or training. US: $207000 - $301000 (USD)...TrainingWorldwide$274k - $304k
...AI Platform Engineer - Training & Inference Saviynt's AI-powered identity platform manages and... ...and cloud LLMs • Build RL training infrastructure: define Flyte workflows for RL pipelines... ...(nice to have): INT8/INT4/FP8 post-training quantization (GPTQ, AWQ, or...Training$236k - $330k
...Robotics initiatives and roadmap, engineering execution with business... ...liquid-cooled AI hardware infrastructure.Oversee the end-to-end deployment... ...bipedal/wheeled humanoid platforms for flexible, human-centric... ..., and relevant education or training. US: $236000 - $330000 (USD)...TrainingContract workRemote workWorldwideFlexible hours$280k - $380k
...TVRoku is the #1 TV streaming platform in the U.S., Canada, and... ...reliability and automation, engineering systems that perform under stress... ...can, and turning complex infrastructure into reliable, well‑documented... ...management processes and post-incident reviews Identify performance...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours$184k - $287.5k
...for a Senior Software Engineer to lead the bring-up,... ...optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales... ...-scale AI clusters, infrastructure, and end-to-end... ...benchmark AI pre-training, post-training, and inference...TrainingFull timeRemote work$245k - $295k
...As the only vertically integrated AI infrastructure company built from the ground up, we... ...seeking a Senior Manager, Infrastructure Platform Engineering to lead a team building core systems... ...challenges of GPU clusters, AI training, and inference workloads Working knowledge...TrainingTemporary workImmediate start$184k - $287.5k
...architects work across product, engineering, sales, developer relations,... ..., deploy and optimize AI infrastructure.This role will focus on... ...accelerated infrastructure for training, fine-tuning, inference, retrieval... ...until July 20, 2026.This posting is for an existing vacancy....TrainingFull timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Post-Training Platform Infrastructure Engineer. Be the first to apply!
- platform engineer San Jose, CA
- senior platform engineer San Jose, CA
- platform developer San Jose, CA
- senior infrastructure engineer San Jose, CA
- infrastructure engineer San Jose, CA
- infrastructure developer San Jose, CA
- remote infrastructure engineer San Jose, CA
- platform product manager San Jose, CA
- platform manager San Jose, CA
- power platform San Jose, CA

