Post-Training Platform Infrastructure Engineer (San Jose)
AMD
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a systems-minded engineer who lives at the intersection of large-scale model inference, distributed systems, and performance optimization. This role focuses on post-training and inference infrastructure, with particular emphasis on P/D disaggregation, KV cache lifecycle management, and efficient offloading mechanisms across both inference and reinforcement learning (RL) systems. THE PERSON:You enjoy reverse-engineering modern ML infrastructure, reasoning about memory and compute tradeoffs, and turning research insights into production-grade features. You are comfortable diving into unfamiliar frameworks, understanding their architectural choices, identifying bottlenecks, and improving them through principled engineering. KEY RESPONSIBILITIES:Research and deeply understand modern LLM inference frameworks, including: Architecture and design tradeoffs of P/D (prefill / decode) disaggregation KV cache lifecycle, memory layout, eviction strategies, and reuse KV cache offloading mechanisms across GPU, CPU, and storage backends Analyze and compare inference execution paths to identify: Performance bottlenecks (latency, throughput, memory pressure) Inefficiencies in scheduling, cache management, and resource utilization Develop and implement infrastructure-level features to: Improve inference latency, throughput, and memory efficiency Optimize KV cache management and offloading strategies Enhance scalability across multi-GPU and multi-node deployments Apply the same research-driven approach to RL frameworks: Study post-training and RL systems (e.g., policy rollout, inference-heavy loops) Debug performance and correctness issues in distributed RL pipelines Optimize inference, rollout efficiency, and memory usage during training Collaborate with research and applied ML teams to: Translate model-level requirements into infrastructure capabilities Validate performance gains with benchmarks and real workloads Document findings, architectural insights, and best practices to guide future system design PREFERRED EXPERIENCE:Strong background in systems engineering, distributed systems, or ML infrastructure Hands-on experience with GPU-accelerated workloads and memory-constrained systems Solid understanding of: LLM inference workflows (prefill vs decode) Attention mechanisms and KV cache behavior Multi-process / multi-GPU execution models Proficiency in Python and C++ (or similar systems languages) Experience debugging performance issues using profiling tools (GPU, CPU, memory) Ability to read, understand, and modify complex open-source codebases Strong analytical skills and comfort working in research-heavy, ambiguous problem spaces Direct experience with LLM inference frameworks or serving stacks Familiarity with: GPU memory hierarchies (HBM, pinned memory, NUMA considerations) KV cache compression, paging, or eviction strategies Storage-backed offloading (NVMe, object stores, distributed file system) Experience with distributed RL or post-training pipelines Knowledge of scheduling systems, async execution, or actor-based runtimes Contributions to open-source ML or systems projects Experience designing benchmarking suites or performance evaluation frameworks ACADEMIC CREDENTIALS: Bachelor’s or master's degree in computer science, computer engineering, electrical engineering, or equivalent LOCATION:San Jose, CA (Hybrid). May consider other US locations.#LI-CJ3#HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$172.5k - $306.63k
...Machine Learning Engineer to join our Applied... ...you’ll build the infrastructure that powers large... ..., multimodalAI training and inference.You... ...developtools and platforms that help teams train... ...on the job posting), the application... ...SummaryLocation: San Jose; SeattleType: Full...TrainingFull timeTemporary workPart timeLocal areaWorldwide- ...leader in AI cloud infrastructure serving tens of thousands... ...presence in our San Francisco, San Jose, or Bellevue office... ...is currently Tuesday.Engineering at Lambda is responsible... ...and RBAC across the platform.Lead incident... ...cause analysis, and post-mortems for platform...SuggestedPart timeWork at officeLocal areaWork from homeFlexible hours
$245k - $350k
...native Zero Trust Exchange platform. This innovation... ...three days a week in San Jose, reporting to the Director... ...,Software Development Engineering ZIdentity team. You... ...displayed on each job posting reflects the minimum and... ...relevant education or training.The base salary range...TrainingFull timePart timeWork at officeLocal area3 days per week$182k - $260k
...Zero Trust Exchange platform. This innovation... ...for a Sr. Engineering Manager, API Platform... ...role going into the San Jose, CA office 3 days... ..., CI/CD, and Infrastructure-as-Code across large... ...on each job posting reflects the minimum... ...relevant education or training.The base salary...TrainingFull timePart timeWork at officeLocal area3 days per week$250k - $310k
...aerospace company based in San Jose, California building... ...: AI-native engineering will let us ship certified... ...the cross-domain platform — by machines, for machines... ...MLOps and simulation infrastructure, code-to-flight CI/CD... ...processing, model training, evaluation, validation...TrainingPart timeLocal area$195.2k - $391.2k
...looking for a Senior Engineering Manager to lead... ...Kubernetes platform powering enterprise... ...ML workloads, GPU infrastructure, and enterprise applications... ...GPU scheduling, training, and inference... ...applies (i.e. San Jose, Durham, Mexico... ...from the date of posting. In good faith,...TrainingPart timeWork at officeLocal areaRemote workRelocation package3 days per week- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per... ....What the Role is AboutThe Senior Platform Engineer plays a key role in accelerating...Part timeWork at officeLocal areaWork from homeFlexible hours
- ...working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data... ...setting, and the governed data and AI infrastructure everything else depends on. We hire on... ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix...Part time
$255k - $340k
...Cloud, is a leader in AI cloud infrastructure serving tens of thousands of... ...requires presence in our San Jose office location 4 days per week... ...currently Tuesday.Hardware Engineering at Lambda is responsible for... ...for state-of-the-art platforms, new product introduction (NPI...Part timeWork at officeLocal areaWork from homeFlexible hours$248k - $310k
...native Zero Trust Exchange platform. This innovation... ...looking for a Director, Engineering, Common Platform... ...team. This is a Hybrid (San Jose, CA office 3 days a week... ...displayed on each job posting reflects the minimum and... ...relevant education or training.The base salary range...TrainingFull timePart timeWork at officeLocal area3 days per week$124k - $271.2k
...level, security IAM and Infrastructure SRE. A solid SRE &... ...ticketing to proactive engineering. You will not just... ...GCP, and critical SaaS platforms like Okta and... ...will lead blameless post-mortems and define SLOs... ...data.SummaryLocation: San Jose (CA)Type: Full time...Full timePart timeWork at officeRemote work$172.5k - $306.63k
...industry-level pre-training and mid-training... ...researchers and engineers building the future... ...PyTorch, and ML infrastructure tools.Excellent... ...through innovative platforms and tools that... ...apply to this job posting because Adobe accepts... ...SummaryLocation: San Jose; San Francisco;...TrainingFull timeTemporary workPart timeLocal areaWorldwide$218k - $323.95k
...the completion of payments on our platform on behalf of our customers.We offer... ...for Venmo’s data platform infrastructure and performance engineering. The Principal Engineer sets the technology... ...is:Primary Location | Pay Range:San Jose, California | ($218,000.00 - $323,...Full timePart timeWork at officeLocal areaImmediate startFlexible hours- ..., is a leader in AI cloud infrastructure serving tens of thousands... ...requires presence in our San Francisco, San Jose, or Bellevue office location... ...the world of distributed AI training and inference, raw GPU and... ...The Lambda Infrastructure Engineering organization forges the...TrainingPart timeWork at officeLocal areaWork from homeFlexible hours
- ...leader in AI cloud infrastructure serving tens of... ...presence in our San Francisco/Bellevue... ...groundbreaking AI training and inference... ...Lambda Infrastructure Engineering organization... ...enterprise or HPC storage platforms: Vast Data, Weka,... ...Office; San Jose Office (First St)...TrainingPart timeWork at officeLocal areaWork from homeFlexible hours
$114.1k - $214.95k
...Software Development Engineer to help build and maintain... ...across cloud platforms, desktop environments... ...backend services and infrastructure.Proficiency in Python... ...as listed on the job posting), the application window... ...liability.SummaryLocation: San Jose; SeattleType: Full...Full timeContract workTemporary workPart timeLocal areaWorldwide$280k - $380k
...TVRoku is the #1 TV streaming platform in the U.S., Canada, and... ...reliability and automation, engineering systems that perform under stress... ...can, and turning complex infrastructure into reliable, well‑documented... ...management processes and post-incident reviews Identify performance...Part timeWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$136.5k - $276.5k
...complex data and to engineer new ideas and methods... ...across multiple systems, platforms, and applications.... ...CI/CD pipelines and Infrastructure as Code (IaC).Strong... ...experience, education/training, and/or skill level.... ...immediately.SummaryLocation: San Jose, California, United...TrainingFull timeContract workPart timeWork experience placementWork at officeLocal areaImmediate start2 days per week- ...thrive.The Sr Principal, Infrastructure Business Operations... ...of Infrastructure Engineering. The individual in this... ...the time of the job posting and is subject to change... ...promotion, benefits, training, discipline, and... ...com.SummaryLocation: San Jose; SeattleType: Full time...TrainingFull timePart timeLocal area
$151.8k - $265.35k
...content effortlessly. The AI for Engineering team builds a scalable, production-grade AI platform that powers creativity across... ...Colorado (as listed on the job posting), the application window will remain... ...liability.SummaryLocation: San Jose; San FranciscoType: Full time...Full timeTemporary workPart timeLocal areaWorldwide- ...thrive.The Senior Infrastructure Capacity Planner... ...and Infrastructure Engineering, and operate with... ...planning platforms. The tool matters... ...time of the job posting and is subject to... ...promotion, benefits, training, discipline, and... ...Field-VA; Field-TX; San Jose; Field-MAType: Full...TrainingFull timePart timeLocal areaShift work
$212k - $265k
...native Zero Trust Exchange platform. This innovation... ...a Principal Software Engineer (Service Platform & Orchestration... ...transformation of our infrastructure from legacy automation... ...displayed on each job posting reflects the minimum... ...relevant education or training.The base salary range...TrainingFull timeTemporary workPart timeWork at officeLocal area- ...Qualifications:Bachelor’s degree in Engineering, Business, Marketing, or... ...with AI infrastructure for training and inference workloadsExperience... ...please see the Benefits Guide posted on micron.com/careers/... ...Technology, Inc.SummaryLocation: San Jose, CA; Austin, TXType: Full...TrainingFull timePart timeLocal areaImmediate start
$182.5k - $220k
...Archer is an aerospace company based in San Jose, California building an all-electric... ...OverviewAs a Senior Staff Backend Engineer on the Platform Core team, you will build the... ...system connectors, and the integration infrastructure that every application depends on. This...Part timeLocal area$173.5k - $331.05k
...creativity by building SDKs and platform libraries that power data‑... ...Cloud. We are seeking a software engineer with strong development and... ...(as listed on the job posting), the application window will... ...civil liability.SummaryLocation: San Jose; SeattleType: Full time...Full timeTemporary workPart timeLocal areaWorldwide$210k - $336k
...team has the tools, data, and training necessary to maximize their... ...onsite presence at our San Jose, CA headquarters in alignment... ...security of the sales technology infrastructure.Troubleshoot and resolve... ...You BringBachelor’s degree in Engineering, Information Technology, or...TrainingPart timeFlexible hours$210.6k - $305.1k
...close on: 08/31/2026Job posting may be removed... ...building next-generation infrastructure software for enterprise... ...fabrics. These platforms must turn complex GPU... ...hands-on Senior Software Engineering Manager to lead a team... ..., and/or training. The full salary range...TrainingFull timeTemporary workPart timeWork at officeLocal areaFlexible hours3 days per week$98.9k - $228.7k
...DevOps team within the Engineering Operations Group. The... ...resilient VoIP infrastructure and services.About the... ...cloud or colocation platforms. Build, operate, and... ...the job description/posting.BenefitsAs part of our... ...data.SummaryLocation: San Jose (CA)Type: Full time...Full timePart timeCasual workWork at officeRemote workWorldwide$21 - $25 per hour
...Description Job Description Description: Service Locations: San Jose, Santa Clara, Milpitas Behavior Technician Serving... ...matters. No experience is necessary to apply, because on the job training is provided! ——— Position: Behavior Technician / Registered...TrainingHourly payFull timePart timeFlexible hoursAfternoon shift$200k
...Outside Sales Rep — $200K+ Potential | Warm Leads + Company Truck (San Jose / Bay Area) Location: Based in San Jose — you'll run the... ...Warm leads every week · Company truck · Uniforms · Full paid training · All tools & materials Read this if you can sell We hand...TrainingWeekly payFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Post-Training Platform Infrastructure Engineer (San Jose). Be the first to apply!
- platform developer San Jose, CA
- senior platform engineer San Jose, CA
- platform engineer San Jose, CA
- senior infrastructure engineer San Jose, CA
- infrastructure engineer San Jose, CA
- infrastructure developer San Jose, CA
- remote infrastructure engineer San Jose, CA
- digital platform specialist San Jose, CA
- platform product manager San Jose, CA
- power platform San Jose, CA















