AI/ML Platform Engineer
Advanced Micro Devices , Inc.
WHAT YOU DO AT AMD CHANGES EVERYTHING
At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you'll discover the real differentiator is our culture. We push the limits of innovation to solve the world's most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE We are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. This role focuses on the infrastructure and platform systems that support large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation, artifact management, and GPU cluster utilization. You will work closely with ML Systems Research Engineers, AI Research Scientists, Applied AI Engineers, and hardware domain experts to operationalize the Blueprint framework across kernel optimization, RTL/PPA optimization, ECO fixing, verification, simulation, and debugging workflows. This is a platform engineering role, not a pure research role. The focus is to build robust shared systems that allow researchers and engineers to run more experiments, compare results reliably, reduce manual orchestration, and move successful workflows into production engineering use. THE PERSON You are a strong systems engineer who enjoys building reliable platforms for AI researchers and applied engineers. You understand distributed systems, ML workloads, GPU infrastructure, experiment management, and production reliability. You can turn messy research workflows into reusable services, APIs, dashboards, job systems, and automation. You care about reproducibility, observability, performance, and developer experience. You are comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without requiring every requirement to be fully specified upfront. KEY RESPONSIBILITIES
#LI-Hybrid Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. This posting is for an existing vacancy.
At AMD, our mission is to build great products that accelerate next-generation computing experiences-from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you'll discover the real differentiator is our culture. We push the limits of innovation to solve the world's most important challenges-striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE We are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. This role focuses on the infrastructure and platform systems that support large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation, artifact management, and GPU cluster utilization. You will work closely with ML Systems Research Engineers, AI Research Scientists, Applied AI Engineers, and hardware domain experts to operationalize the Blueprint framework across kernel optimization, RTL/PPA optimization, ECO fixing, verification, simulation, and debugging workflows. This is a platform engineering role, not a pure research role. The focus is to build robust shared systems that allow researchers and engineers to run more experiments, compare results reliably, reduce manual orchestration, and move successful workflows into production engineering use. THE PERSON You are a strong systems engineer who enjoys building reliable platforms for AI researchers and applied engineers. You understand distributed systems, ML workloads, GPU infrastructure, experiment management, and production reliability. You can turn messy research workflows into reusable services, APIs, dashboards, job systems, and automation. You care about reproducibility, observability, performance, and developer experience. You are comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without requiring every requirement to be fully specified upfront. KEY RESPONSIBILITIES
- Build and operate the shared AI platform for agentic engineering workflows, including job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking.
- Develop reliable infrastructure for distributed training, distributed inference, batch evaluation, and large-scale agent rollout across GPU clusters.
- Build platform services for benchmark execution, correctness checking, profiling, regression tracking, and reproducible evaluation.
- Maintain artifact systems for generated kernels, RTL edits, traces, logs, profiler outputs, benchmark results, simulator outputs, and formal verification artifacts.
- Support scalable integrations with compilers, ROCm/HIP tooling, profilers, simulators, EDA tools, vLLM, SGLang, and internal engineering systems.
- Improve GPU cluster utilization, scheduling efficiency, reliability, quota management, and workload isolation.
- Build dashboards and observability systems for experiment status, resource usage, failure modes, benchmark trends, regression detection, and team productivity.
- Partner with ML Systems Research Engineers to productionize research workflows for RL systems, inference systems, quantification systems, and evaluation pipelines.
- Partner with Applied AI Engineers to make Blueprint harnesses reusable across kernel optimization, RTL optimization, verification, firmware, and CPU/GPU performance workflows.
- Establish platform standards for reproducibility, data retention, run metadata, artifact lineage, access control, and operational reliability.
- Distributed ML platform infrastructure for training, inference, evaluation, and agent execution.
- GPU cluster scheduling, utilization, reliability, quota management, and multi-user workload isolation.
- Experiment tracking, artifact management, run lineage, dashboards, and reproducible workflow management.
- Benchmark and evaluation automation for correctness, performance, regression detection, and reproducibility.
- Integration with ROCm/HIP, Triton, compilers, profilers, vLLM, SGLang, simulators, formal tools, and EDA flows.
- Caching and parallelization for expensive feedback loops, including simulator, compiler, verifier, and benchmark workloads.
- Production-quality APIs, services, workflow engines, and developer tooling for research and engineering teams.
- Strong programming skills in Python and one or more systems languages such as C++, Go, or Rust.
- Experience building ML platforms, AI infrastructure, distributed systems, workflow orchestration, experiment platforms, GPU cluster infrastructure, or developer platforms.
- Strong understanding of job scheduling, distributed workloads, logging, monitoring, reliability, storage systems, and production operations.
- Experience with Kubernetes, Ray, Slurm, workflow engines, containerization, CI/CD, data pipelines, or large-scale compute orchestration.
- Ability to build reliable services, APIs, dashboards, and developer tools used by researchers and engineers.
- Strong debugging skills across distributed systems, GPU workloads, storage systems, networking, containers, and production infrastructure.
- Good collaboration skills with AI researchers, ML systems researchers, applied engineers and hardware domain experts.
- Experience with GPU platforms, ROCm/HIP, CUDA, profiling, kernel benchmarking, model serving, or distributed training/inference.
- Experience with vLLM, SGLang, Triton, PyTorch, JAX, Ray, Kubernetes, Slurm, MLflow, Weights & Biases, or similar systems.
- Experience building experiment tracking systems, artifact stores, workflow engines, benchmark automation platforms, or developer productivity tools.
- Experience supporting LLM agents, tool-use systems, large-scale sampling, or automated program optimization workflows.
- Familiarity with compiler, profiler, simulator, formal verification, EDA, firmware, or hardware performance workflows is a strong plus.
- Experience operating shared GPU clusters or high-performance ML infrastructure in a multi-user research environment.
#LI-Hybrid Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here. This posting is for an existing vacancy.
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the AI/ML Platform Engineer in Santa Clara, CA vacancy
$197.3k - $225.1k
Lead AI/ML Engineer (Platform, kubeflow) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences...SuggestedFull timePart timeLocal area- ...An innovative AI startup is seeking an experienced Machine Learning Engineer to design and deploy production-grade ML systems. This role involves using AI to enhance sales data insights through real-time audio understanding. Candidates should have a relevant degree and...Suggested
- ...ML Engineer Santa Clara, California, United States About the Job Our client is a rapidly growing Tier 1 VC backed startup based... ...revolutionizing how outside sales and service teams work. Their AI technology captures and analyzes real-world conversations, providing...SuggestedFull time
$189.4k - $300.6k
...of autonomous driving? Join the Embodied AI team at General Motors. Our team is developing... ...in autonomous vehicle development. We engineer high-performance tools that identify top-performing... ...models and partner with data-intensive ML teams to drive rapid innovation. In...SuggestedLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- A tech-driven AI company is seeking an ML Engineer to design and deploy production-grade ML systems with full ownership of the model lifecycle. The ideal candidate will possess a degree in Computer Science and 1-6 years of experience in ML engineering, with robust programming...Suggested
- ...ML Systems Engineer Role Overview: We're seeking an experienced engineer to build our ML data infrastructure platform. You'll create the systems and tools that enable efficient data preparation... ...Storage for data lakes ~ Vertex AI Feature Store ~ Cloud Composer (...Local area
$150k - $350k
...of chip design and verification with agentic AI workflows. Our platform leverages cutting‑edge generative AI to assist engineers in RTL design, simulation, and verification,... ...chips. Position Overview We are seeking an ML Systems Engineer to optimize the performance...- ...About the Role We’re looking for an Applied ML Engineer to design, evaluate, and scale... ...builders from unicorn startups and top UGC platforms, backed by Andreessen Horowitz SR04 and top angels across the video AI space. We’re still early, and this is your chance...Full time
$150k
...research, nurture the next generation of AI builders, and drive transformative contributions... ...-class researchers, data scientists, and engineers, tackling the most fundamental and... ...including experience with Machine Learning (ML) models, ML infrastructure, Natural Language...Full timeWorldwideVisa sponsorship- ...development, and deployment of advanced AI agents and agentic systems. Architect and... ...product managers, UX designers, and other engineers to define requirements and deliver impactful... ...or PyTorch. Experience with cloud platforms (AWS) and containerization technologies (...Full timeWork experience placement
- ...Wayve is the leading developer of Embodied AI technology. Our advanced AI software and... ...your career! The Role As an ML Engineer within the Application Engineering team,... ...deployment. You’ll collaborate deeply with AI Platform, Simulation, Robot SW and Model Release...Full timeWork at officeWork from home
$150k
...nurture the next generation of AI builders, and drive... ...researchers, data scientists, and engineers, tackling the most fundamental... ...The Role The Distributed ML Engineer will play a role at the... ...the-art hardware and software platforms to improve their efficiency with...Full timeWork experience placementVisa sponsorship- AI /Client Engineer Santa Clara, CA/ San Jose, CA Pay Rate: $65/hr on c2c Client: Paypal Job Description: Strong in NLP Trend in analysis and anomaly detection BS in CS 5+ yrs experience or MS 3+ yrs experience
$157.2k - $254.1k
...Integrity, and Inclusion. We weave AI into the fabric of everything... ...comprehensive AI security platform. Organizations are increasingly... ...Machine Learning Inference Engineer, you will serve as a technical... ...strategy of our AI platform - ML inference. Beyond individual contribution...Full timeWork at office- ...Job Title: AI/ML Engineer Location: Cupertino, CA Onsite/ Remote: Onsite JD: Build and optimize RAG pipelines, vector databases... ...AI workflows with PLM systems or enterprise knowledge platforms REQUIRED SKILLS Strong experience with LLMs and agent frameworks...Remote work
- ...years of experience with a Master's degree in AI/ML Job Description (Brief): We are looking for a talented AI/ML Engineer with at least 3 years of experience to join... ...with Python, TensorFlow, PyTorch, and cloud platforms is preferred. Key Skills: Machine...
- ...Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale, CA 94085 (Hybrid) Technical Hiring Criteria (Must Haves) Top 3 Required skills: Machine Learning, AI, Python Years of experience in each of the must-have skills: 7 Years Job description...
- ...We are seeking an innovative and results-oriented Mid-Level AI/ML Engineer to join our dynamic team. This role is crucial for transforming... ...Agentic AI systems to solve complex business problems. Platform Expertise: Leverage and integrate core generative AI platforms...Permanent employmentContract workLocal area
$125k - $222k
...Software Engineer Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating... ...can be used for perception, world modeling, and ML driven autonomy Test and evaluate your algorithms...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift$272k - $431.25k
## Principal Platform Software Engineer - RASApplylocations: US, CA, Santa Clara: US, Remotetime type: Full... ..., we are increasingly known as “the AI computing company.” We are looking to grow... ...Confidential Compute. Experience with ML and multi-variable optimization...$275.8k - $340.5k
...About the team: The AV ML Infra team at GM builds ML infrastructure... ...to meet the unique demands of AI and ML innovation, supporting... ...the productivity of ML engineers, and drive the adoption of cutting... ...Experience with Google Cloud Platform, Microsoft Azure, or Amazon...Local areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...We are seeking a Principal Engineer, AI Safety to lead the architecture and technical strategy for safeguarding large language models (LLMs) and multimodal AI systems from misuse. This role focuses on developing scalable defenses against jailbreaks, prompt injection attacks...
$148.7k - $297.3k
...executives, and scientists. THE OPPORTUNITY This Principal AI/ML Engineer position can work out of our Santa Clara, CA location.... ...vital in designing, developing and maintaining a robust AI platform, establishing production-grade MLOps capabilities, and collaborating...Shift work$129k - $198.4k
...Job Description Role: As an AI/ML Engineer on the Metrics Frameworks team, part of the Simulation, Evaluation, and Data organization, you will be an individual contributor focused on developing and optimizing infrastructure to accelerate autonomous vehicle development...Local areaWork from home$300k
...nurture the next generation of AI builders, and drive... ...researchers, data scientists, and engineers, tackling the most fundamental... ...'re looking for a distributed ML infrastructure engineer to help... ...• Experience working on an ML platform/ infrastructure, and/or distributed...Full timeFlexible hours$268.6k - $395k
...discovery experiences. As a Principal Engineer, you will lead the technical direction for AI-first experiences, including... ...understanding. Rigorously evaluate ML and LLM models using a... ...re one of the fastest growing Ads platforms in the world and we're looking to...Hourly payWork at officeLocal areaRemote workFlexible hours$137.1k - $201.6k
...subscription. We are forming a new team that will leverage AI and advanced ML to power decision making in real-time – from personalized... ...About the Role We’re looking for a Staff Machine Learning Engineer to drive the design and development of large-scale ML/...Hourly payWork at officeLocal areaRemote workFlexible hours- ...enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across hybrid multi... ...the world running. Our Team's Vision: Our Engineering team is shaping the future of cybersecurity. We thrive...Immediate start
$229.5k - $360k
...Roku is the #1 TV streaming platform in the U.S., Canada, and Mexico... ...depends on a robust, flexible ML platform built for experimentation... .... Our work blends innovation, engineering excellence, and a deep... ...off. How will I use AI at Roku? At Roku, we don'...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours$150k - $240k
..., Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion... ...family commitments. Meet our software engineers! Meet some of our software engineers who are... ...Construct optimized data pipelines to run ML models Evolve our data engine architecture...Full timeFor contractorsWork experience placementFor subcontractorCasual workWork at officeRemote workDay shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI/ML Platform Engineer. Be the first to apply!
Related searches
- ai developer Santa Clara, CA
- ai engineer Santa Clara, CA
- machine learning engineer Santa Clara, CA
- computer vision machine learning engineer Santa Clara, CA
- platform engineer Santa Clara, CA
- platform developer Santa Clara, CA
- platform product manager Santa Clara, CA
- platform manager Santa Clara, CA
- power platform Santa Clara, CA
- machine learning research scientist Santa Clara, CA



