AI/ML Platform Engineer (Santa Clara)
AMD
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. This role focuses on the infrastructure and platform systems that support large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation, artifact management, and GPU cluster utilization.You will work closely with ML Systems Research Engineers, AI Research Scientists, Applied AI Engineers, and hardware domain experts to operationalize the Blueprint framework across kernel optimization, RTL/PPA optimization, ECO fixing, verification, simulation, and debugging workflows.This is a platform engineering role, not a pure research role. The focus is to build robust shared systems that allow researchers and engineers to run more experiments, compare results reliably, reduce manual orchestration, and move successful workflows into production engineering use.THE PERSONYou are a strong systems engineer who enjoys building reliable platforms for AI researchers and applied engineers. You understand distributed systems, ML workloads, GPU infrastructure, experiment management, and production reliability. You can turn messy research workflows into reusable services, APIs, dashboards, job systems, and automation.You care about reproducibility, observability, performance, and developer experience. You are comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without requiring every requirement to be fully specified upfront.KEY RESPONSIBILITIESBuild and operate the shared AI platform for agentic engineering workflows, including job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking.Develop reliable infrastructure for distributed training, distributed inference, batch evaluation, and large-scale agent rollout across GPU clusters.Build platform services for benchmark execution, correctness checking, profiling, regression tracking, and reproducible evaluation.Maintain artifact systems for generated kernels, RTL edits, traces, logs, profiler outputs, benchmark results, simulator outputs, and formal verification artifacts.Support scalable integrations with compilers, ROCm/HIP tooling, profilers, simulators, EDA tools, vLLM, SGLang, and internal engineering systems.Improve GPU cluster utilization, scheduling efficiency, reliability, quota management, and workload isolation.Build dashboards and observability systems for experiment status, resource usage, failure modes, benchmark trends, regression detection, and team productivity.Partner with ML Systems Research Engineers to productionize research workflows for RL systems, inference systems, quantification systems, and evaluation pipelines.Partner with Applied AI Engineers to make Blueprint harnesses reusable across kernel optimization, RTL optimization, verification, firmware, and CPU/GPU performance workflows.Establish platform standards for reproducibility, data retention, run metadata, artifact lineage, access control, and operational reliability.TECHNICAL FOCUS AREASDistributed ML platform infrastructure for training, inference, evaluation, and agent execution.GPU cluster scheduling, utilization, reliability, quota management, and multi-user workload isolation.Experiment tracking, artifact management, run lineage, dashboards, and reproducible workflow management.Benchmark and evaluation automation for correctness, performance, regression detection, and reproducibility.Integration with ROCm/HIP, Triton, compilers, profilers, vLLM, SGLang, simulators, formal tools, and EDA flows.Caching and parallelization for expensive feedback loops, including simulator, compiler, verifier, and benchmark workloads.Production-quality APIs, services, workflow engines, and developer tooling for research and engineering teams.PREFERRED QUALIFICATIONSStrong programming skills in Python and one or more systems languages such as C++, Go, or Rust.Experience building ML platforms, AI infrastructure, distributed systems, workflow orchestration, experiment platforms, GPU cluster infrastructure, or developer platforms.Strong understanding of job scheduling, distributed workloads, logging, monitoring, reliability, storage systems, and production operations.Experience with Kubernetes, Ray, Slurm, workflow engines, containerization, CI/CD, data pipelines, or large-scale compute orchestration.Ability to build reliable services, APIs, dashboards, and developer tools used by researchers and engineers.Strong debugging skills across distributed systems, GPU workloads, storage systems, networking, containers, and production infrastructure.Good collaboration skills with AI researchers, ML systems researchers, applied engineers and hardware domain experts.PREFERRED EXPERIENCEExperience with GPU platforms, ROCm/HIP, CUDA, profiling, kernel benchmarking, model serving, or distributed training/inference.Experience with vLLM, SGLang, Triton, PyTorch, JAX, Ray, Kubernetes, Slurm, MLflow, Weights & Biases, or similar systems.Experience building experiment tracking systems, artifact stores, workflow engines, benchmark automation platforms, or developer productivity tools.Experience supporting LLM agents, tool-use systems, large-scale sampling, or automated program optimization workflows.Familiarity with compiler, profiler, simulator, formal verification, EDA, firmware, or hardware performance workflows is a strong plus.Experience operating shared GPU clusters or high-performance ML infrastructure in a multi-user research environment.EDUCATIONBachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience. Master's preferred; PhD is a plus, especially with work in ML systems, reinforcement learning, distributed systems, GPU computing, or AI infrastructure.LOCATION: Santa Clara, CA#LI-AG2 #LI-HybridBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$100k
...industry on cutting-edge AI technology,... ...software models, compilers, platforms, networking, and semiconductors... ...an Physical Design Engineer to lead cross-functional... ...integrate, and deploy AI/ML-driven solutions into... ...is hybrid, based out of Santa Clara, CA or Austin, TX or...SuggestedPermanent employmentPart time$176k - $276k
...operate the Kubernetes-based platform and shared services used to provision... ...for a hands-on senior engineer to own the lifecycle and... ...existing vacancy. NVIDIA uses AI tools in its recruiting processes... ...law.SummaryLocation: US, CA, Santa Clara; US, IL, Remote; US, WA,...SuggestedFull timePart timeRemote workWeekend work$148.7k - $297.3k
...Location: United States - California - Santa Clara, United States of AmericaTime Type: Full... ....THE OPPORTUNITYThis Principal AI/ML Engineer position can work out of our Santa Clara... ...developing and maintaining a robust AI platform, establishing production-grade MLOps capabilities...SuggestedPart timeShift work$200k - $322k
...tapping into the unlimited potential of AI to define the next era of computing. An... ...on the world.This Senior Staff Client Platform Engineer role is a high-impact technical leadership... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...SuggestedFull timePart time$272k - $431.25k
...increasingly known as “the AI computing company.” We... ...are looking for expert engineers to come and help design... ...AI supercomputing platforms.Join us at the forefront... ...Compute. Experience with ML and multi-variable... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...SuggestedFull timePart time- ...unleashing the potential of generative AI to power the transformation of... ...Hybrid, working onsite at our Santa Clara, CA, headquarters 3 days per... ...: PrincipalSystem Software Engineer - AI Inference ExecutionWhat... ...closely with other software (ML and compilers) and hardware...Part time3 days per week
$184k - $287.5k
...into the unlimited potential of AI to define the next era of... ...Group (SCG) is seeking Senior AI Platform Engineers. They will set the technical... ...at the intersection of ML infrastructure and large-scale... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time...Full timePart time$100k
...industry on cutting-edge AI technology,... ...software models, compilers, platforms, networking, and semiconductors... ...out of Toronto, ON; Santa Clara, CA; Austin, TX; Belgrade... ...or infrastructure engineer with experience building... ...workload orchestration, and ML services.Develop APIs...Permanent employmentPart time$184k - $253k
...a global leader in materials engineering solutions used to produce virtually... ...connect our world – like AI and IoT. If you want to push the... ...,000.00 - $253,000.00Location:Santa Clara,CAYou’ll benefit from a... ...learning engineer with a deep ML foundation who has actively kept...Full timePart time$195.2k - $361.2k
...journey is to transform AI into something safer, more... ...a Machine Learning Engineer / Data Scientist to join... ...engineering, data science or ML research.Experiences... ...Location: US, California, Santa ClaraAdditional... ...: US, California, Santa Clara; US, Oregon, Hillsboro;...Full timePart timeInternshipLocal areaImmediate startShift work- ...next-generation computing experiences—from AI and data centers, to PCs, gaming and... ...beyond. Together, we advance your career. PLATFORM THERMAL LEADER THE ROLE: We are looking... ...degree in mechanical engineeringLocation: Santa Clara, CAThis role is not eligible for visa sponsorship...Part time
$168k - $264.5k
...are looking for hardworking engineers who will craft FPGA prototypes... ...on standard FPGA prototyping platforms.We are now looking for a Senior... ...our Emulation team onsite in Santa Clara, CA.What you'll be doing:... ...deep learning ignited modern AI — the next era of computing....Full timePart time$232k - $310k
...is a global leader in AI-native enterprise talent... ...500 organizations. Our platform is built from the ground... ...high standards. Our engineers, product leaders, and go... ...world works.About AII/ML TeamOur AI/ML team is building... ...or hybrid in Zone A: Santa Clara, CA The base salary...Part timeWork experience placementWork at officeRemote workFlexible hours3 days per week$224k - $356.5k
...boundaries of multimodal AI, simulation, and world... ...meta-layer of modern ML: the agents, tooling, pipelines... ...for exceptional engineers who are passionate about... ...and scale evaluation platforms that combine automated... ...SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$224k - $356.5k
...tapping into the unlimited potential of AI to define the next era of computing. An era... ....NVIDIA is searching for a world-class engineer in graphics and AI to join our neural graphics... ...by law.SummaryLocation: US, CA, Santa Clara; US, NC, Durham; US, NC, RemoteType: Full...Full timePart time$272k - $431.25k
...seeking a Senior MLOps Engineering Manager to join our Autonomous... ...organization in Santa Clara, CA. This role offers an... ..., end‑to‑end data and ML pipelines that power NVIDIA... ..., CI/CD, and data platforms.What We Need to See:Bachelor... ...vacancy. NVIDIA uses AI tools in its recruiting...Full timePart time$176k - $276k
...intelligence.Join our team of innovative engineers who develop and maintain... ...highly motivated EngOps and Platform Engineers to boost execution... ...existing vacancy. NVIDIA uses AI tools in its recruiting... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$224k - $356.5k
...unlimited potential of AI to define the next era... ...experience. We own the platform — performance, CI/CD pipelines... ...Science, Computer Engineering, Electrical Engineering... ...in GPU computing, ML systems, or high-performance... ...: US, CA, Santa Clara; US, MA, Westford; US,...Full timePart timeLocal area$184k - $287.5k
...Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA is an engineering role to compose... ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes.... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart time$184k - $287.5k
...deep learning ignited modern AI — the next era of computing —... ...OpenBMC Firmware for GPU Server platforms focus on but not limited to... ...(or higher) in Electrical Engineering or Computer Science or equivalent... ...law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart time$272k - $431.25k
...into the unlimited potential of AI to define the next era of... ...a Principal Systems Software Engineer to join our team, comprised of... ...performance, reliability, cross-platform compatibility, and... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, Remote; US, WA, SeattleType...Full timePart timeRemote work$184k - $287.5k
...We are looking for a Senior Software Engineer to become part of our storage management plane... ..., GPU deep learning ignited modern AI — the next era of computing — with the GPU... ...characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart time$184k - $287.5k
...recently, GPU deep learning ignited modern AI—the next era of computing—with the GPU... .... We’re hiring a Senior Backend/Platform Engineer to build and maintain the core infrastructure... ...characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart timeRemote work$196k - $310.5k
...tapping into the unlimited potential of AI to define the next era of computing. An... ...world.NVIDIA is seeking an outstanding Platform Design Engineer to join our Circuit Solutions Group... ...characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$176.1k - $308.2k
...action on what data.Veza's Access Graph platform maps an organization's entire identity ecosystem... ...6.The combination brings together Veza's AI-native Access Graph with ServiceNow's AI... ..., cloud environments, and AI agents. For engineers joining Veza today, this means the scale...Part timeWork at officeRemote workFlexible hours$184k - $287.5k
...seeking a Senior System Software Engineer to lead the evolution of our... ...Data & Observability Platform. We serve and collaborate directly... ...with NVIDIA’s rapidly growing AI, HW, and SW engineering and... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, RemoteType...Full timePart time$208k - $327.75k
...Product Manager to lead strategic AI platform initiatives across... ...workflows. Partnering closely with engineering, architecture, and platform... ...products or AI/ML systems at scale.Experience... ...law.SummaryLocation: US, CA, Santa Clara; US, WA, Remote; US, CA, Remote...Full timeTemporary workPart timeRemote work$184k - $287.5k
...into the unlimited potential of AI to define the next era of... ...looking for a senior software engineer to join our Hardware Infrastructure... ...delivering highly available platform services and performant web applications... ...by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$184k - $287.5k
...seeking a Senior Machine Learning Engineer to join our end‑to‑end... ...into the unlimited potential of AI to define the next era of computing... ...large‑scale data pipelines for ML, including data quality monitoring... ...by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$184k - $287.5k
...driving vehicle powered by AI can meander through a... ...best Machine Learning Engineers with a background in... ...and coordinate entire ML workflows, covering data... ...on the NVIDIA DRIVE AV platform, developing highly efficient... ...: US, CA, Santa ClaraType: Full time...Full timePart timeWorldwideNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI/ML Platform Engineer (Santa Clara). Be the first to apply!
- ai developer Santa Clara, CA
- ai engineer Santa Clara, CA
- machine learning engineer Santa Clara, CA
- senior ml engineer Santa Clara, CA
- platform developer Santa Clara, CA
- platform engineer Santa Clara, CA
- platform product manager Santa Clara, CA
- power platform Santa Clara, CA
- platform manager Santa Clara, CA
- internship machine learning Santa Clara, CA









