Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. AI/ML Platform Engineer

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. This role focuses on the infrastructure and platform systems that support large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation, artifact management, and GPU cluster utilization.You will work closely with ML Systems Research Engineers, AI Research Scientists, Applied AI Engineers, and hardware domain experts to operationalize the Blueprint framework across kernel optimization, RTL/PPA optimization, ECO fixing, verification, simulation, and debugging workflows.This is a platform engineering role, not a pure research role. The focus is to build robust shared systems that allow researchers and engineers to run more experiments, compare results reliably, reduce manual orchestration, and move successful workflows into production engineering use.THE PERSON:You are a strong systems engineer who enjoys building reliable platforms for AI researchers and applied engineers. You understand distributed systems, ML workloads, GPU infrastructure, experiment management, and production reliability. You can turn messy research workflows into reusable services, APIs, dashboards, job systems, and automation.You care about reproducibility, observability, performance, and developer experience. You are comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without requiring every requirement to be fully specified upfront.KEY RESPONSIBILITIES:Build and operate the shared AI platform for agentic engineering workflows, including job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking.Develop reliable infrastructure for distributed training, distributed inference, batch evaluation, and large-scale agent rollout across GPU clusters.Build platform services for benchmark execution, correctness checking, profiling, regression tracking, and reproducible evaluation.Maintain artifact systems for generated kernels, RTL edits, traces, logs, profiler outputs, benchmark results, simulator outputs, and formal verification artifacts.Support scalable integrations with compilers, ROCm/HIP tooling, profilers, simulators, EDA tools, vLLM, SGLang, and internal engineering systems.Improve GPU cluster utilization, scheduling efficiency, reliability, quota management, and workload isolation.Build dashboards and observability systems for experiment status, resource usage, failure modes, benchmark trends, regression detection, and team productivity.Partner with ML Systems Research Engineers to productionize research workflows for RL systems, inference systems, quantification systems, and evaluation pipelines.Partner with Applied AI Engineers to make Blueprint harnesses reusable across kernel optimization, RTL optimization, verification, firmware, and CPU/GPU performance workflows.Establish platform standards for reproducibility, data retention, run metadata, artifact lineage, access control, and operational reliability.TECHNICAL FOCUS AREAS:Distributed ML platform infrastructure for training, inference, evaluation, and agent execution.GPU cluster scheduling, utilization, reliability, quota management, and multi-user workload isolation.Experiment tracking, artifact management, run lineage, dashboards, and reproducible workflow management.Benchmark and evaluation automation for correctness, performance, regression detection, and reproducibility.Integration with ROCm/HIP, Triton, compilers, profilers, vLLM, SGLang, simulators, formal tools, and EDA flows.Caching and parallelization for expensive feedback loops, including simulator, compiler, verifier, and benchmark workloads.Production-quality APIs, services, workflow engines, and developer tooling for research and engineering teams.PREFERRED QUALIFICATIONS:Strong programming skills in Python and one or more systems languages such as C++, Go, or Rust.Experience building ML platforms, AI infrastructure, distributed systems, workflow orchestration, experiment platforms, GPU cluster infrastructure, or developer platforms.Strong understanding of job scheduling, distributed workloads, logging, monitoring, reliability, storage systems, and production operations.Experience with Kubernetes, Ray, Slurm, workflow engines, containerization, CI/CD, data pipelines, or large-scale compute orchestration.Ability to build reliable services, APIs, dashboards, and developer tools used by researchers and engineers.Strong debugging skills across distributed systems, GPU workloads, storage systems, networking, containers, and production infrastructure.Good collaboration skills with AI researchers, ML systems researchers, applied engineers and hardware domain experts.PREFERRED EXPERIENCE:Experience with GPU platforms, ROCm/HIP, CUDA, profiling, kernel benchmarking, model serving, or distributed training/inference.Experience with vLLM, SGLang, Triton, PyTorch, JAX, Ray, Kubernetes, Slurm, MLflow, Weights & Biases, or similar systems.Experience building experiment tracking systems, artifact stores, workflow engines, benchmark automation platforms, or developer productivity tools.Experience supporting LLM agents, tool-use systems, large-scale sampling, or automated program optimization workflows.Familiarity with compiler, profiler, simulator, formal verification, EDA, firmware, or hardware performance workflows is a strong plus.Experience operating shared GPU clusters or high-performance ML infrastructure in a multi-user research environment.EDUCATIONBachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience. Master's preferred; PhD is a plus, especially with work in ML systems, reinforcement learning, distributed systems, GPU computing, or AI infrastructure.LOCATION: Santa Clara, CA#LI-BW1 #LI-HybridBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Sr. AI/ML Platform Engineer in Santa Clara, CA vacancy
  •  ...scientific discovery to powering AI and the technologies people...  ...has a broad portfolio of several ML accelerator devices with NPUs (...  ...product portfolio. The Vitis unified platform (Vitis Unified Software...  ...based of AMD's proprietary AI-engine architecture.THE PERSON:The ideal... 
    Senior

    AMD

    San Jose, CA
    6 days ago
  • We are seeking a talented Machine Learning Engineer with expertise in software engineering to join our team. As a Machine Learning Engineer...  ...- programming languages (GO/GRPC, Python) GCP Vertex AI Components and SDK Experience Experience with React UI is needed... 
    Senior

    Trinity Technology Solutions

    San Jose, CA
    5 days ago
  • $183.6k - $297k

     ..., Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to...  ...incidents at scale. As a Principal Software Engineer, you will own the technical vision for our AI-powered platform, shaping how agentic workflows and LLM-driven automation... 
    Senior
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    6 days ago
  • $203k - $258.6k

     ...on: 09/07/2026Meet the TeamCX AI Incubation team is part of CX...  ...requires strong expertise in AI/ML, software development, and a...  ...collaborating with product management and engineering teams to deliver impactful,...  ...on various AI cloud platforms such as AWS SageMaker, Google... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    19 days ago
  • An innovative firm is seeking a Senior Machine Learning Engineer to lead the development of cutting-edge machine learning models that enhance...  ...with diverse teams, you will drive the implementation of AI solutions while mentoring junior engineers. If you are passionate... 
    Senior

    Jobleads-US

    San Jose, CA
    4 days ago
  •  ...machine learning products. Build ML infrastructure, frameworks,...  ...projects and deliver ML platform capabilities. Requirements...  ...related field plus 3+ years of engineering experience and 5+ years of machine...  ...training, privacy-preserving ML, and agentic AI familiarity.... 
    Senior
    Full time
    Work experience placement

    Apple

    Cupertino, CA
    3 days ago
  • $175k - $263k

     ...has evolved from storage pioneer to data platform, closing fiscal 2026 with $3.7 billion in...  ...the next era of technology - hyperscalers, AI labs, the AI hardware supply chain, data...  ...capabilities. Availability-Focused Engineering: Understanding of non-disruptive upgrade... 
    Senior
    Work at office
    Flexible hours

    Everpure, Inc.

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...harnessing the boundless capabilities of AI to build the next era of computing. An era...  ...systems at scale. We seek a Senior ML Engineer to compose and deliver next-generation Metropolis...  ...Cosmos, NVIDIA's world foundation model platform. Your main task will be constructing... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    18 days ago
  • $193.3k - $261.5k

     ...tools used for building Generative AI on AWS. The Inferentia chip delivers best-in-class ML inference performance at the...  ...multiple disciplines including silicon engineering, hardware design and verification...  ...leap in performance.You: As a Sr. Machine Learning Compiler Engineer... 
    Senior
    Internship
    Local area
    Work from home
    Relocation
    Flexible hours

    Amazon

    Cupertino, CA
    13 days ago
  • $193.3k - $261.5k

    Do you want to be part of AI revolution? At AWS our vision is to...  ...the performance of complex ML models executed on AWS Inferentia...  ...role is for a senior software engineer in the Compiler team for AWS...  ...Storage, Internet of Things (Iot), Platform, and Productivity Apps... 
    Senior
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    a month ago
  •  ...enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across hybrid multi...  ...that keep the world running. Our Team's Vision Our Engineering team is shaping the future of cybersecurity. We... 
    Senior
    Immediate start

    Illumio

    Sunnyvale, CA
    5 days ago
  • $184.7k - $324.8k

    Sr. Machine Learning Engineer, Siri Speech Join the team redefining what a deeply personal...  ...world's most widely used AI assistants, powered by our...  ...strong background in applied ML research and development, particularly...  ...features across Apple platforms. Collaborate closely with... 
    Senior
    Worldwide
    Relocation

    Apple

    Cupertino, CA
    5 days ago
  • Sr. Machine Learning Engineer It started with a simple idea: what if surgery could be...  ...is Intuitive's new robotic platform for minimally invasive biopsy...  ...environments. Expert knowledge of ML frameworks and software...  ...knowledge of the development of AI/ML features in strict... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive

    Sunnyvale, CA
    5 days ago
  • $193.3k - $261.5k

     ...forefront of maximizing performance for AWS's custom ML accelerators. Working at the hardware-software boundary, our engineers craft high-performance kernels for ML functions,...  ...to push the boundaries of what's possible in AI acceleration.The AWS Neuron SDK, developed by... 
    Senior
    Internship
    Local area
    Work from home
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $153.2k - $234.1k

     ...autonomous driving? Join the Embodied AI team at General Motors. Our team is developing...  ...real-world scenarios. As a Senior ML Infra Engineer, you will work on the core systems that...  ..., and high‑performance training platforms you help create. What You’ll Do:Design,... 
    Senior
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    a month ago
  • $150k - $225k

     ...overall quality of life. As a Senior AI / Embedded Engineer, you will be responsible for the full...  ...building reliable, low-power, real-time ML systems that operate at the edge. In...  ...following: ~• Experience with RTOS platforms such as FreeRTOS or Zephyr • Familiarity... 
    Senior
    Full time
    Work at office
    Immediate start
    Visa sponsorship
    Night shift

    E-Space

    Saratoga, CA
    1 day ago
  •  ...Who we are Moveworks:  the Agentic AI Assistant platform that empowers the entire workforce....  ...workflow automation with Moveworks’ Reasoning Engine and natural language capabilities, we...  ...it isn't a testing role. It's applied ML at a point where the methodology genuinely... 
    Senior
    Work at office
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    24 days ago
  •  ...development, and deployment of advanced AI agents and agentic systems. Architect and...  ...product managers, UX designers, and other engineers to define requirements and deliver impactful...  ...or PyTorch. Experience with cloud platforms (AWS) and containerization technologies (... 
    Senior
    Full time
    Work experience placement

    Eightfold

    Santa Clara, CA
    1 day ago
  • $140k - $230k

     ...Arene, our software development platform for software-defined vehicles;...  ...for mobility; and Cloud & AI, the digital infrastructure powering...  ...efficient machine learning (ML) training and evaluation...  ...looking for a skilled Software Engineer or Machine Learning Engineer to... 
    Senior
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    Woven by Toyota

    Palo Alto, CA
    25 days ago
  •  ...Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley...  ...Design efficient and effective solutions to a wide range of engineering challenges. Ship complex software in fast-paced environments... 
    Senior
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    4 days ago
  • Job Title: Senior AI/ML Engineer Work Location with ZIP: Sunnyvale, CA 94085 (Hybrid) Technical Hiring Criteria (Must Haves) Top 3 Required skills: Machine Learning, AI, Python Years of experience in each of the must-have skills: 7 Years Job description: - Strong programming... 
    Senior

    eTeam

    Sunnyvale, CA
    5 days ago
  • $140k - $224.25k

    We are looking for a highly motivated AI/ML Software Engineer to join the Enterprise Agentic AI Platform team within IT. You will work closely with Business Analysts, and Engineering teams to design, develop, and deploy enterprise AI solutions that improve productivity... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    6 days ago
  •  ...Capital One is seeking a Senior Distinguished Engineer to architect and scale a multi-tenant AI/ML platform. You will develop Ray and Spark-based compute engines, drive operational excellence, and lead a portfolio of advanced ML initiatives across the Capital One ecosystem... 
    Senior
    Remote job

    Jobleads-US

    San Jose, CA
    4 days ago
  • $170k - $277k

     ...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use...  ...looking for a visionary Senior Principal Engineer/Architect to serve as the technical authority for our global SRE and Platform Engineering initiatives across the US and India... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Work visa
    Flexible hours

    Jobleads-US

    Santa Clara, CA
    6 days ago
  •  ..., high-performance computing, cloud, and AI. Whether you’re designing next-gen processors...  ...CPU, GPU, AI, and adaptive computing platforms. This technology provides software-executable...  ...a Virtual Platform Functional Modeling Engineer, you will play a key role in defining and... 
    Senior

    AMD

    San Jose, CA
    6 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good...  ...their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed... 
    Senior
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    1 day ago
  •  ...We Are  Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading silicon...  ...tomorrow.  You Are You are a strong platform engineer with a passion for building... 
    Senior

    Synopsys Inc

    Sunnyvale, CA
    9 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (GenAI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years...  ...their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    5 days ago
  •  ...We Are  Synopsys is the leader in engineering solutions from silicon to systems, enabling customers to rapidly innovate AI-powered products. We deliver industry-leading...  ...thoughtful defaults—because you know that the best platforms are the ones that make the right thing the... 
    Senior
    Work at office
    Relocation

    Synopsys Inc

    Sunnyvale, CA
    20 days ago
  • $174k - $253k

     ...accessible technologies. Google's software engineers develop the next-generation technologies...  ..., and enhance software solutions.The AI and Infrastructure team is redefining what...  ...global services, and providing the essential platforms that enable developers to build the... 
    Senior
    Worldwide

    Google

    Sunnyvale, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. AI/ML Platform Engineer. Be the first to apply!