Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Infrastructure Engineer

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE: We are seeking a DevOps / Platform Engineer to join our team building and operating large-scale GPU compute infrastructure that powers AI and ML workloads. THE PERSON: The ideal candidate should be passionate about software engineering and possess leadership skills to independently deliver on multi-quarter projects. They should be able to communicate effectively and work optimally with their peers within our larger organization. Finally, you aren't afraid of a team in more of a startup mode at a larger company and willing to jump in to help in areas adjacent to your main project as needed.KEY RESPONSIBILITIES:Build and extend platform capabilities to enable new classes of workloads (e.g., interactive development pods, CI pipelines, inference services, benchmarking jobs).Design and operate scalable orchestration systems using Kubernetes across both on-prem and multi-cloud environments.Develop platform features such as secret management, configuration management, and deployment automation for customers.Partner with development teams to extend the GPU developer platform with features, APIs, templates, and self-service workflows that streamline job orchestration and environment management.Manage service lifecycle within Kubernetes using Helm and GitOps workflows (e.g., ArgoCD or Flux).Apply expertise in storage and networking to design and integrate CSI drivers, persistent volumes, and network policies that enable high-performance GPU workloads.PREFERRED EXPERIENCE:Experience in DevOps, Platform, or Infrastructure Engineering.Deep hands-on experience with Kubernetes and container orchestration at scale.Proven ability to design and deliver platform features that serve internal customers or developer teamsExperience building developer-facing platforms or internal developer portals (e.g.custom workflow tooling).Hands-on experience in storage or network engineering within Kubernetes environments (e.g., CSI drivers, dynamic provisioning, CNI plugins, or network policy).Experience with Infrastructure as Code tools like Terraform.Background in HPC, Slurm, or GPU-based compute systems for ML/AI workloads.Practical experience with monitoring and observability tools (Prometheus, Grafana, Loki, etc.).Understanding of machine learning frameworks (PyTorch, vLLM, SGLang, etc.).PREFERRED ACADEMIC CREDENTIALS: Bachelors or Masters degree in Computer Science, Computer Engineering, or a related field.LOCATION: San Jose, CA #LI-G11 #LI-HYBRIDBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 14 days ago
Similar jobs that could be interesting for youBased on the AI Infrastructure Engineer in San Jose, CA vacancy
  •  ...small smart high-end electric cars with the FIREFLY brand. About the Position   We are looking for a senior AI Inference Infrastructure Software Engineer with strong hands-on experience building, optimizing, and deploying high-performance, scalable inference systems... 
    Suggested
    Full time
    Temporary work
    Immediate start
    Flexible hours

    NIO USA, INC

    San Jose, CA
    4 days ago
  • $274k - $304k

     ...Saviynt's AI-powered identity platform manages and governs human and non-human access...  ..., please visit AI Platform Engineer - Training & Inference Saviynt's AI...  ...SLMs and cloud LLMs • Build RL training infrastructure: define Flyte workflows for RL pipelines... 
    Suggested

    Saviynt

    Milpitas, CA
    2 days ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (Gen AI Platform Services, Agentic AI) Overview: At Capital One, we are creating responsible and reliable AI...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    1 day ago
  • $229.9k - $262.4k

     ...Senior Lead AI Engineer (Gen AI Platform Services) Overview: At Capital One, we are creating responsible and reliable AI systems...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Suggested
    Full time
    Part time
    Local area

    Capital One Financial Corp

    San Jose, CA
    2 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (GenAI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    San Jose, CA
    1 day ago
  • $221.2k - $387.1k

     ...Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She...  ...so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform... 
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    4 days ago
  •  ...Job Title: AI Infrastructure Systems Engineer Location: San Jose, CA Contract JD: Highly skilled AI Infrastructure Systems Engineer to lead the deployment, provisioning, and optimization of our next-generation NVIDIA GB300 NVL72 AI supercomputer... 
    Contract work

    VDart Inc

    San Jose, CA
    2 days ago
  •  ...Capital One is seeking a Senior Distinguished Engineer to architect and scale a multi-tenant AI/ML platform. You will develop Ray and Spark-based compute engines, drive operational excellence, and lead a portfolio of advanced ML initiatives across the Capital One ecosystem... 
    Remote job

    Jobleads-US

    San Jose, CA
    5 days ago
  •  ...using it to some extent and at TransPerfect, we are no exception. We’re building an advanced voice processing platform that leverages AI for natural voice recognition, synthesis, and interaction. As part of the team, you’ll help design and develop scalable solutions... 
    Full time

    TransPerfect

    San Jose, CA
    a month ago
  •  ...that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded...  ...career. THE ROLE:We are hiring a AI Research Scientist - Infrastructure Engineer, Reinforcement Learning, to own reinforcement learning infrastructure... 

    AMD

    Santa Clara, CA
    a month ago
  • $178k - $321k

     ...Audit function has an early but working AI-native capability: a multi-agent platform...  ...setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated...  ...just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your... 

    OKX

    San Jose, CA
    a month ago
  • $184k - $287.5k

     ...tapping into the unlimited potential of AI to define the next era of computing. An...  ...Group (SCG) is seeking Senior AI Platform Engineers. They will set the technical direction...  ...foundational platforms at the intersection of ML infrastructure and large-scale systems, this is your... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...Group Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive...  ...seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role, you... 
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    26 days ago
  • $40 - $85 per hour

     ...across a wide variety of genres and languages. The AI Platform team builds the infrastructure that Netflix's ML and AI systems run on, from large-...  ...Systems, Systems, Networking, Machine Learning, Computer Engineering, or a related field Research or applied experience... 
    Hourly pay
    Full time
    Internship
    Immediate start
    Remote work
    Flexible hours

    Netflix

    Los Gatos, CA
    21 hours ago
  • $240k - $260k

     ...SAVIYNT Saviynt is a leader in identity security, delivering an AI-powered platform that governs and secures access to applications...  ...is governed across the AI Platform. You define the standards ML engineers and scientists build on, and ensure every training signal is... 

    Saviynt

    Milpitas, CA
    26 days ago
  • $183.6k - $297k

     ...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and...  ...of integrating AI into cybersecurity infrastructure — building intelligent systems that...  ...incidents at scale. As a Principal Software Engineer, you will own the technical vision for... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    7 days ago
  •  ...are builders, makers, and innovators helping enterprises move beyond AI experimentation and into real-world impact through enterprise AI platforms, products, and services. We combine deep engineering expertise with AI innovation to help clients modernize, build intelligent... 

    Publicis Media

    San Jose, CA
    5 days ago
  •  ...next-generation computing experiences—from AI and data centers, to PCs, gaming and...  ...THE ROLE:We are hiring AI / ML Platform Engineers to build the platform layer that makes AI...  ...reproducible. This role focuses on the infrastructure and platform systems that support large-... 

    AMD

    Santa Clara, CA
    a month ago
  • $190k - $260k

     ...has developed an artificial intelligence (AI) powered technology stack purpose-built...  ...large-scale world models – depends on infrastructure that turns thousands of hours of multimodal...  ...training throughput. We are looking for engineers who make model training fast: streaming... 
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    14 days ago
  • $140k - $165k

     ...most advanced electronic devices and IT infrastructure, enabling enhanced performance and user...  ...Why Join Us? Build foundational AI infrastructure that powers next-gen enterprise...  ...the Role: We are seeking a hands-on AI Engineer to design, deploy, and maintain on-prem... 

    SK hynix memory solutions America Inc.

    San Jose, CA
    29 days ago
  •  ...Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to providing comprehensive Bitcoin mining...  ...We are seeking a Senior AI Storage Infrastructure Engineer to build the critical data-delivery fabric of our AI-native... 
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    20 days ago
  •  ...Title: Prinicipal AI Engineer Location: Sunnyvale, California. Duration: 6 to 12+ Months Job Description:...  ...Proficient PostgreSQL, embeddings/vector search Cloud/Infrastructure Proficient GCP (BigQuery, GCS), Kubernetes, CCM Integration... 
    Contract work

    Redolent

    Sunnyvale, CA
    4 days ago
  •  ...wants to learn fast, take ownership, and grow beyond limits, you'll feel right at home here.   Role Overview The AI-Native Software Engineer, Cloud (AWS)will join a security engineering team responsible for building tools that discover, mitigate, and report on security... 

    Vailexa

    San Jose, CA
    13 days ago
  • $184k - $287.5k

     ...leading cloud product that powers innovative AI research and developers. We focus on...  ...workloads, as well as developing scalable AI infrastructure services globally. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $120k - $220k

     ...news and information powered by advanced AI, recommendation systems, and adtech....  ...team to fulfill our mission: building the infrastructure layer for content intelligence. If you...  ...hiring our  first dedicated Agent Platform engineer to own this layer end-to-end. You'll... 
    Full time
    Local area
    Work from home

    NewsBreak

    Mountain View, CA
    18 days ago
  • $255.65k

     ...Immigration sponsorship is not available for this position What you can expect: As a Senior AI Software Engineer, you will collaborate to design, implement, and optimize AI algorithms and software applications. You will ensure AI training, inference, deployment,... 
    Work at office
    Remote work

    Zoom

    San Jose, CA
    21 hours ago
  •  ...Forward Deployed Engineer As a Forward Deployed Engineer at TENEX, you are a customer-embedded problem solver. You live inside strategic...  ...what is broken or missing, and build solutions — combining AI, data, and security expertise to ship things that matter. Your work... 

    TenEx

    San Jose, CA
    3 days ago
  •  ...manufacturer, responsible for shaping the vision and implementation of AI-powered capabilities across our product portfolio. We're...  ...intersection of intelligent systems and real-world networking infrastructure. Candidates must be located in the Bay Area or San Diego, CA... 

    Allied Telesis

    San Jose, CA
    3 days ago
  • $165.2k - $223.6k

    AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the...  ...’ll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers... 
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon.com Services LLC

    Santa Clara, CA
    -127
  • $150k - $210k

     ...measurable business value.About the Position – Embodied AI EngineerWe are seeking an exceptional Embodied AI Engineer to build the foundation of LG's vision of...  ...of embodied AI engineering, building the infrastructure, pipelines, and training systems that let research... 
    Full time
    Temporary work
    For contractors
    Local area
    Immediate start

    LG Electronics

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Infrastructure Engineer. Be the first to apply!