Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal GenAI Inference Optimization Engineer

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative AI inference workloads on AMD GPU platforms. You will contribute to optimizing latency, throughput, and cost efficiency for real-world deployment of large-scale models, working across the software-hardware stack.THE PERSONThe ideal candidate is a strong technical contributor with expertise in GenAI inference optimization, GPU performance, and large-scale serving systems. You have a solid understanding of GPU architecture, memory systems, and communication patterns, and can apply this knowledge to improve inference efficiency.You are comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently drive optimization efforts while collaborating with cross-functional teams.KEY RESPONSIBILITIES- Optimize performance of GenAI inference workloads on AMD GPU platforms across single-node and distributed environments.- Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.- Analyze and resolve bottlenecks across compute, memory, and communication (e.g., kernel efficiency, KV-cache usage, memory bandwidth, scheduling).- Contribute to cross-stack optimizations spanning kernels, runtimes, communication libraries, and inference/serving frameworks (e.g., vLLM, SGLang, Triton, or similar systems).- Implement and evaluate inference optimization techniques such as batching strategies, quantization, prefix caching, and speculative decoding.- Support development and optimization of scalable serving systems, including request scheduling and resource utilization.- Develop and use profiling, benchmarking, and performance analysis tools for inference workloads.- Collaborate with hardware, compiler, and framework teams to improve overall system performance.- Contribute to internal tools and, where applicable, open-source projects for inference optimization on AMD platforms.- Document best practices and contribute to performance guidelines for GenAI deployment.PREFERRED EXPERIENCE- Strong understanding of GPU architecture and performance fundamentals (compute, memory hierarchy, interconnects such as PCIe/Infinity Fabric/RDMA).- Experience with GenAI inference optimization techniques (e.g., quantization, KV-cache optimization, batching).- Hands-on experience with inference/serving frameworks such as vLLM, SGLang, Triton, TensorRT-LLM, or similar.- Experience working on LLM or multimodal inference workloads.- Familiarity with distributed systems and serving architectures.- Experience with ML frameworks (PyTorch, JAX, or TensorFlow), especially for inference.- Proficiency in Python and at least one systems language (C++/CUDA/HIP).- Experience with profiling, debugging, and performance tuning tools.- Ability to work collaboratively across teams and deliver impactful optimizations.ACADEMIC CREDENTIALS- B.S., M.S. or Ph.D. in Computer Science, Computer Engineering, or a related field preferred, or equivalent industry experience.LOCATION- San Jose, CA#LI-MV1#HYBRIDThis role is not eligible for visa sponsorship.Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Principal GenAI Inference Optimization Engineer in San Jose, CA vacancy
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $195.2k - $361.2k

     ...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment... 
    Suggested
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  • $182.5k - $260.5k

     ...Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility...  ...roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient,... 
    Principal

    Netskope

    Santa Clara, CA
    3 days ago
  • $206.4k - $379.1k

     ...AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain. In this high...  ....Design and architect inference infrastructure for enterprise...  ...(training, inference, and/or optimization).Proven track record of leading... 
    Principal
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  •  ...specialized Sr. Staff or Principal level engineer who is passionate about enabling...  ...on scaling training and inference for the latest Generative...  ...training and inference optimizations across a variety of applications...  ...with state-of-the-art GenAI algorithms and software... 
    Suggested

    AMD

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who...  ...are mindful of performance analysis and optimization to help us squeeze every last clock...  ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices... 
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  •  ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated...  ...movement.AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging... 

    AMD

    Santa Clara, CA
    2 days ago
  •  ...Jose, California, United StatesProducts - Engineering /Fulltime /HybridWe are seeking an...  ...user-friendly end-to-end solutions for GenAI applications, combining cutting-edge backend...  ...and accessible interfaces. Develop and optimize backend services and APIs (REST, GraphQL... 
    Principal
    Full time

    Extreme Networks, Inc.

    San Jose, CA
    8 hours ago
  • $150k - $250k

     ...a significant plus.Need to work closely with system and test engineers to develop high speed interface, package/board, and system clocks...  ...Layout design and support. Need to get involved into layout optimizations for high speed or high precision performance directly.•... 
    Principal

    Omnivision Technologies

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class...  ...generation and computational graph optimizations for next-generation NVIDIA GPUs.Advance...  ...compilation problems for AI workloads (both inference and training) and successfully... 
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks,... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Santa Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be...  ...and understand the nuances of what it takes to optimize and trade-off various aspects of hardware-software co... 
    Principal
    3 days per week

    d-Matrix

    Santa Clara, CA
    2 days ago
  •  ....THE PERSON: The ideal candidate is passionate about software engineering and the craft of training performance. You lead sophisticated...  ...throughput, memory efficiency, and stability across data, model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication... 
    Principal

    AMD

    San Jose, CA
    4 days ago
  • $200k

    Role Overview We are looking for a Principal Packaging Engineer to own package architecture and development for Velaura’s next-generation SoCs...  ...substrate and assembly technologies, and cross-functional system optimization. \n Responsibilities ● Own package architecture from... 
    Principal
    Full time
    Flexible hours

    Velaura

    Santa Clara, CA
    4 hours ago
  • $152k - $241.5k

     ...NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other...  ...requirementsContribute to performance optimization and benchmarking efforts for specialized automotive... 
    Full time

    Nvidia

    Santa Clara, CA
    8 hours ago
  • $210.16k - $271.98k

    Senior Principal Systems Development EngineerHelp architect and deliver Dell's L11 rack-scale...  .... As a Principal Systems Development Engineer, you'll own the system-level architecture...  ..., and applies that insight to optimize complete AI rack solutions.Directs the application... 
    Principal

    Dell Technologies

    Santa Clara, CA
    3 days ago
  •  ...Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy... 

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full...  ...frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs,... 
    Principal
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $182k - $319k

     ...City: San Jose General OverviewFunctional Area: Engineering (ENG)Career Stream: Engineering (ENG)Role: Senior Principal (SPR)Job Title: Senior Principal, Design...  ...VR13 design.Well understand power efficiency optimization.Experience of products development from prototype... 
    Principal
    Local area

    Celestica

    San Jose, CA
    1 day ago
  • $150k - $190k

     ...Principal Test Engineer \\ \ \ \ San Jose, CA | Hybrid (3 days onsite \/ 2 remote)\\ \ Base Salary: $150,000–$190,000 (DOE)\\ \...  ...ramps at OSAT partners (domestic and offshore)\ . \\ Optimize yield, test time, and multisite efficiency \ to improve cost... 
    Principal
    Full time
    Work experience placement
    Remote work

    Talentry

    San Jose, CA
    11 hours ago
  •  ...Integrations is seeking an ambitious, highly motivated Principal AI/Data Center Systems Engineer to help define and develop next-generation high-...  ...technical challenges, evaluate design tradeoffs, and optimize system performance, density, and reliability. Serve as... 
    Principal
    Worldwide
    Shift work

    Power Integrations

    San Jose, CA
    2 days ago
  • $160.2k - $240k

     ...thrive, learn, and lead. Your Team, Your ImpactAs the Senior Principal test engineer within the Operations, You’lldefines the IO Chiplet ATE...  ...testability and test strategies, contributing to improved yield, optimized test methodologies, and cross‑functional innovation.Convert... 
    Principal
    Permanent employment
    Full time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    4 days ago
  • $182k - $260k

     ...shape the future of cybersecurity.RoleWe are looking for a Principal GenAI Data Engineer to join our team. This is a Hybrid role based in San Jose...  ..., securing, or positioning AI-driven solutions to optimize outcomes within your functional domainExpert-level Python... 
    Principal
    Full time
    Work at office
    Local area
    Remote work

    Zscaler

    San Jose, CA
    1 day ago
  • $184k - $287.5k

     ...best work. Come join the team and see how you can make a lasting impact on the worldWe are looking for an experienced Compiler Optimization Engineer for an exciting role in our Compute Compiler Team. We deliver features and improvements to CUDA and other compute compilers... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...deliver industry-leading training and inference speeds; over 10 times faster than GPU-based...  ...RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an...  ...SGLang, TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton), and an... 
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...THE ROLE:  We are looking for a dynamic, energetic Lead / Principal Systems Design Engineer to join our growing team. As a key contributor to the...  ...test execution to make sure all features are validated and optimized on time PREFERRED EXPERIENCE:  Programming/scripting... 
    Principal

    AMD

    San Jose, CA
    8 hours ago
  • $158.6k - $234.65k

     ...reliability and performance.Role Overview:We are seeking a Sr. Principal Engineer, Manufacturing Infrastructure to join our Advanced...  ...internal manufacturing infrastructure remains at the cutting edge.Optimize solutions for manufacturing yield, repeatability, and throughput... 
    Principal
    Permanent employment
    Full time
    Internship
    Work from home
    Shift work

    Marvell

    Santa Clara, CA
    3 days ago
  • $170k

     ...A leading chip and silicon IP provider is seeking a NPI Principal Product Engineer to join our Operations team in San Jose. In this role, you...  ...support. Experience in Failure Analysis and RMA support. Optimize Yield, Cost, Cycle Time and Quality. Data Analysis proficiency... 
    Principal
    Remote work
    Relocation package
    3 days per week

    OSI Engineering

    San Jose, CA
    2 days ago
  • $144k - $216k

     ..., quality, and reliability issues. Onto Innovation strives to optimize customers’ critical path of progress by making them smarter, faster...  ...: Lead and mentor a multidisciplinary team of hardware engineers, fostering a culture of technical growth, innovation, and accountability... 
    Principal
    Permanent employment
    Full time

    Onto Innovation

    Milpitas, CA
    3 days ago
  •  ...we are redefining the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,...  ...Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in... 
    Principal
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal GenAI Inference Optimization Engineer. Be the first to apply!