Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal GenAI Inference Optimization Engineer

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative AI inference workloads on AMD GPU platforms. You will contribute to optimizing latency, throughput, and cost efficiency for real-world deployment of large-scale models, working across the software-hardware stack.THE PERSONThe ideal candidate is a strong technical contributor with expertise in GenAI inference optimization, GPU performance, and large-scale serving systems. You have a solid understanding of GPU architecture, memory systems, and communication patterns, and can apply this knowledge to improve inference efficiency.You are comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently drive optimization efforts while collaborating with cross-functional teams.KEY RESPONSIBILITIES- Optimize performance of GenAI inference workloads on AMD GPU platforms across single-node and distributed environments.- Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.- Analyze and resolve bottlenecks across compute, memory, and communication (e.g., kernel efficiency, KV-cache usage, memory bandwidth, scheduling).- Contribute to cross-stack optimizations spanning kernels, runtimes, communication libraries, and inference/serving frameworks (e.g., vLLM, SGLang, Triton, or similar systems).- Implement and evaluate inference optimization techniques such as batching strategies, quantization, prefix caching, and speculative decoding.- Support development and optimization of scalable serving systems, including request scheduling and resource utilization.- Develop and use profiling, benchmarking, and performance analysis tools for inference workloads.- Collaborate with hardware, compiler, and framework teams to improve overall system performance.- Contribute to internal tools and, where applicable, open-source projects for inference optimization on AMD platforms.- Document best practices and contribute to performance guidelines for GenAI deployment.PREFERRED EXPERIENCE- Strong understanding of GPU architecture and performance fundamentals (compute, memory hierarchy, interconnects such as PCIe/Infinity Fabric/RDMA).- Experience with GenAI inference optimization techniques (e.g., quantization, KV-cache optimization, batching).- Hands-on experience with inference/serving frameworks such as vLLM, SGLang, Triton, TensorRT-LLM, or similar.- Experience working on LLM or multimodal inference workloads.- Familiarity with distributed systems and serving architectures.- Experience with ML frameworks (PyTorch, JAX, or TensorFlow), especially for inference.- Proficiency in Python and at least one systems language (C++/CUDA/HIP).- Experience with profiling, debugging, and performance tuning tools.- Ability to work collaboratively across teams and deliver impactful optimizations.ACADEMIC CREDENTIALS- B.S., M.S. or Ph.D. in Computer Science, Computer Engineering, or a related field preferred, or equivalent industry experience.LOCATION- San Jose, CA#LI-MV1#HYBRIDThis role is not eligible for visa sponsorship.Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Principal GenAI Inference Optimization Engineer in San Jose, CA vacancy
  •  ...and create.We’re seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of GenAI Products & Knowledge Innovations...  ..., scalable services.Optimize latency, throughput, and resource...  ...-scale model training, inference pipelines, or search/retrieval... 
    Principal
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    4 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $124.5k - $272k

     ...Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility...  ...roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient,... 
    Principal

    Netskope

    Santa Clara, CA
    3 days ago
  • $150k - $190k

     ...Principal Test Engineer \\ \ \ \ San Jose, CA | Hybrid (3 days onsite \/ 2 remote)\\ \ Base Salary: $150,000–$190,000 (DOE)\\ \...  ...ramps at OSAT partners (domestic and offshore)\ . \\ Optimize yield, test time, and multisite efficiency \ to improve cost... 
    Principal
    Full time
    Work experience placement
    Remote work

    Talentry

    San Jose, CA
    7 hours ago
  • $272k - $431.25k

     ...: models, adaptation workflows, and inference pipelines built for real inspection...  ...everyday constraints.We’re seeking a Principal Software Engineer for Systems Inspection in Santa Clara...  ..., model compression, and deployment optimization. You’ll join a small, high-impact... 
    Principal
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    6 days ago
  • $206.4k - $379.1k

     ...AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain. In this high...  ....Design and architect inference infrastructure for enterprise...  ...(training, inference, and/or optimization).Proven track record of leading... 
    Principal
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Jose, CA
    2 days ago
  •  ...Santa Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be...  ...and understand the nuances of what it takes to optimize and trade-off various aspects of hardware-software co... 
    Principal
    3 days per week

    d-Matrix

    Santa Clara, CA
    a month ago
  •  ...NVIDIA Corporation in Santa Clara is seeking a Principal Software Engineer for Systems Inspection to advance AI products for semiconductor analysis...  ...vision, multimodal AI, anomaly detection, and deployment optimization. Lead architecture, collaborate across teams, and... 
    Principal

    Jobleads-US

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency...  ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks,... 
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...Principal Packaging EngineerWe are looking for a Principal Packaging Engineer to own package architecture and development for Velaura's next-generation SoCs. This role...  ...technologies, and cross-functional system optimization.ResponsibilitiesOwn package architecture from... 
    Principal
    Flexible hours

    Auradine

    Santa Clara, CA
    3 days ago
  •  ...is seeking a full-time Deep Learning Algorithm Engineer passionate about pioneering DL, foundation models, and GenAI for image processing in the semiconductor...  ...experience with DL model development, training, optimization, and deployment, with emphasis on performance... 
    Full time

    KLA-Belgium

    Milpitas, CA
    2 days ago
  •  ...performance computing, AI, and semiconductor products. As a Principal Mechanical Engineer, you will lead the development of advanced package...  ...thermal, reliability, manufacturing, and supplier teams to optimize product performance and manufacturability.Conduct and lead... 
    Principal

    AMD

    San Jose, CA
    14 hours ago
  • $205k - $260k

     ...Responsibilities Architect and optimize cloud-native CI/CD pipelines and shared automation...  ..., and timing-closure workflows for engineering throughput and cloud cost efficiency....  ...and benefits. The role offers a broad principal-level mandate spanning cloud-native CAD... 
    Principal
    Full time

    Astera Labs

    San Jose, CA
    11 days ago
  •  ...Role Overview  We are looking for an experienced Principal Systems Integration Engineer to join the Arycs’ GTM team. In this role, you will serve...  ...and execute comprehensive customer test plans to ensure optimal product performance.  Qualifications  ~7+ years of... 
    Principal
    Flexible hours

    Arycs Technologies, Inc.

    Los Gatos, CA
    more than 2 months ago
  •  ...and integration limits. Our systems are engineered for deployment in real networks — not...  ...optical technologies. About the Role As a Principal Hardware Engineer, Signal Processing,...  ...for developing DSP algorithmic models, optimizing signal chains, validating performance... 
    Principal
    Work at office

    Arycs Technologies, Inc.

    Los Gatos, CA
    more than 2 months ago
  •  ...Description Renesas is seeking a Principal Test Architect for our Analog and Mixed...  ..., test infrastructure, and process optimization across Renesas' analog and mixed-signal...  ...closely with design, product, and test engineering teams to drive test innovation from new... 
    Principal
    Full time
    Flexible hours

    Renesas Electronics

    San Jose, CA
    3 days ago
  •  ...Adobe Applied Science & Machine Learning (ASML) seeks a Principal Scientist, ML – Overall Architect to own end-to-end architecture across training, inference, model building and data pipelines. You will serve as the principal technical owner bridging infrastructure, modeling... 
    Principal

    Jobleads-US

    San Jose, CA
    6 days ago
  •  ...Role Overview We are seeking a Wafer Bonding Integration Engineer to lead the development, integration, and manufacturing transfer...  ...achieved. Wafer Bonding Process Integration ~ Develop and optimize advanced wafer bonding processes including: ~ Fusion Bonding... 
    Principal

    nEye.ai

    Santa Clara, CA
    23 days ago
  • $148.32k - $203.94k

     ...For more information, visit: . Job Summary   The Principal Test Engineer is responsible for the development of production test solutions...  ...collaborate with cross-functional teams to develop highly optimized test solutions enabling SiTime’s game-changing timing... 
    Principal
    Work experience placement

    SiTime Corporation

    Santa Clara, CA
    18 days ago
  • $217k - $359k

     ...and quality of enterprise-class storage hardware systems that power mission-critical data environments. You will bridge complex engineering concepts with real-world deployment by leading cross-functional alignment across internal engineering, Joint Development Manufacturing... 
    Principal
    Work at office
    Flexible hours

    Everpure, Inc.

    Santa Clara, CA
    4 days ago
  • $182k - $319k

     ...and the top 3 non-X86 server providers engineering solutions for this generation and the next...  ...the future together. The Senior Principal, Design Engineering will be responsible...  .... ~ Well understand power efficiency optimization. ~ Experience of products development... 
    Principal
    Local area

    Celestica International LP

    San Jose, CA
    more than 2 months ago
  • $248k - $396.75k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing...  ...of production environments.As a Principal SRE, you will shape the technical direction...  ...AI/ML platforms, GPU infrastructure, inference systems, training environments, or high... 
    Principal
    Full time

    NVIDIA

    Santa Clara, CA
    14 hours ago
  • $184k - $287.5k

     ...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who...  ...are mindful of performance analysis and optimization to help us squeeze every last clock...  ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices... 
    Full time

    Nvidia

    Santa Clara, CA
    7 hours ago
  •  ...index serving, query processing, and ML inference. Turn advances in information retrieval...  ..., compilation, and hardware-aware optimization.Lead step-change initiatives across multiple...  ...fragmented systems, mentor senior engineers, and align technical and product leaders... 
    Principal
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    4 days ago
  • $200k - $250k

     ...multilayer stacking, and advanced cell architectures. We are seeking a Principal Engineer – Thin-Film & Process Engineering to serve as a technical leader for the development, optimization, and scale-up of our thin-film and cell manufacturing processes. This is... 
    Principal

    ENSURGE

    San Jose, CA
    19 days ago
  •  ...Job Overview We are seeking a Systems Packaging & Assembly Engineer to lead the development and integration of advanced packaging and...  ...requirements and system interfaces. Assembly Process Development Optimize assembly processes including die attach, flip-chip, wire... 
    Principal
    Contract work

    nEye.ai

    Santa Clara, CA
    25 days ago
  •  ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated...  ...movement.AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging... 

    AMD

    Santa Clara, CA
    2 days ago
  • $136.88k - $205k

     ...Your ImpactMarvell is seeking a highly motivated and talented Principal Test Engineer to join our dynamic NPI product engineering team. You will...  ...ATE format, bridging the gap between design and production.Optimize and Innovate: Play a lead role in optimizing test flows,... 
    Principal
    Permanent employment
    Full time
    Internship
    Work from home

    Marvell

    Santa Clara, CA
    a month ago
  •  ...Position Overview We are seeking a Principal Kubernetes Control Plane Engineer to architect the foundational...  ...management and orchestration. Optimize etcd performance and ensure rock-solid...  ...they impact customer training or inference jobs. Qualifications ~... 
    Principal
    Remote job
    Full time
    Local area

    Bitdeer Technologies Group

    San Jose, CA
    19 days ago
  • $160k - $240k

     ...quality, and reliability issues. Onto Innovation strives to optimize customers' critical path of progress by making them...  ...Summary & Responsibilities We are seeking a highly experienced Principal Systems Engineer to lead the architecture, integration, and validation of... 
    Principal
    Permanent employment

    Onto

    Milpitas, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal GenAI Inference Optimization Engineer. Be the first to apply!