Principal GenAI Inference Optimization Engineer
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative AI inference workloads on AMD GPU platforms. You will contribute to optimizing latency, throughput, and cost efficiency for real-world deployment of large-scale models, working across the software-hardware stack.THE PERSONThe ideal candidate is a strong technical contributor with expertise in GenAI inference optimization, GPU performance, and large-scale serving systems. You have a solid understanding of GPU architecture, memory systems, and communication patterns, and can apply this knowledge to improve inference efficiency.You are comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently drive optimization efforts while collaborating with cross-functional teams.KEY RESPONSIBILITIES- Optimize performance of GenAI inference workloads on AMD GPU platforms across single-node and distributed environments.- Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.- Analyze and resolve bottlenecks across compute, memory, and communication (e.g., kernel efficiency, KV-cache usage, memory bandwidth, scheduling).- Contribute to cross-stack optimizations spanning kernels, runtimes, communication libraries, and inference/serving frameworks (e.g., vLLM, SGLang, Triton, or similar systems).- Implement and evaluate inference optimization techniques such as batching strategies, quantization, prefix caching, and speculative decoding.- Support development and optimization of scalable serving systems, including request scheduling and resource utilization.- Develop and use profiling, benchmarking, and performance analysis tools for inference workloads.- Collaborate with hardware, compiler, and framework teams to improve overall system performance.- Contribute to internal tools and, where applicable, open-source projects for inference optimization on AMD platforms.- Document best practices and contribute to performance guidelines for GenAI deployment.PREFERRED EXPERIENCE- Strong understanding of GPU architecture and performance fundamentals (compute, memory hierarchy, interconnects such as PCIe/Infinity Fabric/RDMA).- Experience with GenAI inference optimization techniques (e.g., quantization, KV-cache optimization, batching).- Hands-on experience with inference/serving frameworks such as vLLM, SGLang, Triton, TensorRT-LLM, or similar.- Experience working on LLM or multimodal inference workloads.- Familiarity with distributed systems and serving architectures.- Experience with ML frameworks (PyTorch, JAX, or TensorFlow), especially for inference.- Proficiency in Python and at least one systems language (C++/CUDA/HIP).- Experience with profiling, debugging, and performance tuning tools.- Ability to work collaboratively across teams and deliver impactful optimizations.ACADEMIC CREDENTIALS- B.S., M.S. or Ph.D. in Computer Science, Computer Engineering, or a related field preferred, or equivalent industry experience.LOCATION- San Jose, CA#LI-MV1#HYBRIDThis role is not eligible for visa sponsorship.Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time$195.2k - $361.2k
...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment...SuggestedFull timeInternshipLocal areaImmediate startShift work$182.5k - $260.5k
...Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility... ...roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient,...Principal$206.4k - $379.1k
...AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain. In this high... ....Design and architect inference infrastructure for enterprise... ...(training, inference, and/or optimization).Proven track record of leading...PrincipalFull timeTemporary workLocal areaWorldwide- ...specialized Sr. Staff or Principal level engineer who is passionate about enabling... ...on scaling training and inference for the latest Generative... ...training and inference optimizations across a variety of applications... ...with state-of-the-art GenAI algorithms and software...Suggested
$184k - $287.5k
...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...are mindful of performance analysis and optimization to help us squeeze every last clock... ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices...Full time- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...movement.AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging...
- ...Jose, California, United StatesProducts - Engineering /Fulltime /HybridWe are seeking an... ...user-friendly end-to-end solutions for GenAI applications, combining cutting-edge backend... ...and accessible interfaces. Develop and optimize backend services and APIs (REST, GraphQL...PrincipalFull time
$150k - $250k
...a significant plus.Need to work closely with system and test engineers to develop high speed interface, package/board, and system clocks... ...Layout design and support. Need to get involved into layout optimizations for high speed or high precision performance directly.•...Principal$152k - $241.5k
...NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...generation and computational graph optimizations for next-generation NVIDIA GPUs.Advance... ...compilation problems for AI workloads (both inference and training) and successfully...Full time$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks,...Full time- ...Santa Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be... ...and understand the nuances of what it takes to optimize and trade-off various aspects of hardware-software co...Principal3 days per week
- ....THE PERSON: The ideal candidate is passionate about software engineering and the craft of training performance. You lead sophisticated... ...throughput, memory efficiency, and stability across data, model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication...Principal
$200k
Role Overview We are looking for a Principal Packaging Engineer to own package architecture and development for Velaura’s next-generation SoCs... ...substrate and assembly technologies, and cross-functional system optimization. \n Responsibilities ● Own package architecture from...PrincipalFull timeFlexible hours$152k - $241.5k
...NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other... ...requirementsContribute to performance optimization and benchmarking efforts for specialized automotive...Full time$210.16k - $271.98k
Senior Principal Systems Development EngineerHelp architect and deliver Dell's L11 rack-scale... .... As a Principal Systems Development Engineer, you'll own the system-level architecture... ..., and applies that insight to optimize complete AI rack solutions.Directs the application...Principal- ...Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy...
$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full... ...frameworks, and performance engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs,...PrincipalFull time$182k - $319k
...City: San Jose General OverviewFunctional Area: Engineering (ENG)Career Stream: Engineering (ENG)Role: Senior Principal (SPR)Job Title: Senior Principal, Design... ...VR13 design.Well understand power efficiency optimization.Experience of products development from prototype...PrincipalLocal area$150k - $190k
...Principal Test Engineer \\ \ \ \ San Jose, CA | Hybrid (3 days onsite \/ 2 remote)\\ \ Base Salary: $150,000â$190,000 (DOE)\\ \... ...ramps at OSAT partners (domestic and offshore)\ . \\ Optimize yield, test time, and multisite efficiency \ to improve cost...PrincipalFull timeWork experience placementRemote work- ...Integrations is seeking an ambitious, highly motivated Principal AI/Data Center Systems Engineer to help define and develop next-generation high-... ...technical challenges, evaluate design tradeoffs, and optimize system performance, density, and reliability. Serve as...PrincipalWorldwideShift work
$160.2k - $240k
...thrive, learn, and lead. Your Team, Your ImpactAs the Senior Principal test engineer within the Operations, You’lldefines the IO Chiplet ATE... ...testability and test strategies, contributing to improved yield, optimized test methodologies, and cross‑functional innovation.Convert...PrincipalPermanent employmentFull timeInternshipWork from home$182k - $260k
...shape the future of cybersecurity.RoleWe are looking for a Principal GenAI Data Engineer to join our team. This is a Hybrid role based in San Jose... ..., securing, or positioning AI-driven solutions to optimize outcomes within your functional domainExpert-level Python...PrincipalFull timeWork at officeLocal areaRemote work$184k - $287.5k
...best work. Come join the team and see how you can make a lasting impact on the worldWe are looking for an experienced Compiler Optimization Engineer for an exciting role in our Compute Compiler Team. We deliver features and improvements to CUDA and other compute compilers...Full time- ...deliver industry-leading training and inference speeds; over 10 times faster than GPU-based... ...RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an... ...SGLang, TensorRT-LLM), GPU kernel-level optimization toolchains (CUDA, Triton), and an...Contract workShift work
- ...THE ROLE: We are looking for a dynamic, energetic Lead / Principal Systems Design Engineer to join our growing team. As a key contributor to the... ...test execution to make sure all features are validated and optimized on time PREFERRED EXPERIENCE: Programming/scripting...Principal
$158.6k - $234.65k
...reliability and performance.Role Overview:We are seeking a Sr. Principal Engineer, Manufacturing Infrastructure to join our Advanced... ...internal manufacturing infrastructure remains at the cutting edge.Optimize solutions for manufacturing yield, repeatability, and throughput...PrincipalPermanent employmentFull timeInternshipWork from homeShift work$170k
...A leading chip and silicon IP provider is seeking a NPI Principal Product Engineer to join our Operations team in San Jose. In this role, you... ...support. Experience in Failure Analysis and RMA support. Optimize Yield, Cost, Cycle Time and Quality. Data Analysis proficiency...PrincipalRemote workRelocation package3 days per week$144k - $216k
..., quality, and reliability issues. Onto Innovation strives to optimize customers’ critical path of progress by making them smarter, faster... ...: Lead and mentor a multidisciplinary team of hardware engineers, fostering a culture of technical growth, innovation, and accountability...PrincipalPermanent employmentFull time- ...we are redefining the architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our hardware organization,... ...Processing Unit) to solve the industry's most massive LLM inference challenges.Required Qualifications• Education: BS/MS in...Principal3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal GenAI Inference Optimization Engineer. Be the first to apply!
- chief engineer San Jose, CA
- engineering director San Jose, CA
- project engineer assistant project manager San Jose, CA
- senior director engineering San Jose, CA
- director of product engineering San Jose, CA
- senior chief engineer San Jose, CA
- hotel chief engineer San Jose, CA
- principal developer San Jose, CA
- chief design engineer San Jose, CA
- principal engineer San Jose, CA


