Principal GenAI Inference Optimization Engineer
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLEWe are seeking a Principal GenAI Inference Optimization Engineer to join our Models and Applications team. This role focuses on improving performance, efficiency, and scalability of generative AI inference workloads on AMD GPU platforms. You will contribute to optimizing latency, throughput, and cost efficiency for real-world deployment of large-scale models, working across the software-hardware stack.THE PERSONThe ideal candidate is a strong technical contributor with expertise in GenAI inference optimization, GPU performance, and large-scale serving systems. You have a solid understanding of GPU architecture, memory systems, and communication patterns, and can apply this knowledge to improve inference efficiency.You are comfortable working across multiple layers—from kernels and runtimes to frameworks and serving systems—and can independently drive optimization efforts while collaborating with cross-functional teams.KEY RESPONSIBILITIES- Optimize performance of GenAI inference workloads on AMD GPU platforms across single-node and distributed environments.- Improve latency, throughput, and cost efficiency for LLM and multimodal model serving in production.- Analyze and resolve bottlenecks across compute, memory, and communication (e.g., kernel efficiency, KV-cache usage, memory bandwidth, scheduling).- Contribute to cross-stack optimizations spanning kernels, runtimes, communication libraries, and inference/serving frameworks (e.g., vLLM, SGLang, Triton, or similar systems).- Implement and evaluate inference optimization techniques such as batching strategies, quantization, prefix caching, and speculative decoding.- Support development and optimization of scalable serving systems, including request scheduling and resource utilization.- Develop and use profiling, benchmarking, and performance analysis tools for inference workloads.- Collaborate with hardware, compiler, and framework teams to improve overall system performance.- Contribute to internal tools and, where applicable, open-source projects for inference optimization on AMD platforms.- Document best practices and contribute to performance guidelines for GenAI deployment.PREFERRED EXPERIENCE- Strong understanding of GPU architecture and performance fundamentals (compute, memory hierarchy, interconnects such as PCIe/Infinity Fabric/RDMA).- Experience with GenAI inference optimization techniques (e.g., quantization, KV-cache optimization, batching).- Hands-on experience with inference/serving frameworks such as vLLM, SGLang, Triton, TensorRT-LLM, or similar.- Experience working on LLM or multimodal inference workloads.- Familiarity with distributed systems and serving architectures.- Experience with ML frameworks (PyTorch, JAX, or TensorFlow), especially for inference.- Proficiency in Python and at least one systems language (C++/CUDA/HIP).- Experience with profiling, debugging, and performance tuning tools.- Ability to work collaboratively across teams and deliver impactful optimizations.ACADEMIC CREDENTIALS- B.S., M.S. or Ph.D. in Computer Science, Computer Engineering, or a related field preferred, or equivalent industry experience.LOCATION- San Jose, CA#LI-MV1#HYBRIDThis role is not eligible for visa sponsorship.Benefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
- ...and create.We’re seeking a Principal Machine Learning Systems Engineer (P60) to lead technical directions of GenAI Products & Knowledge Innovations... ..., scalable services.Optimize latency, throughput, and resource... ...-scale model training, inference pipelines, or search/retrieval...PrincipalWork at officeLocal area
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SuggestedFull time$124.5k - $272k
...Fortune 100, trust the Netskope One platform, its Zero Trust Engine, and the powerful NewEdge network to gain full visibility... ...roleAs a Senior Staff Machine Learning Scientist, you own the inference and optimization layer that makes AI in agentic workflows fast, efficient,...Principal$150k - $190k
...Principal Test Engineer \\ \ \ \ San Jose, CA | Hybrid (3 days onsite \/ 2 remote)\\ \ Base Salary: $150,000â$190,000 (DOE)\\ \... ...ramps at OSAT partners (domestic and offshore)\ . \\ Optimize yield, test time, and multisite efficiency \ to improve cost...PrincipalFull timeWork experience placementRemote work$272k - $431.25k
...: models, adaptation workflows, and inference pipelines built for real inspection... ...everyday constraints.We’re seeking a Principal Software Engineer for Systems Inspection in Santa Clara... ..., model compression, and deployment optimization. You’ll join a small, high-impact...PrincipalFull timeShift work$206.4k - $379.1k
...AI Services team is seeking a Principal Service Engineer to serve as the technical lead for our GenAI Services domain. In this high... ....Design and architect inference infrastructure for enterprise... ...(training, inference, and/or optimization).Proven track record of leading...PrincipalFull timeTemporary workLocal areaWorldwide- ...Santa Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be... ...and understand the nuances of what it takes to optimize and trade-off various aspects of hardware-software co...Principal3 days per week
- ...NVIDIA Corporation in Santa Clara is seeking a Principal Software Engineer for Systems Inspection to advance AI products for semiconductor analysis... ...vision, multimodal AI, anomaly detection, and deployment optimization. Lead architecture, collaborate across teams, and...Principal
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks,...Full time- ...Principal Packaging EngineerWe are looking for a Principal Packaging Engineer to own package architecture and development for Velaura's next-generation SoCs. This role... ...technologies, and cross-functional system optimization.ResponsibilitiesOwn package architecture from...PrincipalFlexible hours
- ...is seeking a full-time Deep Learning Algorithm Engineer passionate about pioneering DL, foundation models, and GenAI for image processing in the semiconductor... ...experience with DL model development, training, optimization, and deployment, with emphasis on performance...Full time
- ...performance computing, AI, and semiconductor products. As a Principal Mechanical Engineer, you will lead the development of advanced package... ...thermal, reliability, manufacturing, and supplier teams to optimize product performance and manufacturability.Conduct and lead...Principal
$205k - $260k
...Responsibilities Architect and optimize cloud-native CI/CD pipelines and shared automation... ..., and timing-closure workflows for engineering throughput and cloud cost efficiency.... ...and benefits. The role offers a broad principal-level mandate spanning cloud-native CAD...PrincipalFull time- ...Role Overview We are looking for an experienced Principal Systems Integration Engineer to join the Arycs’ GTM team. In this role, you will serve... ...and execute comprehensive customer test plans to ensure optimal product performance. Qualifications ~7+ years of...PrincipalFlexible hours
- ...and integration limits. Our systems are engineered for deployment in real networks — not... ...optical technologies. About the Role As a Principal Hardware Engineer, Signal Processing,... ...for developing DSP algorithmic models, optimizing signal chains, validating performance...PrincipalWork at office
- ...Description Renesas is seeking a Principal Test Architect for our Analog and Mixed... ..., test infrastructure, and process optimization across Renesas' analog and mixed-signal... ...closely with design, product, and test engineering teams to drive test innovation from new...PrincipalFull timeFlexible hours
- ...Adobe Applied Science & Machine Learning (ASML) seeks a Principal Scientist, ML – Overall Architect to own end-to-end architecture across training, inference, model building and data pipelines. You will serve as the principal technical owner bridging infrastructure, modeling...Principal
- ...Role Overview We are seeking a Wafer Bonding Integration Engineer to lead the development, integration, and manufacturing transfer... ...achieved. Wafer Bonding Process Integration ~ Develop and optimize advanced wafer bonding processes including: ~ Fusion Bonding...Principal
$148.32k - $203.94k
...For more information, visit: . Job Summary The Principal Test Engineer is responsible for the development of production test solutions... ...collaborate with cross-functional teams to develop highly optimized test solutions enabling SiTime’s game-changing timing...PrincipalWork experience placement$217k - $359k
...and quality of enterprise-class storage hardware systems that power mission-critical data environments. You will bridge complex engineering concepts with real-world deployment by leading cross-functional alignment across internal engineering, Joint Development Manufacturing...PrincipalWork at officeFlexible hours$182k - $319k
...and the top 3 non-X86 server providers engineering solutions for this generation and the next... ...the future together. The Senior Principal, Design Engineering will be responsible... .... ~ Well understand power efficiency optimization. ~ Experience of products development...PrincipalLocal area$248k - $396.75k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline focused on designing... ...of production environments.As a Principal SRE, you will shape the technical direction... ...AI/ML platforms, GPU infrastructure, inference systems, training environments, or high...PrincipalFull time$184k - $287.5k
...now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...are mindful of performance analysis and optimization to help us squeeze every last clock... ...Implement language and multimodal model inference as part of NVIDIA Inference Microservices...Full time- ...index serving, query processing, and ML inference. Turn advances in information retrieval... ..., compilation, and hardware-aware optimization.Lead step-change initiatives across multiple... ...fragmented systems, mentor senior engineers, and align technical and product leaders...PrincipalWork at officeLocal area
$200k - $250k
...multilayer stacking, and advanced cell architectures. We are seeking a Principal Engineer – Thin-Film & Process Engineering to serve as a technical leader for the development, optimization, and scale-up of our thin-film and cell manufacturing processes. This is...Principal- ...Job Overview We are seeking a Systems Packaging & Assembly Engineer to lead the development and integration of advanced packaging and... ...requirements and system interfaces. Assembly Process Development Optimize assembly processes including die attach, flip-chip, wire...PrincipalContract work
- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...movement.AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging...
$136.88k - $205k
...Your ImpactMarvell is seeking a highly motivated and talented Principal Test Engineer to join our dynamic NPI product engineering team. You will... ...ATE format, bridging the gap between design and production.Optimize and Innovate: Play a lead role in optimizing test flows,...PrincipalPermanent employmentFull timeInternshipWork from home- ...Position Overview We are seeking a Principal Kubernetes Control Plane Engineer to architect the foundational... ...management and orchestration. Optimize etcd performance and ensure rock-solid... ...they impact customer training or inference jobs. Qualifications ~...PrincipalRemote jobFull timeLocal area
$160k - $240k
...quality, and reliability issues. Onto Innovation strives to optimize customers' critical path of progress by making them... ...Summary & Responsibilities We are seeking a highly experienced Principal Systems Engineer to lead the architecture, integration, and validation of...PrincipalPermanent employment
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal GenAI Inference Optimization Engineer. Be the first to apply!
- director data engineering San Jose, CA
- principal engineer San Jose, CA
- senior chief engineer San Jose, CA
- principal infrastructure engineer San Jose, CA
- principal data engineer San Jose, CA
- senior civil engineer project manager San Jose, CA
- chief engineer San Jose, CA
- director software engineering San Jose, CA
- principal network engineer San Jose, CA
- director systems engineering San Jose, CA





