Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior GPU Inference Performance Engineer

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, and AI serving frameworks, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.THE PERSON:A hands-on performance engineer who is equally comfortable reading a GPU trace and briefing executives. You are curious, evidence-driven, rigorous and you don't stop at "X is faster," you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, and AI serving framework teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, systems, and AI.KEY RESPONSIBILITIES:Full-stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and PCIe/Infinity Fabric data movement.AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging serving runtimes. Understand KV-cache management, continuous batching, PagedAttention, speculative decoding, and quantization (FP8, MXFP4, INT4) effects on throughput and latency.Competitive performance analysis: Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why performance differs — attributing gaps to specific hardware features (HBM bandwidth, compute density, interconnect topology), software maturity (kernel libraries, operator fusion, graph compilation), or configuration differences.Multi-server inference networking: Profile and optimize distributed inference topologies including prefill-decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, and NIC-level counters (Pensando, ConnectX). Quantify the impact of network latency, bandwidth, and congestion on end-to-end inference SLAs.GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold-start, and resource contention issues in production serving environments.Tooling and automation: Build reproducible benchmarking harnesses, profiling scripts, and performance regression dashboards. Automate trace collection and analysis to support continuous performance validation across driver, firmware, and framework updates.PREFERRED EXPERIENCE:Background in GPU performance engineering, HPC, or systems performance analysisHands-on proficiency with either AMD (ROCm, rocProfiler, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep understanding of GPU architecture: warp/wavefront execution, memory hierarchy (registers → LDS/shared → L2 → HBM), occupancy, and instruction-level parallelismExperience profiling vLLM, SGLang, or equivalent LLM serving frameworks, including quantization workflows (FP8, MXFP4, INT4, AWQ, GPTQ) and their performance implicationsExperience with multi-GPU and multi-node inference — tensor parallelism, pipeline parallelism, or PD disaggregation over RDMA/RoCE — including RCCL/NCCL profiling and network tools (perftest, ib_write_bw, tcpdump, Memory Fabric counters)Demonstrated ability to explain performance differences in written reports or presentations — not just "X is faster" but why, rooted in hardware and software evidenceStrong Python and C/C++ skills; comfort reading GPU kernel code (HIP/CUDA)Experience with Kubernetes GPU scheduling, MIG, and GPU operator performance, or contributions to open-source inference or profiling projectsACADEMIC CREDENTIALS:Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desiredThis role is not eligible for visa sponsorship.#LI-TB1#LI-HybridBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior GPU Inference Performance Engineer in Santa Clara, CA vacancy
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $184k - $287.5k

    We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers...  ...who are mindful of performance analysis and optimization...  ...hardware/software stack from GPU architecture to Deep...  ...language and multimodal model inference as part of NVIDIA... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This...  ...-speed inference.About The RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an expert on state... 
    Senior
    Performance
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $160k - $253k

     ...accelerated computing is the engine of artificial...  ...platforms integrate high performance compute, networking, and...  .... We are looking for a Senior Technical Marketing Engineer...  ...in showcasing NVIDIA's GPU architecture, server-...  ...and efficiency for AI inference & training.What you’ll... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with...  ...efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...advance your career. THE ROLE: Drive the performance of post‑training workloads on AMD...  ...candidate is passionate about software engineering and the craft of training performance. You...  ...model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication patterns... 
    Senior
    Performance

    AMD

    San Jose, CA
    1 day ago
  • $184k - $287.5k

    We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team. This...  ...and implementation of a real‑time, GPU‑accelerated propagation engine that predicts...  ...to see:PhD in computer graphics, high‑performance computing, computational... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...can make a lasting impact on the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining how developers write high-performance GPU software by bringing the safety, expressiveness, and modern... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195.2k - $361.2k

     ...hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H10...  ...Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to... 
    Senior
    Performance
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    3 days ago
  • $139k - $208.4k

     ...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer GPU RTL Design Engineer, you will contribute to the design and... 
    Senior
    Performance
    Hourly pay
    Full time
    Worldwide
    Relocation

    Samsung Semiconductor

    San Jose, CA
    5 hours ago
  • $124k - $208.4k

     ...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Architect, you will work on the analysis, verification, and... 
    Senior
    Performance
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    4 days ago
  • $136k - $218.5k

     ...seeking best-in-class ASIC Verification Engineers to verify the world’s leading GPUs. In this...  ...and system interface hardware of the GPU. This position will be working across many...  ...the challenge of creating the highest performance chips in the industry? If so, we want to... 
    Senior
    Performance
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

    We are now looking for a Senior Formal Verification Engineer for GPU Kernels! Modern AI performance relies on highly optimized GPU kernels — performance-critical code where bugs can be hard to catch and expensive to miss. NVIDIA's Deep Learning Safety Team is hiring engineers... 
    Senior
    Performance
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $124k - $208.4k

     ...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Modeling Architect, you will work on the design and development... 
    Senior
    Performance
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    4 days ago
  • $168k - $264.5k

    NVIDIA is seeking outstanding Senior Design Verification Engineers with a specialty in tools and automation to...  ...challenge of crafting the highest performance & lowest power silicon possible? If...  ...want to hear from you. Come, join our GPU ASIC team and help build the real-time... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $124k - $208.4k

     ...for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior GPU Design Verification Engineer - GCDV, you will contribute to the verification of top... 
    Senior
    Performance
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    3 days ago
  • $184k - $287.5k

    NVIDIA's invention of the GPU 1999 sparked the growth of the PC gaming market...  ...company”.We're looking for a Senior Performance Compiler Engineer to join our team and work on the open...  ...applications, accelerating both training and inference. You will be immersed in a diverse,... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking an NCX Senior Engineer to join our DSX team, collaborating...  ...customers realize efficient performance from NVIDIA's AI platform...  ...distributed training, inference optimization, and MLOps pipelines...  ...Kubernetes, containers, and GPU scheduling systems aligned to... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...night and find its way. NVIDIA's GPU runs Deep Learning algorithms...  ...the best Machine Learning Engineers with a background in computer...  ..., model metrics, continuous performance instrumentation, and reporting...  ...optimization for real-time inference on embedded or automotive platforms... 
    Senior
    Performance
    Full time
    Worldwide
    Night shift

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...computing. An era in which our GPU acts as the brains of...  ...for a Deep Learning Compiler Engineer. NVIDIA is hiring software engineers...  ...been the backbone of NVIDIA inference engine, spanning across data...  ...must deliver leading inference performance, fast build time, reduced... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...transforming every industry. GPU-accelerated deep learning...  ...are seeking an exceptional Senior Perception Engineer to help design and productize...  ...to quantify perception performance; analyze large-scale real and...  ...and optimizing training or inference pipelines through custom CUDA... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...transforming every industry. GPU-accelerated deep learning...  ...seeking an exceptional Senior Radar Perception Engineer to help design and productize...  ...quantify radar perception performance; analyze large-scale real...  ...optimizing training or inference pipelines through custom CUDA... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $138k - $206k

     ...consumption and maximizing performance. To achieve this goal,...  ...hardware and software engineers to identify and...  ...workloads.We are seeking a Senior LLM Systems...  ...reasoning, disaggregated inference, and Mixture-of-Experts...  ...understanding of NVIDIA GPU architecture and performance... 
    Senior
    Performance
    Work experience placement
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    2 days ago
  • $184k - $287.5k

     ...millions globally. We seek a Senior Engineer to lead technical efforts in...  ...By combining powerful local inference (Nemotron models) with...  ...(Ollama, Llama.cpp, vLLM), GPU-accelerated computing (CUDA,...  ...languages, particularly C++ (for performance-critical systems/OS integration... 
    Senior
    Performance
    Full time
    Local area
    Shift work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual...  ...for a motivated Deep Learning engineer to bring advanced communication technologies...  ...on scales up to 100K GPUs to inference down at microsecond latency.... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

     ...computing. An era in which our GPU acts as the brains of...  ...world.As a Developer Technology Engineer, you will be at the forefront...  ...resulting in suboptimal runtime performance.Conduct hands-on trainings, develop...  ...with GPU-accelerated AI inference driven by NVIDIA APIs and... 
    Senior
    Performance
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...computing. An era in which our GPU acts as the brains of computers...  ...at the forefront of AI and high-performance computing. As a Senior / Principal Deep Learning Engineer — Model Evaluation & AI Systems...  ...Work alongside model training, inference, and product divisions to... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $175k - $275k

     ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference...  ...We are looking for a hands-on Senior Quality Engineer to drive manufacturing quality...  ...inspection data, and parametric performance to proactively identify and correct... 
    Senior
    Performance
    Contract work

    Cerebras Systems

    Sunnyvale, CA
    4 days ago
  • $240k

     ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud...  ...multiple openings for Senior Quality Assurance...  ...integration, regression, and performance test strategies for AI/...  ...monitoring tools, and engineering best practices.Document... 
    Senior
    Performance

    Cerebras Systems

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior GPU Inference Performance Engineer. Be the first to apply!