Senior GPU Inference Performance Engineer
Advanced Micro Devices Inc
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack from GPU silicon, communication libraries, networking fabrics, and operating systems through the software runtime and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, networking, and AI infrastructure, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.THE PERSON:A hands-on performance engineer who is equally comfortable reading a GPU trace, debugging distributed systems performance issues, and briefing executives. You are curious, evidence-driven, rigorous, and you don't stop at "X is faster" and you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, networking, and AI infrastructure teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, operating systems, communication libraries, networking, and AI. Experience with Linux systems, distributed GPU infrastructure, RDMA/RoCE networking, or communication libraries such as NCCL/RCCL is highly valued.KEY RESPONSIBILITIES:Full-stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data movement, and GPU runtime behavior.Systems and runtime performance analysis: Profile and diagnose performance interactions between GPU runtimes, Linux operating systems, device drivers, container runtimes, memory subsystems, CPU scheduling, NUMA topology, and I/O pathways. Identify system-level bottlenecks that impact throughput, latency, and GPU utilization.Competitive performance analysis: Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized AI and LLM workloads. Produce clear, data-backed explanations of why performance differs attributing gaps to hardware architecture, networking topology, communication libraries, software maturity, runtime behavior, or configuration differences.Multi-server inference networking: Profile and optimize distributed inference topologies including prefill-decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, GPUDirect RDMA, NIC-level counters (Pensando, ConnectX), and network performance tools. Quantify the impact of latency, bandwidth, congestion, and topology on end-to-end inference SLAs.GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold-start, resource contention, and infrastructure inefficiencies in production environments.Tooling and automation: Build reproducible benchmarking harnesses, profiling scripts, and performance regression dashboards. Automate trace collection and analysis to support continuous performance validation across firmware, drivers, networking stacks, runtimes, and AI frameworks.PREFERRED EXPERIENCE:Background in GPU performance engineering, HPC, distributed systems, networking, operating systems, or systems performance analysis.Hands-on proficiency with either AMD (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep understanding of GPU architecture: warp/wavefront execution, memory hierarchy, occupancy, and instruction-level parallelism.Experience analyzing GPU communication and networking performance including NCCL/RCCL, RDMA/RoCE, GPUDirect RDMA, UCX, MPI, ConnectX, Pensando, or similar high-performance networking technologies.Experience with multi-GPU and multi-node inference, training, or HPC environments including tensor parallelism, pipeline parallelism, distributed communication libraries, and network performance analysis tools.Experience with Linux systems performance analysis, operating systems, device drivers, virtualization, container runtimes, or low-level systems software development.Demonstrated ability to explain performance differences in written reports or presentations—not just "X is faster" but why, rooted in hardware and software evidence.Strong Python and C/C++ skills; comfort reading GPU kernel code (HIP/CUDA), runtime code, or systems-level software.Experience with Kubernetes GPU scheduling, MIG, GPU operator performance, or contributions to open-source infrastructure, systems, networking, inference, or profiling projects.ACADEMIC CREDENTIALS:Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desired.This role is not eligible for visa sponsorship.#LI-CB1#LI-HybridBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SeniorPerformanceFull time- ...NVIDIA Corporation in Santa Clara, CA, seeks a Sr. Inference Engineer to accelerate LLM inference through GPU kernel optimization. You will lead kernel benchmarking, model-level performance analysis, and AI-driven optimization workflows across silicon and software stacks...SeniorPerformance
- ...industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference... ....About The RoleWe’re hiring a Senior Frontend Engineer to own and scale critical... ...real-time vs batch processing, performance, and cost.Drive projects from...SeniorPerformanceShift work
$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational... ...latency trade-offs• Develop and tune high-performance kernels for critical operations where... ...session load• Implement multi-GPU inference (tensor parallelism, collective...SeniorPerformanceInternshipLocal areaFlexible hours$160k - $253k
...accelerated computing is the engine of artificial... ...platforms integrate high performance compute, networking, and... .... We are looking for a Senior Technical Marketing Engineer... ...in showcasing NVIDIA's GPU architecture, server-... ...and efficiency for AI inference & training.What you’ll...SeniorPerformanceFull time- ...Position Overview We are seeking a Senior GPU Systems & Fabric Engineer to serve as the critical bridge... ...responsible for creating the high-performance 'hardware foundation' that makes AI... ...) to enable efficient multi-tenant inference workloads and maximize cluster utilization...SeniorPerformanceFull timeLocal area
- ...advance your career. THE ROLE: Drive the performance of post‑training workloads on AMD... ...candidate is passionate about software engineering and the craft of training performance. You... ...model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication patterns...SeniorPerformance
- ...NVIDIA Corporation’s Local AI team is building software that optimizes LLM inference on edge AI hardware. You’ll evaluate open-source frameworks, map models to GPU architecture, and drive performance characterization across multi-node configurations. You will own...SeniorPerformanceLocal area
$184k - $287.5k
We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team. This... ...and implementation of a real‑time, GPU‑accelerated propagation engine that predicts... ...to see:PhD in computer graphics, high‑performance computing, computational...SeniorPerformanceFull time$152k - $241.5k
...can make a lasting impact on the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining how developers write high-performance GPU software by bringing the safety, expressiveness, and modern...SeniorPerformanceFull timeRemote work- ...NVIDIA Corporation is seeking a Senior System Software Engineer to advance Dynamo, its GPU-accelerated deep learning inference platform. You will develop open-source software for AI model inference, expand disaggregated serving, and optimize latency and throughput across...Senior
$139k - $208.4k
...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom... ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer GPU RTL Design Engineer, you will contribute to the design and...SeniorPerformanceHourly payFull timeWorldwideRelocation$139k - $208.4k
...for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other... ...world. Come build with us!Role and ResponsibilitiesAs a Senior GPU Front-End Integration Engineer, you will contribute to top-level RTL integration of complex...SeniorPerformanceHourly payFull timeWorldwideRelocation- ...NVIDIA in Santa Clara is seeking software engineers to design and implement libraries and tools to accelerate... ...and global teams to define features, optimize performance, and deliver production-ready code for scalable GPU-accelerated software. #J-18808-Ljbffr Jobleads...SeniorPerformance
$136k - $218.5k
...seeking best-in-class ASIC Verification Engineers to verify the world’s leading GPUs. In this... ...and system interface hardware of the GPU. This position will be working across many... ...the challenge of creating the highest performance chips in the industry? If so, we want to...SeniorPerformanceFull timeShift work- ...NVIDIA is seeking a Senior ASIC Power Engineer to design power-efficient hardware accelerators for mobile, embedded, and datacenter platforms. You will develop innovative HW, GPU and system designs to push performance and efficiency while driving power reductions across...SeniorPerformance
$124k - $208.4k
...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom... ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Architect, you will work on the analysis, verification, and...SeniorPerformanceHourly payFull timeRelocation$124k - $208.4k
...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom... ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Modeling Architect, you will work on the design and development...SeniorPerformanceHourly payFull timeRelocation$224k - $356.5k
...systems, and software to solve some of the world’s most exciting computing problems. We are looking for a Senior System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team. You will be responsible for architecting production software for...SeniorPerformanceFull timeRemote work$184k - $287.5k
...and software to solve some of the world’s most exciting computing problems. We are looking for a System Software Engineer to join NVIDIA’s GPU Performance and Power Management Software team. You will help design, develop, and debug production software for performance states...SeniorPerformanceFull time$140k - $224.25k
...seeking a creative, and hands-on software engineer with a test to failure approach who is a... ...and optimize the testing workflows in GPU domain.Write maintainable, reliable, and... ...peer code reviews including feedback on performance, scalability, and correctnessOptimally estimate...SeniorPerformanceFull time- ...Samsung Electronics America, based in the United States, is seeking a Senior Engineer GPU Graphics Core Performance Verification Engineer to lead pre-silicon verification for the next-generation GPU and system IP, collaborating with architecture, RTL, modeling and software...SeniorPerformance
- ...NVIDIA Corporation in Santa Clara, CA is seeking a software engineer with a strong background in parallel processing and GPU architecture to push the performance envelope at the intersection of AI, high‑performance computing, and financial systems. You will design and...SeniorPerformance
- ...NVIDIA Corporation is seeking a Senior GPU Platforms Engineer to drive GPU/SOC software across platforms, focusing on kernel driver development, firmware interactions, health monitoring and performance optimization. You will build high‑performance observability systems...SeniorPerformance
$124k - $208.4k
...for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other... ...world. Come build with us!Role and ResponsibilitiesAs a Senior GPU Design Verification Engineer - GCDV, you will contribute to the verification of top...SeniorPerformanceHourly payFull timeRelocation- ...implement, and maintain production GPU system software across kernel... ..., tracing, profiling, and performance measurement. Own... ..., validation, and production engineering teams. Improve resilience... ...with distributed training or inference and topology-aware communication...SeniorPerformanceFull time
$184k - $287.5k
NVIDIA's invention of the GPU 1999 sparked the growth of the PC gaming market... ...company”.We're looking for a Senior Performance Compiler Engineer to join our team and work on the open... ...applications, accelerating both training and inference. You will be immersed in a diverse,...SeniorPerformanceFull timeRemote work$152k - $241.5k
NVIDIA's invention of the GPU 1999 sparked the growth of the PC... ...for versatile software engineers for our XLA team. NVIDIA is at... .... Come join us to build high-performance, production-grade software that... ...workloads. You will optimize inference and training performance for...SeniorPerformanceFull timeRemote work$184k - $287.5k
NVIDIA is seeking an NCX Senior Engineer to join our DSX team, collaborating... ...customers realize efficient performance from NVIDIA's AI platform... ...distributed training, inference optimization, and MLOps pipelines... ...Kubernetes, containers, and GPU scheduling systems aligned to...SeniorPerformanceFull timeRemote work$184k - $287.5k
...transforming every industry. GPU-accelerated deep learning... ...are seeking an exceptional Senior Perception Engineer to help design and productize... ...to quantify perception performance; analyze large-scale real and... ...and optimizing training or inference pipelines through custom CUDA...SeniorPerformanceFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior GPU Inference Performance Engineer. Be the first to apply!
- senior service designer Santa Clara, CA
- senior mulesoft developer Santa Clara, CA
- senior business manager Santa Clara, CA
- senior linux systems engineer Santa Clara, CA
- senior mainframe developer Santa Clara, CA
- senior cloud security engineer Santa Clara, CA
- senior level Santa Clara, CA
- senior project manager Santa Clara, CA
- senior network engineer remote Santa Clara, CA
- senior manager product development Santa Clara, CA


