Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior GPU Inference Performance Engineer

Advanced Micro Devices Inc

WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack from GPU silicon, communication libraries, networking fabrics, and operating systems through the software runtime and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, networking, and AI infrastructure, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.THE PERSON:A hands-on performance engineer who is equally comfortable reading a GPU trace, debugging distributed systems performance issues, and briefing executives. You are curious, evidence-driven, rigorous, and you don't stop at "X is faster" and you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, networking, and AI infrastructure teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, operating systems, communication libraries, networking, and AI. Experience with Linux systems, distributed GPU infrastructure, RDMA/RoCE networking, or communication libraries such as NCCL/RCCL is highly valued.KEY RESPONSIBILITIES:Full-stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data movement, and GPU runtime behavior.Systems and runtime performance analysis: Profile and diagnose performance interactions between GPU runtimes, Linux operating systems, device drivers, container runtimes, memory subsystems, CPU scheduling, NUMA topology, and I/O pathways. Identify system-level bottlenecks that impact throughput, latency, and GPU utilization.Competitive performance analysis: Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized AI and LLM workloads. Produce clear, data-backed explanations of why performance differs attributing gaps to hardware architecture, networking topology, communication libraries, software maturity, runtime behavior, or configuration differences.Multi-server inference networking: Profile and optimize distributed inference topologies including prefill-decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, GPUDirect RDMA, NIC-level counters (Pensando, ConnectX), and network performance tools. Quantify the impact of latency, bandwidth, congestion, and topology on end-to-end inference SLAs.GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold-start, resource contention, and infrastructure inefficiencies in production environments.Tooling and automation: Build reproducible benchmarking harnesses, profiling scripts, and performance regression dashboards. Automate trace collection and analysis to support continuous performance validation across firmware, drivers, networking stacks, runtimes, and AI frameworks.PREFERRED EXPERIENCE:Background in GPU performance engineering, HPC, distributed systems, networking, operating systems, or systems performance analysis.Hands-on proficiency with either AMD (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep understanding of GPU architecture: warp/wavefront execution, memory hierarchy, occupancy, and instruction-level parallelism.Experience analyzing GPU communication and networking performance including NCCL/RCCL, RDMA/RoCE, GPUDirect RDMA, UCX, MPI, ConnectX, Pensando, or similar high-performance networking technologies.Experience with multi-GPU and multi-node inference, training, or HPC environments including tensor parallelism, pipeline parallelism, distributed communication libraries, and network performance analysis tools.Experience with Linux systems performance analysis, operating systems, device drivers, virtualization, container runtimes, or low-level systems software development.Demonstrated ability to explain performance differences in written reports or presentations—not just "X is faster" but why, rooted in hardware and software evidence.Strong Python and C/C++ skills; comfort reading GPU kernel code (HIP/CUDA), runtime code, or systems-level software.Experience with Kubernetes GPU scheduling, MIG, GPU operator performance, or contributions to open-source infrastructure, systems, networking, inference, or profiling projects.ACADEMIC CREDENTIALS:Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desired.This role is not eligible for visa sponsorship.#LI-TB1#LI-HybridBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior GPU Inference Performance Engineer in Santa Clara, CA vacancy
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • CoreWeave is hiring a Senior Engineer for its Benchmarking & Performance team to write, profile, and optimize GPU kernels on the LLM inference path. You will improve latency and throughput and collaborate with product, orchestration, and hardware teams to achieve strict... 
    Senior
    Performance

    CoreWeave

    Sunnyvale, CA
    4 days ago
  • $184k - $287.5k

    We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers...  ...who are mindful of performance analysis and optimization...  ...hardware/software stack from GPU architecture to Deep...  ...language and multimodal model inference as part of NVIDIA... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This...  ...-speed inference.About The RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an expert on state... 
    Senior
    Performance
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    5 days ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational...  ...latency trade-offs• Develop and tune high-performance kernels for critical operations where...  ...session load• Implement multi-GPU inference (tensor parallelism, collective... 
    Senior
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    5 days ago
  • $184k - $356.5k

     ...leading technology company in California is seeking a Senior DL Algorithms Engineer to drive inference performance for Deep Learning workloads. The role involves...  ...have over 5 years of experience in deep learning and GPU programming. This position offers a competitive base... 
    Senior
    Performance

    NVIDIA Corporation

    Santa Clara, CA
    4 days ago
  • NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable inferencing software and contribute to high-performance GPU-accelerated deployments. Join a cross-disciplinary team to push... 
    Senior
    Performance

    NVIDIA Corporation

    Santa Clara, CA
    2 days ago
  • $160k - $253k

     ...accelerated computing is the engine of artificial...  ...platforms integrate high performance compute, networking, and...  .... We are looking for a Senior Technical Marketing Engineer...  ...in showcasing NVIDIA's GPU architecture, server-...  ...and efficiency for AI inference & training.What you’ll... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • $184k - $287.5k

     ...seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with...  ...efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    5 days ago
  • Advanced Micro Devices is seeking a Senior GPU Inference Performance Engineer to own end-to-end profiling of GPU-accelerated AI inference workloads. You will analyze workloads across AMD Instinct and NVIDIA GPUs, profile AI serving frameworks, and explain performance gaps... 
    Senior
    Performance

    Advanced Micro Devices

    Santa Clara, CA
    4 days ago
  •  ...advance your career. THE ROLE: Drive the performance of post‑training workloads on AMD...  ...candidate is passionate about software engineering and the craft of training performance. You...  ...model, and optimizer steps.Optimize multi‑GPU/multi‑node training and communication patterns... 
    Senior
    Performance

    AMD

    San Jose, CA
    2 days ago
  • $184k - $287.5k

    We are seeking a self‑motivated senior engineer for the Aerial Omniverse Digital Twin team. This...  ...and implementation of a real‑time, GPU‑accelerated propagation engine that predicts...  ...to see:PhD in computer graphics, high‑performance computing, computational... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...can make a lasting impact on the world.NVIDIA is hiring a Senior Compiler Engineer to join our team driving the next generation of GPU systems programming. We are redefining how developers write high-performance GPU software by bringing the safety, expressiveness, and modern... 
    Senior
    Performance
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $195.2k - $361.2k

     ...hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H10...  ...Benchmark across hardware tiers and publish honest performance comparisonsUpstream fixes and patches to... 
    Senior
    Performance
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    5 days ago
  • $139k - $208.4k

     ...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer GPU RTL Design Engineer, you will contribute to the design and... 
    Senior
    Performance
    Hourly pay
    Full time
    Worldwide
    Relocation

    Samsung Semiconductor

    San Jose, CA
    1 day ago
  •  ..., Mirantis empowers platform engineering teams to deliver composable,...  ...Mirantis delivers the automation, GPU orchestration, and policy-...  ..., model the real TCO, prove performance, and de-risk a customer's move...  ...stood up training and inference workloads, argued interconnect... 
    Senior
    Performance
    Full time

    Mirantis

    San Jose, CA
    8 days ago
  •  ...Technologies, Inc. in Santa Clara is seeking a GPU Engineer to architect, design, implement, verify, and optimize GPU cores for power and performance. You will collaborate with cross-...  ...development, and experience working with senior leadership on technical strategy. #J-18... 
    Senior
    Performance
    Remote job

    Qualcomm

    Santa Clara, CA
    4 days ago
  • Apple Inc. in Santa Clara, California seeks a senior architect for GPU micro-architecture within the Silicon Technologies group. You will help design and build high-performance, power-efficient processors and a next-generation GPU. The role involves specifying new micro... 
    Senior
    Performance

    Apple Inc.

    Santa Clara, CA
    4 days ago
  • $166k - $244k

    Google is looking for a Senior Software Engineer in Sunnyvale, CA to lead GPU performance optimizations for cutting-edge AI and machine learning technologies. This role offers the opportunity to work on innovative projects that impact billions of users around the globe.... 
    Senior
    Performance

    Google

    Sunnyvale, CA
    1 day ago
  • $136k - $218.5k

     ...seeking best-in-class ASIC Verification Engineers to verify the world’s leading GPUs. In this...  ...and system interface hardware of the GPU. This position will be working across many...  ...the challenge of creating the highest performance chips in the industry? If so, we want to... 
    Senior
    Performance
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    5 days ago
  • $124k - $208.4k

     ...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Architect, you will work on the analysis, verification, and... 
    Senior
    Performance
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    5 days ago
  • $124k - $208.4k

     ...excellence for Intellectual Property (IP) that is applied to high-performance computing devices (mobile, automotive, and other custom...  ...world. Come build with us!Role and ResponsibilitiesAs a Senior Engineer, GPU Modeling Architect, you will work on the design and development... 
    Senior
    Performance
    Hourly pay
    Full time
    Relocation

    Samsung Semiconductor

    San Jose, CA
    5 days ago
  • $184k - $287.5k

    We are now looking for a Senior Formal Verification Engineer for GPU Kernels! Modern AI performance relies on highly optimized GPU kernels — performance-critical code where bugs can be hard to catch and expensive to miss. NVIDIA's Deep Learning Safety Team is hiring engineers... 
    Senior
    Performance
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $182.2k - $273.4k

    Qualcomm is seeking a GPU Engineer in Santa Clara, CA to lead the architecture and optimization of GPU cores. This role involves collaborating with cross-functional teams and utilizing advanced knowledge in GPU technology to fulfill customer needs. Candidates must possess... 
    Senior
    Performance

    Qualcomm

    Santa Clara, CA
    4 days ago
  • $83k - $132k

    CoreWeave is seeking a Bare Metal Support Engineer in Sunnyvale, California. This role involves supporting and maintaining GPU infrastructure to ensure performance and reliability. You will work closely with customers and collaborate with engineering teams to resolve issues... 
    Senior
    Performance

    CoreWeave

    Sunnyvale, CA
    1 day ago
  •  ...delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and architecture decisions, lead performance optimizations, and own the inference engine to scale research and production workloads. This role... 
    Senior
    Performance

    Sanas

    Palo Alto, CA
    2 days ago
  • NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and...  ...drives experimental agents, optimize performance, and contribute to cutting-edge AI research... 
    Senior
    Performance

    Nvidia Corporation in

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise... 
    Senior
    Performance
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H10...  ...Benchmark across hardware tiers and publish honest performance comparisons Upstream fixes and patches to... 
    Senior
    Performance
    Local area
    Shift work

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    4 days ago
  •  ...headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize...  ...proficiency in C/C++/Python on Linux, and experience with distributed, high-performance software. #J-18808-Ljbffr Jobleads-US
    Senior
    Performance

    Jobleads-US

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior GPU Inference Performance Engineer. Be the first to apply!