Senior GPU Inference Performance Engineer (Santa Clara)
AMD
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, and AI serving frameworks, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.THE PERSON:A hands-on performance engineer who is equally comfortable reading a GPU trace and briefing executives. You are curious, evidence-driven, rigorous and you don't stop at X is faster, you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, and AI serving framework teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, systems, and AI.KEY RESPONSIBILITIES:Full-stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, and PCIe/Infinity Fabric data movement.AI serving framework performance: Profile and optimize inference engines including vLLM, SGLang, and emerging serving runtimes. Understand KV-cache management, continuous batching, PagedAttention, speculative decoding, and quantization (FP8, MXFP4, INT4) effects on throughput and latency.Competitive performance analysis: Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why performance differs — attributing gaps to specific hardware features (HBM bandwidth, compute density, interconnect topology), software maturity (kernel libraries, operator fusion, graph compilation), or configuration differences.Multi-server inference networking: Profile and optimize distributed inference topologies including prefill-decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, and NIC-level counters (Pensando, ConnectX). Quantify the impact of network latency, bandwidth, and congestion on end-to-end inference SLAs.GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold-start, and resource contention issues in production serving environments.Tooling and automation: Build reproducible benchmarking harnesses, profiling scripts, and performance regression dashboards. Automate trace collection and analysis to support continuous performance validation across driver, firmware, and framework updates.PREFERRED EXPERIENCE:Background in GPU performance engineering, HPC, or systems performance analysisHands-on proficiency with either AMD (ROCm, rocProfiler, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep understanding of GPU architecture: warp/wavefront execution, memory hierarchy (registers → LDS/shared → L2 → HBM), occupancy, and instruction-level parallelismExperience profiling vLLM, SGLang, or equivalent LLM serving frameworks, including quantization workflows (FP8, MXFP4, INT4, AWQ, GPTQ) and their performance implicationsExperience with multi-GPU and multi-node inference — tensor parallelism, pipeline parallelism, or PD disaggregation over RDMA/RoCE — including RCCL/NCCL profiling and network tools (perftest, ib_write_bw, tcpdump, Memory Fabric counters)Demonstrated ability to explain performance differences in written reports or presentations — not just X is faster but why, rooted in hardware and software evidenceStrong Python and C/C++ skills; comfort reading GPU kernel code (HIP/CUDA)Experience with Kubernetes GPU scheduling, MIG, and GPU operator performance, or contributions to open-source inference or profiling projectsACADEMIC CREDENTIALS:Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desiredThis role is not eligible for visa sponsorship.#LI-TB1#LI-HybridBenefits offered are described: AMD benefits at a glance.AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.This posting is for an existing vacancy.
$184k - $287.5k
...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking... ...who are mindful of performance analysis and optimization... ...software stack from GPU architecture to Deep... ...multimodal model inference as part of NVIDIA... ...SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType:...SeniorPerformanceFull timePart time$160k - $253k
...computing is the engine of artificial intelligence... ...integrate high performance compute,... ...are looking for a Senior Technical Marketing... ...showcasing NVIDIA's GPU architecture, server... ...efficiency for AI inference & training.What you... ...SummaryLocation: US, CA, Santa Clara; US, CA,...SeniorPerformanceFull timePart time$184k - $287.5k
...are seeking a self‑motivated senior engineer for the Aerial Omniverse... ...implementation of a real‑time, GPU‑accelerated propagation engine... ...in computer graphics, high‑performance computing, computational... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time...SeniorPerformanceFull timePart time$195.2k - $361.2k
...actually own. You optimize inference engines (llama.cpp, vLLM) for... ...edge environments — GPU/iGPUs, Vulkan backends... ...and publish honest performance comparisonsUpstream fixes... ...: US, California, Santa ClaraAdditional Locations... ...US, California, Santa Clara; US, Oregon, Hillsboro...SeniorPerformanceFull timePart timeInternshipLocal areaImmediate startShift work$200k - $322k
...are seeking a self‑motivated senior engineer for the Aerial Omniverse... ...you will design and implement GPU kernels that apply time‑varying... ...we need to see:PhD in high‑performance computing, computer... ...law.SummaryLocation: US, CA, Santa Clara; US, CO, RemoteType: Full time...SeniorPerformanceFull timePart time$184k - $287.5k
...skilled and motivated software engineers to join us and build AI inference systems that serve large-... ...architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...SeniorPerformanceFull timePart time$184k - $287.5k
...NVIDIA's invention of the GPU 1999 sparked the growth... ...company”.We're looking for a Senior Performance Compiler Engineer to join our team and work... ...both training and inference. You will be immersed in... ...US, CA, Remote; US, CA, Santa Clara; US, WA, SeattleType: Full...SeniorPerformanceFull timePart timeRemote work$184k - $287.5k
...NVIDIA is seeking an NCX Senior Engineer to join our DSX team,... ...realize efficient performance from NVIDIA's AI... ...distributed training, inference optimization, and MLOps... ...Kubernetes, containers, and GPU scheduling systems... ...: US, CA, Santa Clara; US, Remote; US, WA,...SeniorPerformanceFull timePart timeRemote work$184k - $287.5k
...globally. We seek a Senior Engineer to lead technical efforts... ...powerful local inference (Nemotron models) with... ...Ollama, Llama.cpp, vLLM), GPU-accelerated computing... ...C++ (for performance-critical systems/OS integration... ...: US, CA, Santa Clara; US, WA, RedmondType:...SeniorPerformanceFull timePart timeLocal areaShift work$184k - $287.5k
...are now looking for a Senior GPU & Deep Learning... ...delivering the highest performance in the world for deep... ...workloads, both training and inference, and maintain our... ...Science, Electrical Engineering or Computer Engineering... ...: US, CA, Santa Clara; US, TX, Austin; US,...SeniorPerformanceFull timePart time$184k - $287.5k
...unlock maximum hardware performance for emerging AI... ...systems from a single GPU to supercomputer clusters... ...in both training and inference pipelines.Collaborate... ...Computer Science, Computer Engineering, Electrical... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, TX...SeniorPerformanceFull timePart time$184k - $287.5k
...support in understanding performance aspects related to... ...scale LLM training and inference.Conducting regular technical... ...Electrical/Computer Engineering, Computer Science,... ...Hands-on experience with GPU systems in general... ...SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType:...SeniorPerformanceFull timePart time$272k - $431.25k
...NVIDIA is the engine of modern AI, and robotics is... ...grow, and mentor a high-performing team of robotics... ...training, and on-robot inference on Jetson and edge platforms... ...understanding of GPU-accelerated computing... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, VA...SeniorPerformanceFull timePart time$224k - $356.5k
.... An era in which our GPU acts as the brains of... ...We own the platform — performance, CI/CD pipelines, validated... ...open-source LLM inference frameworks — identify... ...Computer Science, Computer Engineering, Electrical... ...SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US,...SeniorPerformanceFull timePart timeLocal area$184k - $287.5k
...are seeking for an expert Senior Compiler Engineer to join our Compute Compiler... ...effectivelyPartner with architecture, performance, and product teams to... ...or reviewerBackground in GPU architectures, CUDA, or... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, TX,...SeniorPerformanceFull timePart timeRemote work$184k - $287.5k
...This role will be part of an engineering team developing, scaling,... ...designing and developing high-performance software and want to help us... ...and developing and optimizing GPU accelerated algorithms... ...law.SummaryLocation: US, CA, Santa Clara; US, NY, New YorkType: Full...SeniorPerformanceFull timePart time$184k - $287.5k
...computing. An era in which our GPU acts as the brains of... ...enable runs of demanding high performance computing, and computationally... ...Computer Science, Electrical Engineering or related field or... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, AustinType: Full time...SeniorPerformanceFull timePart time$184k - $287.5k
...We are looking for software engineers to join our development efforts... ...algebra kernels for high-performance libraries such as cuSOLVER. Around... ...know our team develops the GPU accelerated libraries and SDKs... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full...SeniorPerformanceFull timePart time$184k - $287.5k
...Libraries team is looking for a senior engineer to join our development... ...know our team develops the GPU accelerated mathematical libraries... ...high quality and performance numerical dense linear algebra... ...law.SummaryLocation: US, CA, Santa Clara; US, PA, Remote; US, WA, Remote...SeniorPerformanceFull timePart timeRemote work$242.35k - $363k
...networks that demand uncompromising performance.As a Senior Distinguished Engineer, you will be one of the most... ...technology that fuels large-scale GPU clusters and AI fabrics.• Our scale... ...requires full onsite collaboration in Santa Clara—where silicon, system, and...SeniorPerformancePermanent employmentPart timeInternshipWork from home$184k - $287.5k
...transforming every industry. GPU-accelerated deep... ...an exceptional Senior Perception Engineer to help design and productize... ...quantify perception performance; analyze large-scale... ...training or inference pipelines through custom... ...SummaryLocation: US, CA, Santa ClaraType: Full time...SeniorPerformanceFull timePart time$224k - $356.5k
...transforming every industry. GPU-accelerated deep... ...an exceptional Senior Radar Perception Engineer to help design and productize... ...radar perception performance; analyze large-scale... ...training or inference pipelines through custom... ...: US, CA, Santa ClaraType: Full time...SeniorPerformanceFull timePart time- ...working onsite at our Santa Clara, CA, headquarters 3... ...days per week.The role: Senior toStaff Runtime... ...memory compute for AI inference in datacenters.This position... ...for runtime software engineering, working on the... ...all aspects of runtime performance of the silicon product...SeniorPerformancePart time3 days per week
$168k - $264.5k
...for hardworking systems engineers who will craft FPGA... ...are now looking for a Senior Systems Prototyping Engineer... ...team onsite in Santa Clara, CA.What you'll be doing... ...place and route.Improve performance of the prototype,... ...NVLINK, USB, CHI and CPU-GPU CoherencyGood debugging...SeniorPerformanceFull timePart time$168k - $264.5k
...for hardworking systems engineers who will craft FPGA... ...are now looking for a Senior Systems Prototyping and... ...Emulation team onsite in Santa Clara, CA.What you’ll be... ...compilation flows.Improve performance of FPGA prototypes and... ..., USB, CHI, and CPU-GPU coherency.Hands-on...SeniorPerformanceFull timePart time$184k - $287.5k
...We’re currently seeking a Senior Developer Technology Engineer, Artificial Intelligence!... ...the best possible performance of computer hardware? Could... ...and develop techniques to GPU accelerate workloads in deep... ...SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX...SeniorPerformanceFull timePart timeWork experience placement$184k - $287.5k
...The role of a Deep Learning Systems Engineer would be to analyze the performance and power consumption of deep... ...this role you will find how CPU, GPU, networking, and IO relate to deep... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, RedmondType: Full time...SeniorPerformanceFull timePart time$184k - $287.5k
...We’re currently seeking a Senior Developer Technology Engineer!NVIDIA's Developer Technology... ...with key customers to perform in-depth analysis and optimization... ...expertise with CPU and GPU architecture fundamentals.... ....SummaryLocation: US, CA, Santa Clara; US, CA, Remote; US, NY,...SeniorPerformanceFull timePart timeWork experience placementRemote work$184k - $287.5k
...We are now seeking a Senior Developer Technology Engineer for the Public Sector!NVIDIA... ...and develop techniques to GPU-accelerate leading applications... ...AI in HPC. You will be performing in-depth analysis and... ...SummaryLocation: US, CA, Santa Clara; US, DC, Remote; US, CA,...SeniorPerformanceFull timePart timeWork experience placementRemote work- ...Senior Director, CTIO Engineering TechnologistsInfrastructure Solutions Group (ISG) builds... ...Team in Austin, Texas or Santa Clara, California.What you’ll... ...technology investigations, performs a strategic analysis of... ...infra including large scale GPU servers, liquid cooling and...SeniorPerformancePart time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior GPU Inference Performance Engineer (Santa Clara). Be the first to apply!
- senior lighting artist Santa Clara, CA
- senior hvac project manager Santa Clara, CA
- senior technical product manager Santa Clara, CA
- senior medical science liaison Santa Clara, CA
- senior accountant remote Santa Clara, CA
- senior app developer Santa Clara, CA
- senior marketing account manager Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- sr project manager Santa Clara, CA
- senior compensation manager Santa Clara, CA





