Engineering Manager, Inference Benchmarking — AI Perf (Santa Clara)
$224k - $356.5kNvidia
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s an outstanding legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.NVIDIA’s open-source benchmarking platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud providers, and enterprises use AIPerf to inform decisions on production inference. This includes choosing GPUs, optimizing costs, reducing latency, improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo organization. Your responsibility is to build and advance the platform so AIPerf becomes the leading benchmarking tool for datacenter, local, and edge use cases. This span LLM, multimodal, diffusion, and computer vision inference. This position combines hands-on leadership with expertise in systems engineering, inference infrastructure, and open-source communities. It has a direct effect on how AI performance is measured and pushed forward.What you'll be doing:Driving the technical roadmap for AIPerf's core infrastructure: load generation, ZMQ-based microservices, GPU telemetry (DCGM/PyNVML, Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment.Taking ownership for the accuracy and statistical soundness of benchmark results that engineering groups throughout the industry depend on to inform production infrastructure decisions.Advising upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams to maintain AIPerf's relevance across emerging hardware, workload categories, and inference configurations.Hiring, mentoring, and growing a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide.What we need to see:Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.8+ overall years of software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.3+ years of engineering leadership experience as a tech lead, TLM, or engineering manager.Deep understanding of LLM inference mechanics — TTFT, ITL, KV caching, Prefill/Decode, speculative decoding — and the ability to reason about measurement correctness and reproducibility.Proven track record of collaborating across multi-functional groups and delivering production-quality output in high-velocity, high-external-visibility environments.Ways to stand out from the crowd:Extensive experience with vLLM, TRT-LLM or SGLang internals along with contributions to their upstream projects.Experience building Kubernetes-native infrastructure including operators, Helm charts, and GPU observability tooling (DCGM, dcgm-exporter, PyNVML).Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard evaluation systems.History leading or making meaningful contributions to active open-source projects with external communities.Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until June 1, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, AL, Remote; US, CO, Remote; US, WA, Remote; US, CA, RemoteType: Full time
$184k - $287.5k
...unlimited potential of AI to define the next... ...a highly motivated engineer to lead performance benchmarking and optimization... ...world AI training, inference, and HPC workloads... ...profiling tools (Linux perf, NVIDIA Nsight... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:...SuggestedFull timePart time- ...potential of generative AI to power the... ...with LLM inference on heterogeneous... ...applied research and engineering team that moves fast... ...everything from benchmarking and evaluation pipelines... ...decode, KV-cache management.• Work closely... ...and benefits in Santa Clara, CA.Equal...SuggestedPart time
$272k - $431.25k
...unlimited potential of AI to define the next era... ...Networking Codesign and Benchmarking R&D group requires a senior software engineer. In this exciting role,... ...Learning LLM training and inference. Your primary focus... ...SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, CO...SuggestedFull timePart timeRemote work$272k - $431.25k
...infrastructure that stores, manages, and serves exabytes of... ...backbone for NVIDIA's AI infrastructure, enabling researchers and engineers to reliably store... ...accelerating training and inference pipelines.We are... ...SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full...SuggestedFull timePart time$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...kernels and compilers, drive industry benchmarks, and scale workloads across multi-... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...SuggestedFull timePart time$182.5k - $260.5k
...spread across offices in Santa Clara, St. Louis, Bangalore,... ...Scientist, you own the inference and optimization layer that makes AI in agentic workflows... ...quantization, KV-cache and memory management, sparsity, fine-tuning,... ...systems and backend engineers to ship capabilities...Part timeWork at office$195.2k - $361.2k
...journey is to transform AI into something safer,... ...own. You optimize inference engines (llama.cpp, vLLM) for... ...start / stop / health)Benchmark across hardware tiers... ...Location: US, California, Santa ClaraAdditional Locations... ...US, California, Santa Clara; US, Oregon, Hillsboro...Full timePart timeInternshipLocal areaImmediate startShift work$184k - $287.5k
...Senior DL Algorithms Engineer! NVIDIA is seeking senior... ...that leads the AI revolution.What you will... ...and multimodal model inference as part of NVIDIA Inference... ...inference performance.Benchmark state-of-the-art... ...SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full...Full timePart time$184k - $287.5k
...unlimited potential of AI to define the next era... ...of forward‑thinking engineers tackling some of the globe... ...plugins, distributed inference serving, and major... ...performance modeling and benchmarking for large‑scale... ...SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType:...Full timePart timeRemote work$184k - $287.5k
...that empower NVIDIA engineers to improve perf and power efficiency... ...developing tools for AI researchers and SW/HW... ...or networkingCreate benchmarking and simulation technologies... ...training and inference.Knowledge of GPU cluster... ...SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$184k - $287.5k
...seeking outstanding AI Solutions... ...with ISVs, product, engineering, developer relations... ...dives, POCs, benchmarks, demos, and customer... ..., or systems for managing and processing dataFamiliarity... ...processing, or inference and data platform... ...: US, CA, Santa Clara; US, TX, Remote;...Full timePart timeRemote work$192k - $304.75k
...now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited to change the... ...model systems.Build and run agentic benchmarks (e.g., Terminal-Bench ) to... ...SummaryLocation: US, WA, Remote; US, CA, Santa Clara; US, RemoteType: Full time...Full timePart timeRemote work$184k - $287.5k
...Solutions Architect - AI Factory... ...Specialists team in Santa Clara! This role is uniquely... ...workloads and benchmarks on Linux-based GPU... ..., Mathematics, Engineering, Physics, or related... ...of experience managing Linux-based systems... ...training and/or inference workflows using frameworks...Full timePart timeRemote work$224k - $356.5k
...workstation-class AI computer—built on... ...NemoClaw, LLM inference via NIM, Hermes agents... ...systems software engineer who will own AI... ...) to establish benchmarks and identify regression... ..., and memory management for Blackwell... ...SummaryLocation: US, CA, Santa Clara; US, RemoteType:...Full timePart timeLocal area- ...computing experiences—from AI and data centers, to... ...Solution Engineer, you will partner with... ...scale LLM training and inference on AMD Instinct GPUs.... ...and SLURM controllers.Benchmark and optimize LLM inference... ...field requiredLOCATION:Santa Clara, Ca or open to discuss...Part time
$184k - $287.5k
...NVIDIA and help bring AI solutions to our... ...Sales Account Managers and Developer... ...LLM training and inference.Conducting regular... ...Electrical/Computer Engineering, Computer Science... ...performance benchmarks for data center systems... ...: US, CA, Santa Clara; US, WA, SeattleType...Full timePart time$184k - $287.5k
...network stack for distributed inference which will be used to... ...impact on the world by applying AI inference aware technology to... ...Computer Science, Electrical Engineering, Software Engineer, or related... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart timeWork experience placement$184k - $287.5k
...seeking outstanding AI Solutions... ...across product, engineering, sales, developer... ...training, fine-tuning, inference, retrieval, and agentic... ...dives, POCs, benchmarks, demos, and... ..., telemetry, and management tools to improve... ...SummaryLocation: US, CA, Santa Clara; US, TX, Remote;...Full timePart timeRemote work$184k - $287.5k
...technical expert uniting engineering, field teams, and... ...performance of world-class AI, deep learning, and... ...maintain robust performance benchmarking suites to stress-test... ...systems (e.g., Perf, eBPF, Prometheus, Grafana... ...: US, CA, Santa Clara; US, TX, AustinType: Full...Full timePart time$320k
...world leader in physical AI, powering self-driving... ...team is the execution engine behind NVIDIA’s Vision... ...delivering robust, low-latency inference at scale. You have led... ...partners.Performance Benchmarking: Orchestrate efforts to... ...: US, CA, Santa ClaraType: Full time...Full timePart time$272k - $431.25k
...NVIDIA is seeking a Senior MLOps Engineering Manager to join our Autonomous Driving organization in Santa Clara, CA. This role offers an... ...success metrics, and operational benchmarks, and ensure consistent... ...existing vacancy. NVIDIA uses AI tools in its recruiting processes...Full timePart time$272k - $431.25k
...As a Senior Engineering Manager for Agentic Systems & Platform Architecture... ...backed by evaluations, benchmarking, and feedback... ...optimized training and inference workflows.Lead integration of the AI Data Platform into NVIDIA... ...: US, CA, Santa ClaraType: Full time...Full timePart time$184k - $287.5k
...which every new AI-powered application... ...Senior Software Engineer focused on container... ...for NVIDIA Inference Microservices (NIMs... ...strategy, dependency management, and artifact/... ...LLM)Background in benchmarking and optimizing... ...SummaryLocation: US, CA, Santa Clara; US, CA,...Full timePart time$332k
...into the unlimited potential of AI to define the next era of... ...multi-functionally with Product Management, Product Marketing, Developer... ...Define and Implement a leadership Inference go-to-market strategy!... ...#deeplearningSummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart timeWorldwide$272k - $431.25k
...the unlimited potential of AI to define the next era of computing... ...a highly skilled Senior Engineering Manager to help drive the future of... ...workflowsExperience with AI inference engines and frameworksTrack... ...law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full time...Full timePart timeLocal area$184k - $287.5k
...motivated Deep Learning engineer to bring advanced CUDA... ...technologies into AI stacks, including PyTorch... ...up to 100K GPUs to inference down at microsecond latency... ...performance benchmarking on AI clusters. Familiarity... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US,...Full timePart time- ...generation computing experiences—from AI and data centers, to PCs,... ...and SOTA LLM and Multimodal inference at scale across multi-GPU and... ...ecosystem. THE PERSON: Skilled engineer with strong technical and... ...the ability to define goals, manage development efforts, and deliver...Part time
$224k - $356.5k
...We are now looking for an AI Developer Technology Engineering Manager:Join our global Developer Technology (DevTech) team at NVIDIA, where we drive innovation... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, MA, Westford; US, TX, Austin; US, OR, Hillsboro;...Full timeTemporary workPart time$272k - $431.25k
...looking for a Principal Engineer to join our CSP... ...NVIDIA's performance and benchmark teams with a dedicated... ...environmentsUnderstanding of inference workload performance... ...vacancy. NVIDIA uses AI tools in its... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR...Full timePart timeRemote work$200k - $322k
...Physical AI and Artificial Intelligence (AI)... ...Senior Developer Advocate Engineer to own technical... ...generation, evaluation, benchmarking, inference optimization, and... ...day event and account management.Ability to thrive in... ...SummaryLocation: US, CA, Santa Clara; US, DC, Remote; US,...Full timePart timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Engineering Manager, Inference Benchmarking — AI Perf (Santa Clara). Be the first to apply!





