Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Engineering Manager, Inference Benchmarking — AI Perf

$224k - $356.5k

Jobleads-US

About NVIDIA

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's an outstanding legacy of innovation that's fueled by great technology—and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA's open-source benchmarking platform, AIPerf, is the growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud providers, and enterprises use AIPerf to inform decisions on production inference. This includes choosing GPUs, optimizing costs, reducing latency, improving efficiency, and scaling. As Technical Lead Manager, you will lead the engineering team within NVIDIA's Dynamo organization. Your responsibility is to build and advance the platform so AIPerf becomes the leading benchmarking tool for datacenter, local, and edge use cases. This span LLM, multimodal, diffusion, and computer vision inference. This position combines hands-on leadership with expertise in systems engineering, inference infrastructure, and open-source communities. It has a direct effect on how AI performance is measured and pushed forward.

What you'll be doing

  • Driving the technical roadmap for AIPerf's core infrastructure: load generation, ZMQ-based microservices, GPU telemetry (DCGM/PyNVML, Prometheus metrics, statistical confidence intervals, and Kubernetes-native deployment.
  • Taking ownership for the accuracy and statistical soundness of benchmark results that engineering groups throughout the industry depend on to inform production infrastructure decisions.
  • Advising upstream engine integrations involving vLLM, TRT-LLM, and SGLang in partnership with NVIDIA's Dynamo and NIM teams to maintain AIPerf's relevance across emerging hardware, workload categories, and inference configurations.
  • Hiring, mentoring, and growing a team of senior engineers operating in a high-velocity open-source environment with active external contributors worldwide.

What we need to see

  • Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • 8+ overall years of software engineering experience building performance-critical infrastructure, ML tooling, or distributed systems.
  • 3+ years of engineering leadership experience as a tech lead, TLM, or engineering manager.
  • Deep understanding of LLM inference mechanics - TTFT, ITL, KV caching, Prefill/Decode, speculative decoding - and the ability to reason about measurement correctness and reproducibility.
  • Proven track record of collaborating across multi-functional groups and delivering production-quality output in high-velocity, high-external-visibility environments.

Ways to stand out from the crowd

  • Extensive experience with vLLM, TRT-LLM or SGLang internals along with contributions to their upstream projects.
  • Experience building Kubernetes-native infrastructure including operators, Helm charts, and GPU observability tooling (DCGM, dcgm-exporter, PyNVML).
  • Background in competitive benchmarking frameworks such as MLPerf or equivalent industry-standard evaluation systems.
  • History leading or making meaningful contributions to active open-source projects with external communities.

Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until June 1, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

NVIDIA pioneered accelerated computing. Today, our AI infrastructure powers global intelligence, transforming every industry. Learn more about NVIDIA.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Engineering Manager, Inference Benchmarking — AI Perf in Santa Clara, CA vacancy
  • $206k - $333k

     ...The Essential Cloud for AI™. Built for pioneers by...  ...for a Principal Engineer to be the technical lead of CoreWeave's Benchmarking & Performance team. You...  ...If MLPerf (Training & Inference), Working closely with...  ..., and audit trails. Perf Ownership - Lead end-to... 
    Suggested
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    a month ago
  • $224k - $356.5k

     ...the unlimited potential of AI to define the next era of computing...  ...for an experienced Software Engineering Manager to lead the development of...  ...—backed by evaluations, benchmarking, and feedback loops.Providing...  ...GPU-optimized training and inference workflows to deliver best-in... 
    Suggested
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $407.33k

     ...future. Very few people in AI can say this. Every...  ...covers. ️ About our Engineering Teams At Wayve, we...  ..., and training and inference clusters in several regions...  ..., model lifecycle management and infrastructure-as-...  ...dependant): Salaries benchmarked against the market... 
    Suggested
    Full time
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Wayve

    Sunnyvale, CA
    3 days ago
  • $175.8k - $293k

     ...We’re looking for a Principal AI Engineer to architect, build, and harden...  ...tool routing, prompt and context management, orchestration runtimes, and inference serving. Evaluate and adopt emerging...  ...+ online evals, LLM-as-a-judge, benchmarking, CI/CD gates, production... 
    Suggested

    Jobleads-US

    Santa Clara, CA
    1 day ago
  •  ...CoreWeave is seeking an Engineering Manager to lead the marimo team and molab across open-source projects. You will mentor engineers, shape architectural direction, and drive delivery cadence while keeping projects maintainable and responsive to diverse users. You will... 
    Suggested

    Jobleads-US

    Sunnyvale, CA
    1 day ago
  • $130k - $260k

     ...production systems.Establish engineering patterns and best practices for model deployment, feature management, reproducibility, automated testing...  ..., evaluation, deployment, inference, monitoring, and retraining....  ..., the team develops advanced AI-driven solutions that empower... 
    Full time
    Temporary work
    Part time
    Work experience placement

    Walmart

    Sunnyvale, CA
    1 day ago
  • $192k - $278k

    Direct, manage, and set strategy for the power grid distribution and power integrity team...  ...qualifications:Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science,...  ...role, you’ll work to shape the future of AI/ML hardware acceleration. You will have... 
    Worldwide

    Google

    Sunnyvale, CA
    2 days ago
  • $240k - $290k

     ...Sonatus, we're driving the transformation to AI-enabled software-defined vehicles....  ...is looking for an experienced Senior Engineering Manager to build and lead our AI Validation function...  ...scalable evaluation frameworks and benchmarking platforms for AI systems, including multi... 
    Work at office
    Worldwide
    Flexible hours
    Shift work

    Sonatus

    Sunnyvale, CA
    20 days ago
  •  ...We are seeking an Application Engineering Manager to lead a customer-facing systems and application...  ...high-performance compute (HPC), AI, hyperscaler, and advanced digital semiconductor...  ...from requirements definition, benchmarks, and bring-up through production release... 

    Advantest America

    San Jose, CA
    4 days ago
  • $276.1k - $300.4k

     ...run on in the future. Very few people in AI can say this. Every role here, whatever...  ...s what this particular role covers. Engineering Manager, OS, Robot Software Sunnyvale,...  ...Benefits vary by location. Salaries benchmarked against the market annually Meaningful... 
    Full time
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Wayve

    Sunnyvale, CA
    3 days ago
  • $240k - $320k

     ...senior electrical and system integration engineers to define, build, and scale our vehicle...  ...Principal Engineer - ADAS & AV Data Loop & AI Flywheel, you will spearhead the...  ...development framework that accelerates the benchmarking, active learning selection, and continuous... 
    Full time
    Work experience placement
    Local area
    Flexible hours

    Robert Bosch

    Sunnyvale, CA
    3 days ago
  • $234k - $286k

     ...leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers...  ...Principal AI Solutions Engineer: a hands‑on technical...  ...routing Performance and Benchmarking: Benchmark end-to-end...  ...SambaStack deployment, management and integration... 
    Full time
    Local area
    Worldwide

    Jobleads-US

    San Jose, CA
    1 day ago
  • $400k

     ...Principal Research & Engineering, Realtime Voice AI About Inflection AI Inflection AI is a Public...  ...GPU cluster to support performance benchmarking and extensive experimentation. Determine...  ...models, barge‑in, low‑latency inference, or realtime agents. Strong technical... 
    Work at office
    Flexible hours

    Inflection AI, Inc.

    Palo Alto, CA
    1 day ago
  • $278.1k - $417.1k

     ...building the next generation of AI-driven game experiences,...  .... As our Principal Engineer for On-Device AI Inference & Systems, you will be the...  ...path, and frame-budget management alongside the renderer. Architect...  ...and automated on-device benchmarking in CI. Research... 
    Work at office
    Worldwide
    Relocation package

    ironSource

    Mountain View, CA
    8 hours ago
  •  ...intelligence . As the only vertically integrated AI infrastructure company built from the...  ...Crusoe. About the Role: As an Engineering Manager on the Managed AI team at Crusoe, you...  ...with CPU & GPU performance, inference frameworks, or LLM systems is a strong... 
    Temporary work
    Work at office

    Crusoe

    Sunnyvale, CA
    23 days ago
  •  ...investing deeply in Generative AI — pioneering advanced...  ...Principal Machine Learning Systems Engineer (P60) to lead technical directions...  ...large-scale model training, inference pipelines, or search/...  ...NDCG, groundedness, latency benchmarks).Experience with monitoring,... 
    Work at office
    Local area

    Atlassian

    Mountain View, CA
    2 hours ago
  • $262k - $364k

     ...optimization solvers within the GCE Holdback Manager, ensuring real-time machine-drain...  ...Stockroom).Recruit, mentor, and scale an engineering organization of software engineers and Tech...  ...Engine (GCE) core compute and high-priority AI/ML infrastructure.Google Cloud... 

    Google

    Sunnyvale, CA
    2 days ago
  • $50 per hour

     ...capability development and are seeking a motivated Digital Design Manager to join the Engineering & Technology team.What does this role look like?You will...  ...and Missile System design experience.Proficient in AI prompting as it applies to this position and practice.Pay... 
    Full time
    Temporary work
    Work experience placement
    For subcontractor
    Casual work
    Work at office
    Flexible hours

    Lockheed Martin

    Sunnyvale, CA
    1 day ago
  •  ...Tensordyne is seeking a Director of AI Systems Solutions Engineering in Sunnyvale, CA to own and grow our...  ...hardware accelerators, and production inference, plus the ability to lead a small,...  ...and strategic customers to evaluate, benchmark, and deploy Tensordyne systems,... 

    Jobleads-US

    Sunnyvale, CA
    1 day ago
  •  ...Director, AI Systems Solutions Engineering Sunnyvale, CA About Tensordyne...  ...building a new class of AI inference system designed for high-...  ...responsible for customer benchmarking, technical evaluation, NPI...  ...capable of independently managing sophisticated technical engagements... 
    Remote work

    Jobleads-US

    Sunnyvale, CA
    1 day ago
  • $207k - $300k

     ...mentor, and scale a high-performing team of 8 engineers, setting the technical direction and...  ...footprint.Drive the integration of generative AI into the software development lifecycle (...  ...Cloud building or operating large-scale, managed services.3 years of experience in a... 
    Shift work

    Google

    Sunnyvale, CA
    3 days ago
  •  ...the infrastructure behind the AI-driven data economy.As AI...  ...generates data that must be stored, managed, and made accessible over...  ...where we come in.We combine deep engineering expertise with global-scale...  ....Engineer for latency, inference cost, and cloud spend as first... 
    Temporary work
    Immediate start
    Remote work
    Worldwide
    Flexible hours
    Shift work

    Western Digital

    San Jose, CA
    3 days ago
  • $262.7k - $355.4k

    Job Summary:Are you passionate about optimizing AI workloads and delivering real-world performance improvements...  ...on edge devices?We’re looking for an experienced engineer to help customers achieve best-in-class inference performance for production AI models running on Arm... 
    Work at office
    Local area

    ARM

    San Jose, CA
    3 days ago
  • $220k - $300k

     ...is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide...  ...Senior Principal Machine Learning Engineer, you will be responsible for...  ...organizational direction without direct management authority ~ Track record of... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    SambaNova

    San Jose, CA
    4 days ago
  • $140k - $210k

     ...involved. If you want to make an impact on a global scale, come make a difference at Fiserv.Job TitleManager, Engineering (AI/ML)What does a successful Engineering Manager (AI/ML/DE) do at Clover?The Data Science team at Clover is responsible for improving the merchant... 
    Full time

    Fiserv

    Sunnyvale, CA
    3 days ago
  • $219k - $351k

     ...customers, partners, and communities.Principal Engineer, Architecture & Performance Research Engineer for Data Center and Agentic AI CPUWhat You’ll DoArchitecture Research Lab...  ...by our human recruiting team and hiring managers to ensure every candidate is evaluated fairly... 
    Work at office
    Flexible hours
    Shift work

    Samsung Semiconductor

    San Jose, CA
    1 day ago
  • $200k - $275k

     ...Intelligent Edge. ADI combines analog, digital, AI, and software technologies into solutions...  ...are seeking a Principal Forward Deployed Engineer to partner directly with business...  ...patterns across AI productivity, knowledge management, customer support, engineering... 
    Permanent employment
    Full time
    Work at office
    Day shift

    Analog Devices

    San Jose, CA
    2 hours ago
  • $180k - $255k

     ...SambaNova is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide. At the core of SambaNova's technology...  ...SambaNova is looking for a Principal RTL Design Engineer to own the microarchitecture and RTL for our next-... 
    Full time
    Temporary work
    Local area
    Worldwide
    Flexible hours

    SambaNova

    San Jose, CA
    4 days ago
  •  ...the Kubernetes-native AI infrastructure company...  ...Mirantis empowers platform engineering teams to deliver...  ...driven control needed to manage infrastructure with...  ...stood up training and inference workloads, argued interconnect...  ...architectures, TCO/benchmark models, POV playbooks,... 

    Mirantis

    San Jose, CA
    a month ago
  • $180k - $225k

     ...believe the future of work is Human + AI and are building an AI-native...  ...Zscaler.RoleWe are looking for a Senior Manager Software Engineering to join our team. This is a Hybrid (3...  ...#LI-HybridZscaler’s salary ranges are benchmarked and are determined by role and level.... 
    Full time
    Work at office
    Local area

    Zscaler

    San Jose, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Engineering Manager, Inference Benchmarking — AI Perf. Be the first to apply!