Senior Software Engineer, Quantized Inference
$152k - $241.5kNVIDIA
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified variants — unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines.Each new recipe demands corresponding kernel and model-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron-LM, ModelOpt, and vLLM.What you'll be doing:Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream servingBuild prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimizationDevelop data analysis tooling and visualizations for numerics debuggingImprove developer productivity across the team: CI, build systems, training infrastructure, pipeline frictionParticipate in code reviews and incorporate feedbackWhat we need to see:Proficient in Python; familiarity with C++Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted toolingExperience with ML accelerators with a basic understanding of how certain ML layers affect execution timeFamiliarity with PyTorch internals (custom ops, autograd, export) or equivalent frameworkExperience reading, modifying, or contributing to a large open-source codebaseMS/PhD in Computer Science or related field, or equivalent experience. 4+ years in a relevant software engineering roleDemonstrated ability to move fast with ambiguous requirements, with strong written and verbal communicationWays to stand out from the crowd:Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel developmentTrack record of debugging numerical issues across mixed-precision boundariesDeep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsityYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 24, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, WA, Redmond; US, CA, Santa ClaraType: Full time
Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer, Quantized Inference in Santa Clara, CA vacancy
- NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer specializing in Quantized Inference to speed up LLM deployment. You will implement quantized and sparse recipes in inference engines, optimize export pipelines, and build benchmarks for throughput and...Senior
$152k - $241.5k
...most challenging problems. We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the TensorRT team, you will be responsible...SeniorFull time$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA...SeniorFull time$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...SeniorFull time$165k - $242k
Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence...SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$193.3k - $261.5k
...Amazon builds Amazon Neuron, the software development kit used to... ...JAX enabling unparalleled ML inference and training performance.The... ...hardware-software boundary, our engineers build systematic infrastructure... ...-sharing and mentorship. Our senior members enjoy one-on-one...SeniorWork experience placementInternshipLocal areaFlexible hours$139k - $204k
What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to latency... ..., and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role...SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work$152k - $241.5k
...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise...SeniorFull time$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join... ...on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state...SeniorInternshipLocal areaFlexible hours$152k - $204k
...Nasdaq: CRWV) in March 2025. Learn more at What You'll Do: Senior engineers are area owners who lead designs, raise engineering... ...orchestration, and hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale. About the role...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hoursShift work- Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance benchmarking and competitive pricing models for their... ...candidate will have extensive experience with open-source inference frameworks and an understanding of ML systems. This role...Senior
$224k - $356.5k
...Local AI team is building the software stack that makes large... ...in leading open-source LLM inference frameworks — identify performance... ...decoding, multi-token prediction, quantized inference) map onto NVIDIA... ...Computer Science, Computer Engineering, Electrical Engineering, or...SeniorFull timeLocal area- NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute...Senior
- NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models...Senior
- NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache...Senior
- NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines...Senior
$152k - $241.5k
...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team... ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,...SeniorFull timeRemote work$224k - $356.5k
We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers for its GPU-accelerated deep learning... .... We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier...SeniorFull time- ...Inc. is looking for a Sr. Member of Technical Staff to design software features that enhance system resiliency and high availability... ...distributed environments. The role includes developing scalable AI inference services and deploying cloud-based workflows. Ideal candidates...Senior
- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...Senior
- NVIDIA seeks a Senior Product Manager for AI Platform Inference (Finance) in Santa Clara to lead tooling, SDKs, and libraries enabling GPU-based inference... ...technical product management, knowledge of inference software and GenAI concepts, and strong communication. Equity...Senior
- Apple Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will collaborate with research and external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment, owning hard...Senior
$139k - $204k
CoreWeave is seeking a Senior Engineer to lead designs and enhance engineering standards within their Kubernetes-native inference platform. Responsibilities include driving architecture, defining SLIs/SLOs, and mentoring engineers, with 3-8 years of experience preferred...SeniorRemote jobFlexible hours$165k - $242k
A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed...Senior$184k - $287.5k
...exposure to engage with Large-scale model inference architectureContribute to the... ...vehicle technologies.Deep knowledge of E2E AV software integration from perception through control... ...tuning.What We Need To See:We’re looking for engineers who can build systems — not just...SeniorFull time$230k - $250k
Cerebras Systems is seeking a Sr. Member of Technical Staff in Sunnyvale, CA. This role involves designing resilient software features for cloud-based AI inference, leveraging AWS tools and services. Candidates should have a Master’s degree in Computer Science and experience...Senior$152k - $241.5k
...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team... ...automotive and robotics. We build the software stack that enables Large Language, Vision... ...Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of...SeniorFull time$184k - $287.5k
We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate for AI inference. As a member of the team, you'll develop libraries, code generators...SeniorFull timeRemote work$152k - $241.5k
...Language Models and 3D reconstruction? We are looking for a driven Software Engineer to bring ground breaking models into NVIDIA's software... ...you'll be doing:Design, build, and optimize containerized inference execution for the latest 3D VLMs from NVIDIA, turning research...SeniorFull time$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good... .... Design, develop, test, deploy, and support AI software components including foundation model training,...SeniorFull timePart timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, Quantized Inference. Be the first to apply!
Related searches
- cybersecurity software engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software developer fintech Santa Clara, CA
- new graduate software engineer Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- software engineer visa sponsorship Santa Clara, CA
- software qa engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer remote Santa Clara, CA
- part time software developer remote Santa Clara, CA

