Senior Software Engineer, Quantized Inference
$152k - $241.5kNVIDIA
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified variants — unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines.Each new recipe demands corresponding kernel and model-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron-LM, ModelOpt, and vLLM.What you'll be doing:Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream servingBuild prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimizationDevelop data analysis tooling and visualizations for numerics debuggingImprove developer productivity across the team: CI, build systems, training infrastructure, pipeline frictionParticipate in code reviews and incorporate feedbackWhat we need to see:Proficient in Python; familiarity with C++Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted toolingExperience with ML accelerators with a basic understanding of how certain ML layers affect execution timeFamiliarity with PyTorch internals (custom ops, autograd, export) or equivalent frameworkExperience reading, modifying, or contributing to a large open-source codebaseMS/PhD in Computer Science or related field, or equivalent experience. 4+ years in a relevant software engineering roleDemonstrated ability to move fast with ambiguous requirements, with strong written and verbal communicationWays to stand out from the crowd:Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel developmentTrack record of debugging numerical issues across mixed-precision boundariesDeep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsityYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 26, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, WA, Redmond; US, CA, Santa ClaraType: Full time
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer, Quantized Inference in Santa Clara, CA vacancy
$152k - $241.5k
NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in...SeniorFull timeRemote work$152k - $241.5k
We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA...SeniorFull time$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...SeniorFull time$193.3k - $261.5k
...(AWS) builds AWS Neuron, the software development kit used to accelerate... ...JAX enabling unparalleled ML inference and training performance.The... ...-software boundary, our engineers build systematic infrastructure... ...-sharing and mentorship. Our senior members enjoy one-on-one...SeniorWork experience placementInternshipLocal areaFlexible hours$152k - $241.5k
...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise...SeniorFull time- ...RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node... ...You will collaborate across internal GPU software teams and engage with open-source... ...software ecosystem. THE PERSON: Skilled engineer with strong technical and analytical expertise...Senior
$152k - $204k
...Nasdaq: CRWV) in March 2025. Learn more at What You'll Do: Senior engineers are area owners who lead designs, raise engineering... ...orchestration, and hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale. About the role...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hoursShift work$139k - $257.55k
...decisions that minimize impact on the customer experience. As a Senior Software Engineer focused on Data Science and Platform Engineering, you'll... ...engineering, offline model development, and real-time inference — building the models that detect and stop fraud and abuse...SeniorFull timeTemporary workLocal areaWorldwideShift work$224k - $356.5k
...Local AI team is building the software stack that makes large... ...in leading open-source LLM inference frameworks — identify performance... ...decoding, multi-token prediction, quantized inference) map onto NVIDIA... ...Computer Science, Computer Engineering, Electrical Engineering, or...SeniorFull timeLocal area$152k - $241.5k
...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team... ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,...SeniorFull timeRemote work$224k - $356.5k
We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers for its GPU-accelerated deep learning... .... We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier...SeniorFull time$184k - $287.5k
...exposure to engage with Large-scale model inference architectureContribute to the... ...vehicle technologies.Deep knowledge of E2E AV software integration from perception through control... ...tuning.What We Need To See:We’re looking for engineers who can build systems — not just...SeniorFull time$152k - $241.5k
...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team... ...automotive and robotics. We build the software stack that enables Large Language, Vision... ...Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of...SeniorFull time$184k - $287.5k
We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate for AI inference. As a member of the team, you'll develop libraries, code generators...SeniorFull timeRemote work$152k - $241.5k
...Language Models and 3D reconstruction? We are looking for a driven Software Engineer to bring ground breaking models into NVIDIA's software... ...you'll be doing:Design, build, and optimize containerized inference execution for the latest 3D VLMs from NVIDIA, turning research...SeniorFull time$170.6k - $261.3k
Job DescriptionAs a Senior Software Engineer on the SimCore team, you will build and deploy applied AI/ML solutions that directly support simulation... ...models, and excel at building robust, high-performance inference pipelines. This role is not focused on training foundation...SeniorFull timeLocal areaWork from homeFlexible hours$152k - $241.5k
...artificial intelligence.We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus... ..., DMA, high-speed interconnects, and distributed training/inference systems.Experience with server management technologies, data...SeniorFull timeRemote work$224k - $356.5k
...Platform Team is building the software foundation for scalable, high... ...are looking for exceptional engineers who thrive on solving deeply... ...automotive computing.We are seeking a Senior Software Engineer for next-... ...core platform, deep learning inference, TensorRT and related...SeniorFull time$184k - $287.5k
...every new AI-powered application is built. We are seeking a Senior Software Engineer focused on container and cloud infrastructure. You will... ...design and implement our core container strategy for NVIDIA Inference Microservices (NIMs) and our hosted services. You will...SeniorFull time$184k - $287.5k
...infrastructure challenges in the field. RL requires inference, rollout generation, and training... ...do.NVIDIA is building an RL Frameworks engineering team to develop the open-source tools... ...depend on. The team spans the full software stack, from collaborating closely with...SeniorFull time$184k - $287.5k
...lasting impact on the world.We are looking for outstanding Senior Deep Learning Software Engineers to develop and productize NVIDIA's deep learning... ...Developing compiler technologies to accelerate deep learning inference on NVIDIA hardware platforms for Physical AI.Working...SeniorFull timeWork experience placement$2,000 per month
...chain-of-thought reasoning agents. Job Summary Etched’s Inference SW team enables optimal mapping of models to Sohu’s dataflow... ...hosts and racks. We are seeking a highly skilled and motivated engineer to join our team as we work towards enabling Mixture-of-...Full timeWork at officeRelocation package$152k - $241.5k
...applications and industries. Within our software stack, CUTLASS stands out as a popular... ...state-of-the-art deep learning models’ inference and training passes to identify key GPU... ...PhD degree in Computer Science, Computer Engineering, or related field (or equivalent...SeniorFull time$184k - $287.5k
We are now looking for a Senior Software Engineer for AI Resiliency!At NVIDIA, we are pushing the boundaries of what’s possible in AI. We are... ...environments, ensuring seamless operation of AI training and inference workloads.What We Need to See:You've achieved a Bachelor’s...SeniorFull time$152k - $241.5k
...role in crafting the digital landscape.Local AI seeks a Senior Systems Software Engineer interested in solving client-side AI challenges on Windows... ...of AI models, data processing pipelines, and inference runtime features.Identifying, evaluating, and implementing...SeniorFull timeLocal area$170.6k - $261.3k
...Overview As a Senior Software Engineer on the SimCore team, you will build and deploy applied AI/ML solutions that directly support simulation... ...models, and excel at building robust, high‑performance inference pipelines. This role is not focused on training foundation...SeniorFlexible hours$184k - $287.5k
...are looking for a motivated Deep Learning engineer to bring advanced CUDA features and... ...from training on scales up to 100K GPUs to inference down at microsecond latency. CUDA features... ...systems principles (aka systems software fundamentals)Adaptability and passion to...SeniorFull time$143k - $191k
...enterprises make better talent decisions. We’re looking for a Software Engineer to design and build highly scalable, user-facing product... ...workflowsDevelop backend services and APIs that support model inference, orchestration, and tool executionImplement LLM workflows, including...SeniorWork at officeRemote workFlexible hours$193.3k - $261.5k
...fully customized stack of hardware, firmware, and software to deliver unparalleled virtualization at a... ...Supercomputers, optimized for high-performance training and inference workloads.We are looking for an experienced software engineer to drive development for new EC2 machine...SeniorInternshipLocal areaFlexible hours$224k - $356.5k
...platform with high visibility and real-world impact. As a System Software Engineer for Vision AI, you will develop and optimize high-... ...perception algorithms at scale.Profiling and tuning GPU-accelerated inference pipelines to meet strict latency, efficiency, and...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, Quantized Inference. Be the first to apply!
Related searches
- ngo software engineer Santa Clara, CA
- software system engineer Santa Clara, CA
- software engineer - early career Santa Clara, CA
- entry level software engineer remote Santa Clara, CA
- software developer positions Santa Clara, CA
- junior software developer internship Santa Clara, CA
- software development engineer aws Santa Clara, CA
- consulting software engineer Santa Clara, CA
- financial software developer Santa Clara, CA
- software developer Santa Clara, CA


