Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer, Quantized Inference

$152k - $241.5k

NVIDIA

We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified variants — unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines.Each new recipe demands corresponding kernel and model-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron-LM, ModelOpt, and vLLM.What you'll be doing:Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream servingBuild prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimizationDevelop data analysis tooling and visualizations for numerics debuggingImprove developer productivity across the team: CI, build systems, training infrastructure, pipeline frictionParticipate in code reviews and incorporate feedbackWhat we need to see:Proficient in Python; familiarity with C++Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted toolingExperience with ML accelerators with a basic understanding of how certain ML layers affect execution timeFamiliarity with PyTorch internals (custom ops, autograd, export) or equivalent frameworkExperience reading, modifying, or contributing to a large open-source codebaseMS/PhD in Computer Science or related field, or equivalent experience. 4+ years in a relevant software engineering roleDemonstrated ability to move fast with ambiguous requirements, with strong written and verbal communicationWays to stand out from the crowd:Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel developmentTrack record of debugging numerical issues across mixed-precision boundariesDeep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsityYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 26, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, WA, Redmond; US, CA, Santa ClaraType: Full time
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer, Quantized Inference in Santa Clara, CA vacancy
  • $152k - $241.5k

    NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $193.3k - $261.5k

     ...(AWS) builds AWS Neuron, the software development kit used to accelerate...  ...JAX enabling unparalleled ML inference and training performance.The...  ...-software boundary, our engineers build systematic infrastructure...  ...-sharing and mentorship. Our senior members enjoy one-on-one... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $152k - $241.5k

     ...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  •  ...RL training and SOTA LLM and Multimodal inference at scale across multi-GPU and multi-node...  ...You will collaborate across internal GPU software teams and engage with open-source...  ...software ecosystem. THE PERSON:   Skilled engineer with strong technical and analytical expertise... 
    Senior

    AMD

    Santa Clara, CA
    2 days ago
  • $152k - $204k

     ...Nasdaq: CRWV) in March 2025. Learn more at What You'll Do: Senior engineers are area owners who lead designs, raise engineering...  ...orchestration, and hardware teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale. About the role... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    16 days ago
  • $139k - $257.55k

     ...decisions that minimize impact on the customer experience. As a Senior Software Engineer focused on Data Science and Platform Engineering, you'll...  ...engineering, offline model development, and real-time inference — building the models that detect and stop fraud and abuse... 
    Senior
    Full time
    Temporary work
    Local area
    Worldwide
    Shift work

    Adobe

    San Jose, CA
    1 day ago
  • $224k - $356.5k

     ...Local AI team is building the software stack that makes large...  ...in leading open-source LLM inference frameworks — identify performance...  ...decoding, multi-token prediction, quantized inference) map onto NVIDIA...  ...Computer Science, Computer Engineering, Electrical Engineering, or... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team...  ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers for its GPU-accelerated deep learning...  .... We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...exposure to engage with Large-scale model inference architectureContribute to the...  ...vehicle technologies.Deep knowledge of E2E AV software integration from perception through control...  ...tuning.What We Need To See:We’re looking for engineers who can build systems — not just... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team...  ...automotive and robotics. We build the software stack that enables Large Language, Vision...  ...Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

    We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate for AI inference. As a member of the team, you'll develop libraries, code generators... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...Language Models and 3D reconstruction? We are looking for a driven Software Engineer to bring ground breaking models into NVIDIA's software...  ...you'll be doing:Design, build, and optimize containerized inference execution for the latest 3D VLMs from NVIDIA, turning research... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $170.6k - $261.3k

    Job DescriptionAs a Senior Software Engineer on the SimCore team, you will build and deploy applied AI/ML solutions that directly support simulation...  ...models, and excel at building robust, high-performance inference pipelines. This role is not focused on training foundation... 
    Senior
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    9 hours ago
  • $152k - $241.5k

     ...artificial intelligence.We are looking for highly motivated Senior Software Engineers to join our Fabric Networking team with a targeted focus...  ..., DMA, high-speed interconnects, and distributed training/inference systems.Experience with server management technologies, data... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $224k - $356.5k

     ...Platform Team is building the software foundation for scalable, high...  ...are looking for exceptional engineers who thrive on solving deeply...  ...automotive computing.We are seeking a Senior Software Engineer for next-...  ...core platform, deep learning inference, TensorRT and related... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...every new AI-powered application is built. We are seeking a Senior Software Engineer focused on container and cloud infrastructure. You will...  ...design and implement our core container strategy for NVIDIA Inference Microservices (NIMs) and our hosted services. You will... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...infrastructure challenges in the field. RL requires inference, rollout generation, and training...  ...do.NVIDIA is building an RL Frameworks engineering team to develop the open-source tools...  ...depend on. The team spans the full software stack, from collaborating closely with... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...lasting impact on the world.We are looking for outstanding Senior Deep Learning Software Engineers to develop and productize NVIDIA's deep learning...  ...Developing compiler technologies to accelerate deep learning inference on NVIDIA hardware platforms for Physical AI.Working... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $2,000 per month

     ...chain-of-thought reasoning agents. Job Summary Etched’s Inference SW team enables optimal mapping of models to Sohu’s dataflow...  ...hosts and racks. We are seeking a highly skilled and motivated engineer to join our team as we work towards enabling Mixture-of-... 
    Full time
    Work at office
    Relocation package

    Etched

    San Jose, CA
    13 hours ago
  • $152k - $241.5k

     ...applications and industries. Within our software stack, CUTLASS stands out as a popular...  ...state-of-the-art deep learning models’ inference and training passes to identify key GPU...  ...PhD degree in Computer Science, Computer Engineering, or related field (or equivalent... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $184k - $287.5k

    We are now looking for a Senior Software Engineer for AI Resiliency!At NVIDIA, we are pushing the boundaries of what’s possible in AI. We are...  ...environments, ensuring seamless operation of AI training and inference workloads.What We Need to See:You've achieved a Bachelor’s... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...role in crafting the digital landscape.Local AI seeks a Senior Systems Software Engineer interested in solving client-side AI challenges on Windows...  ...of AI models, data processing pipelines, and inference runtime features.Identifying, evaluating, and implementing... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  • $170.6k - $261.3k

     ...Overview As a Senior Software Engineer on the SimCore team, you will build and deploy applied AI/ML solutions that directly support simulation...  ...models, and excel at building robust, high‑performance inference pipelines. This role is not focused on training foundation... 
    Senior
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $184k - $287.5k

     ...are looking for a motivated Deep Learning engineer to bring advanced CUDA features and...  ...from training on scales up to 100K GPUs to inference down at microsecond latency. CUDA features...  ...systems principles (aka systems software fundamentals)Adaptability and passion to... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $143k - $191k

     ...enterprises make better talent decisions. We’re looking for a Software Engineer to design and build highly scalable, user-facing product...  ...workflowsDevelop backend services and APIs that support model inference, orchestration, and tool executionImplement LLM workflows, including... 
    Senior
    Work at office
    Remote work
    Flexible hours

    Eightfold

    Santa Clara, CA
    1 day ago
  • $193.3k - $261.5k

     ...fully customized stack of hardware, firmware, and software to deliver unparalleled virtualization at a...  ...Supercomputers, optimized for high-performance training and inference workloads.We are looking for an experienced software engineer to drive development for new EC2 machine... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

     ...platform with high visibility and real-world impact. As a System Software Engineer for Vision AI, you will develop and optimize high-...  ...perception algorithms at scale.Profiling and tuning GPU-accelerated inference pipelines to meet strict latency, efficiency, and... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer, Quantized Inference. Be the first to apply!