Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Software Engineer, Quantized Inference

$152k - $241.5k

NVIDIA

Senior Software Engineer for Quantized InferenceNVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified variants — unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines.Each new recipe demands corresponding kernel and model-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron-LM, ModelOpt, and vLLM.What You'll Be DoingImplement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)Own model export pipelines (ModelOpt, Megatron-LM HuggingFace), ensuring quantized checkpoints serialize correctly for downstream servingBuild prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimizationDevelop data analysis tooling and visualizations for numerics debuggingImprove developer productivity across the team: CI, build systems, training infrastructure, pipeline frictionParticipate in code reviews and incorporate feedbackWhat We Need To SeeProficient in Python; familiarity with C++Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted toolingExperience with ML accelerators with a basic understanding of how certain ML layers affect execution timeFamiliarity with PyTorch internals (custom ops, autograd, export) or equivalent frameworkExperience reading, modifying, or contributing to a large open-source codebaseMS/PhD in Computer Science or related field, or equivalent experience.4+ years in a relevant software engineering roleDemonstrated ability to move fast with ambiguous requirements, with strong written and verbal communicationWays To Stand Out From The CrowdExperience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel developmentTrack record of debugging numerical issues across mixed-precision boundariesDeep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsityYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 26, 2026.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Senior Software Engineer, Quantized Inference in Santa Clara, CA vacancy
  • $152k - $241.5k

     ...most challenging problems. We're seeking talented and motivated engineers to join our TensorRT team in developing the industry-leading deep learning inference software for NVIDIA AI accelerators. As a Senior Software Engineer in the TensorRT team, you will be responsible... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in Deep Learning by helping build a state-of-the-art inference framework for accelerating Deep Learning models, especially Large Language Models, on NVIDIA... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...NVIDIA is seeking a Senior Software Engineer for Deep Learning Inference to help build a state-of-the-art inference framework on NVIDIA GPUs, accelerating large language models. You will join the TensorRT Workflows team and tackle scalable, real-time inferencing challenges... 
    Senior

    NVIDIA AI

    Santa Clara, CA
    16 hours ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    11 hours ago
  • $193.3k - $261.5k

     ...Amazon builds Amazon Neuron, the software development kit used to...  ...JAX enabling unparalleled ML inference and training performance.The...  ...hardware-software boundary, our engineers build systematic infrastructure...  ...-sharing and mentorship. Our senior members enjoy one-on-one... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • $165k - $242k

     ...Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence... 
    Senior
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    16 hours ago
  • $139k - $204k

     ...What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to...  ...orchestration, and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The... 
    Senior
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    16 hours ago
  • $139k - $204k

    What You’ll Do Senior engineers are area owners who lead designs, raise engineering standards, and deliver measurable improvements to latency...  ..., and hardware teams to evolve our Kubernetes‑native inference platform and meet strict P99 SLAs at scale. About The Role... 
    Senior
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours
    Shift work

    CoreWeave

    Sunnyvale, CA
    2 days ago
  • $170k - $216k

     ...developer velocity. We’re looking for a software engineer to join the team to build and maintain the...  ...will report to the Head of ML Platform- Senior Staff Software Engineer. You will: Develop Waymo's inference platform to make it scalable, high... 
    Senior
    Full time
    Remote work

    Waymo

    Mountain View, CA
    2 days ago
  • $193.3k - $261.5k

    We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Join...  ...on the Trainium hardware.As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard and optimize state... 
    Senior
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $189k - $301k

     ...Conductor in San Jose, CA is seeking a seasoned engineer to lead co-design efforts for optimizing AI model inference performance. The role requires a deep understanding of AI infrastructure, covering everything from model definition to serving. The ideal candidate... 
    Senior

    Conductor

    San Jose, CA
    16 hours ago
  • $224k - $356.5k

     ...Local AI team is building the software stack that makes large...  ...in leading open-source LLM inference frameworks — identify performance...  ...decoding, multi-token prediction, quantized inference) map onto NVIDIA...  ...Computer Science, Computer Engineering, Electrical Engineering, or... 
    Senior
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute... 
    Senior

    Nvidia Corporation in

    Santa Clara, CA
    15 hours ago
  •  ...NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models... 
    Senior

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • Cerebras Systems, Inc. is looking for a Senior Performance Engineer to enhance the performance benchmarking and competitive pricing models for their...  ...candidate will have extensive experience with open-source inference frameworks and an understanding of ML systems. This role... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    1 day ago
  • $224k - $356.5k

    We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers for its GPU-accelerated deep learning...  .... We are a fast-paced team building Generative AI inference platform to make design and deployment of new AI models easier... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines... 
    Senior

    NVIDIA

    Santa Clara, CA
    16 hours ago
  • $152k - $241.5k

     ...company”.We are looking for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI Compiler (DLC) team...  ...of AI, our DLC has been the backbone of NVIDIA’s inference engine, spanning across data centers, personal devices,... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...NVIDIA Corporation is seeking a Senior Software Engineer for the TensorRT Edge-LLM team in the US. You will develop a high-performance inference framework in modern C++ that extends TensorRT for autoregressive model serving, including speculative decoding and KV cache... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    16 hours ago
  • Intel is seeking a seasoned software engineer to accelerate AI inference on edge hardware. You will optimize llama.cpp/vLLM, tune KV cache, batching and scheduling, and push quantization strategies to balance speed and quality. This role focuses on low-latency, privacy... 
    Senior

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    22 hours ago
  • Intel in Santa Clara, CA, seeks an experienced software engineer to optimize local inference for edge devices. You will work on llama.cpp, vLLM, and quantization, improving latency, memory usage, and startup times across hardware tiers. You will profile performance, collaborate... 
    Senior
    Local area

    Intel

    Santa Clara, CA
    22 hours ago
  •  ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will collaborate...  ..., while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The... 
    Senior
    Local area

    NVIDIA

    Santa Clara, CA
    16 hours ago
  •  ...Apple Inc. is seeking a Sr. Machine Learning Engineer for the Foundation Models Inference team in Santa Clara, CA. You will collaborate with research and external partners to bring cutting-edge model architectures from prototype to planetary-scale deployment, owning hard... 
    Senior

    Apple Inc.

    Santa Clara, CA
    16 hours ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 
    Senior

    Jobleads-US

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...exposure to engage with Large-scale model inference architectureContribute to the...  ...vehicle technologies.Deep knowledge of E2E AV software integration from perception through control...  ...tuning.What We Need To See:We’re looking for engineers who can build systems — not just... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $139k - $204k

     ...CoreWeave is seeking a Senior Engineer to lead designs and enhance engineering standards within their Kubernetes-native inference platform. Responsibilities include driving architecture, defining SLIs/SLOs, and mentoring engineers, with 3-8 years of experience preferred... 
    Senior
    Remote work
    Flexible hours

    CoreWeave

    Sunnyvale, CA
    16 hours ago
  • $165k - $242k

     ...A cloud service provider is seeking a Senior Software Engineer II for their Inference team in Sunnyvale, California. In this role, you'll lead design reviews, implement optimizations, and improve service reliability. The ideal candidate has extensive experience with distributed... 
    Senior

    CoreWeave

    Sunnyvale, CA
    16 hours ago
  • $152k - $241.5k

     ...limits of real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM team...  ...automotive and robotics. We build the software stack that enables Large Language, Vision...  ...Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years of... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...Inc. is looking for a Sr. Member of Technical Staff to design software features that enhance system resiliency and high availability...  ...distributed environments. The role includes developing scalable AI inference services and deploying cloud-based workflows. Ideal candidates... 
    Senior

    Cerebras Systems, Inc.

    Sunnyvale, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Software Engineer, Quantized Inference. Be the first to apply!