Senior Software Engineer, AI Inference Systems
$184k - $287.5kNVIDIA
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry benchmarks, and scale workloads across multi-GPU, multi-node, and multi-cloud environments. You’ll collaborate across inference, compiler, scheduling, and performance teams to push the frontier of accelerated computing for AI.What you’ll be doing:Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding, data/tensor/expert/pipeline-parallelism, prefill-decode disaggregation.Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization; build and extend high-level DSLs and compiler infrastructure to boost kernel developer productivity while approaching peak hardware utilization.Define and build inference benchmarking methodologies and tools; contribute both new benchmark and NVIDIA’s submissions to the industry-leading MLPerf Inference benchmarking suite.Architect the scheduling and orchestration of containerized large-scale inference deployments on GPU clusters across clouds.Conduct and publish original research that pushes the pareto frontier for the field of ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into NVIDIA’s software products.What we need to see:Bachelor’s degree (or equivalent expeience) in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE) with 7+ years of experience; alternatively, Master’s degree in CS/CE/SE with 5+ years of experience; or PhD degree with the thesis and top-tier publications in ML Systems, GPU architecture, or high-performance computing.Strong programming skills in Python and C/C++; experience with Go or Rust is a plus; solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, distributed systems, deep learning theories.Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).Experience with containers and orchestration (Docker, Kubernetes, Slurm); familiarity with Linux namespaces and cgroups.Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.Ways to stand out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work with ML compilers and DSLs (e.g., Triton, TorchDynamo/Inductor, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).Experience contributing to containerization/virtualization technologies such as containerd/CRI-O/CRIU.Experience with cloud platforms (AWS/GCP/Azure), infrastructure as code, CI/CD, and production observability.Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential. Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.#LI-HybridYour base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until May 2, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time
$224k - $356.5k
We are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA is hiring software engineers... ...using GPUs to power a revolution in AI, enabling breakthroughs in problems... ...-paced team building Generative AI inference platform to make design and deployment...SeniorFull time- NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute...Senior
- NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models...Senior
- NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines...Senior
- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...Senior
$184k - $287.5k
NVIDIA is the platform upon which every new AI-powered application is built. We are seeking a Senior Software Engineer - AI Inference Performance to advance innovative LLM and... ...performance limits on NVIDIA GPU-accelerated systems. Your work will span models, serving...SeniorFull time$152k - $241.5k
...We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate... ...across the team: CI, build systems, training infrastructure, pipeline... ...concise, well-tested code; fluent with AI-assisted toolingExperience with ML...Senior$193.3k - $261.5k
...Amazon Neuron, the software development kit... ...enabling unparalleled ML inference and training... ...software boundary, our engineers build systematic... ...'s possible in AI acceleration.As part... ...the stack from system level optimizations... ...and mentorship. Our senior members enjoy one-...SeniorWork experience placementInternshipLocal areaFlexible hours$184k - $287.5k
...and highly motivated software professional to work on... ...CUDA and Deep Learning Systems. As the complexity and... ...for emerging AI workloads. You will be... ...in both training and inference pipelines.Collaborate... ...Computer Science, Computer Engineering, Electrical Engineering...SeniorFull time$152k - $241.5k
...eager to work on cutting-edge AI technology for safety-... ...NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of... ...enabling high-performance AI inference solutions for automotive safety... ...of functions, classes, and systems to support certification and...SeniorFull time- ...unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the... ...3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires...3 days per week
$152k - $241.5k
## Senior Software Engineer, Deep Learning Inference - TensorRTApplylocations: US, CA, Santa Claratime type: Full timeposted... ...crowd:*** Experience developing System Software.* Proficiency in Python as... ...an existing vacancy.NVIDIA uses AI tools in its recruiting processes.NVIDIA...Senior$165k - $242k
Apply for the Senior Software Engineer II, Inference role at CoreWeave. CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of... ...industry experience building distributed systems or cloud services. Strong coding in...SeniorPermanent employmentTemporary workCasual workWork at officeRemote workFlexible hoursShift work- ...developer of Embodied AI technology. Our advanced AI software and foundation... ...automated driving systems. Our vision is to... ...as machine learning inference on the edge. Fault... ...Product, Project, and Engineering management to advise... ...for success as a Senior Software Engineer at...Senior
- Adobe Firefly's Generative AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across Adobe products like Firefly,... ...Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency,...Senior
$184k - $287.5k
...latency to enable real-time AI. Modern AI is no longer... ...GPU memory, and scalable, software-defined sensor integration... ...surgical robotics, autonomous systems, industrial automation, and... ...scale.We’re looking for a Senior Systems Software Engineer to help architect and scale...SeniorFull time$152k - $241.5k
...passionate about redefining how software is built in the age of Generative AI? Join NVIDIA’s TensorRT team... ...entry point for out-of-framework inference globally. We are moving beyond... ...scale.If you are a systems-thinking C++ engineer who wants to help scale out an...SeniorFull time- Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture... ...industry-leading training and inference speeds; over 10 times faster than... ...team builds the software systems that power engineering workflows across Cerebras.Our...Senior
$152k - $241.5k
...deep learning ignited modern AI — the next era of... ...& Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning... ...generative AI, recommendation systems, image classification,... ...the backbone of NVIDIA’s inference engine, spanning across data...SeniorFull timeRemote work$184k - $287.5k
...forefront of the generative AI revolution! The... ...diffusion models for maximal inference efficiency using... ...developing an innovative software platform (TRT Model... ...externally by research and engineering teams alike developing... ...are now looking for a Senior Deep Learning Software...SeniorFull time$193.3k - $261.5k
We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud... ....As a Sr. Software Development Engineer on the Inference Model Enablement team, you will onboard... ...excellence in performance optimization and system reliability across the Neuron ecosystem...SeniorInternshipLocal areaFlexible hours$174k - $252k
Write and test product or system development code. Review code developed by other engineers and provide feedback to ensure best practices... ..., maintaining, or launching software products, and 1 year of... ...connectivity, mobile, and now, AI. Google's XR team is at the forefront...Senior$184k - $287.5k
NVIDIA is searching for a creative and highly motivated engineer with expertise in System Software to join the Chips System Software organization. You will... ...26.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to...SeniorFull timeWork at officeLocal area- Cerebras Systems, Inc. is looking for a Sr. Member of Technical Staff to design software features that enhance system resiliency and high availability across distributed... .... The role includes developing scalable AI inference services and deploying cloud-based workflows...Senior
$184k - $287.5k
...into the unlimited potential of AI to define the next era of... ...doing:Develop use cases and system requirements for L3 and L4 autonomous... ...with Data Analytics, Test Engineering, and System Integration &... ...analysis, data analysis, and software architecture.Strong software...SeniorFull time$174k - $252k
...experience.5 years of experience with software development in C/C++.3 years... ...with embedded operating systems.Experience in taking consumer... ...RTOS) systems.Google's software engineers develop the next-generation... ...connectivity, mobile, and now, AI. Google's XR team is at the forefront...Senior$184k - $287.5k
We are looking for a Senior System Software Engineer, Software Defined Networking to design, build, and operate... ...scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated... ...including hyperscale multi-node training, inference, cloud gaming, and cloud functions....SeniorFull time$230k - $250k
Cerebras Systems is seeking a Sr. Member of Technical Staff in Sunnyvale, CA. This role involves designing resilient software features for cloud-based AI inference, leveraging AWS tools and services. Candidates should have a Master’s degree in Computer Science and experience...Senior$174k - $253k
Write and test product or system development code. Review code developed by other engineers and provide feedback to ensure best practices... ..., maintaining, or launching software products, and 1 year of... ...fastest experience possible.The AI and Infrastructure team is redefining...SeniorWorldwide$174k - $252k
Write and test product or system development code. Review code developed by other engineers and provide feedback to ensure best practices... ..., maintaining, or launching software products, and 1 year of... ...enhance software solutions.The AI and Infrastructure team is redefining...SeniorWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer, AI Inference Systems. Be the first to apply!
- cybersecurity software engineer Santa Clara, CA
- graduate software engineer Santa Clara, CA
- software developer fintech Santa Clara, CA
- new graduate software engineer Santa Clara, CA
- senior robotics software engineer Santa Clara, CA
- software engineer visa sponsorship Santa Clara, CA
- software qa engineer Santa Clara, CA
- network software engineer Santa Clara, CA
- software engineer remote Santa Clara, CA
- part time software developer remote Santa Clara, CA

