Senior Software Engineer - AI Inference
$152k - $241.5kNVIDIA
NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer – AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in‑class on NVIDIA GPUs and systems-and by improving the underlying stack that enables high‑throughput, low‑latency inference at scale.This is a hands-on role for an engineer who enjoys digging into performance bottlenecks, designing pragmatic runtime improvements, and shipping high‑quality changes that are broadly useful to the community and production deployments.What you'll be doing:Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion.Implement and optimize inference‑runtime capabilities: batching and scheduling policies, streaming, request lifecycle management, and KV‑cache efficiency (paging/sharding) to improve throughput and tail latency.Profile and improve hot paths across layers-from Python orchestration to C++/CUDA kernels-using data to guide optimization work.Improve multi‑GPU inference performance and reliability: parallelism strategies, communication patterns, and resource utilization across NVIDIA platforms.Build and maintain performance and correctness regression tests to prevent slowdowns and ensure stable behavior across model and hardware configurations.Collaborate with model, platform, and SRE teams to translate production requirements into upstreamable solutions with strong operability and maintainability.What we need to see:5+ years building production software with solid systems engineering fundamentals and a track record of delivering performance or reliability improvements.Experience with LLM inference/serving stacks (e.g., vLLM, SGLang) and an understanding of the tradeoffs that drive real production performance.Strong programming skills in Python plus C++ and/or CUDA; ability to debug and optimize performance‑critical code.Experience with profiling and performance investigation (microbenchmarks, flame graphs, GPU profiling) and a measurement‑driven mindset.Familiarity with distributed systems concepts and concurrency (queues/schedulers, multi‑process/multi‑threading, scaling across GPUs/nodes).Strong communication skills and comfort working with open‑source communities (issues, PR discussions, code review).BS/MS in Computer Science, Computer Engineering, or related field (or equivalent experience).Ways to stand out from the crowd:Open‑source contributions to vLLM, SGLang, PyTorch, Triton, NCCL, Dynamo or adjacent serving/runtime projects.Shipped performance work such as improved attention/KV cache efficiency, speculative decoding, scheduler improvements, quantization-aware serving, or streaming latency reductions.Experience building reproducible benchmarking and performance regression infrastructure for latency/throughput.Systems performance background spanning memory bandwidth, kernel fusion, PCIe/NVLink effects, and network fabrics (e.g., InfiniBand).We are widely considered to be one of the technology world’s most desirable employers. We have some of the most forward‑thinking and creative people in the world working for us. If you're creative and autonomous with a real passion for technology, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until July 20, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, NY, Remote; US, CA, RemoteType: Full time
$152k - $241.5k
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize... ...accelerated software that powers today’s most sophisticated AI applications. Our team is responsible for...SeniorFull timeRemote work$152k - $241.5k
...recently, GPU deep learning ignited modern AI — the next era of computing — with... ...for an AI & Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning & AI... ...has been the backbone of NVIDIA’s inference engine, spanning across data centers...SeniorFull timeRemote work$110k - $150k
...About us Topaz Labs is a full-stack AI company that develops, trains, and deploys... ...rely on. About the role As a Software Engineer supporting our AI Engine, you would report... ...NVIDIA, AMD, Intel, Apple) to optimize inference on their hardware. About you...SuggestedFull timeWork experience placementRelocation$184k - $287.5k
...the forefront of the generative AI revolution, building the software and systems that power the world’... ...workloads. We are looking for a Senior Software Engineer to lead the bring-up, triage, benchmarking... ...of distributed training and inference workloads across NVIDIA GPU...SeniorFull timeRemote work$152k - $241.5k
...computing platforms are powering the AI revolution across many... ...and industries. Within our software stack, CUTLASS stands out as... ...the-art deep learning models’ inference and training passes to identify... ...Computer Science, Computer Engineering, or related field (or equivalent...SeniorFull time$152k - $241.5k
...real-time large language model inference? Join NVIDIA’s TensorRT Edge-LLM... ...shape the next generation of edge AI for automotive and robotics. We build the software stack that enables Large... ...Computer Science, Electrical/Computer Engineering, or a closely related field.4+ years...SeniorFull time$184k - $287.5k
We're looking for outstanding AI systems engineers to develop groundbreaking technologies in the inference systems software stack! We build innovative AI systems software to accelerate for AI inference. As a member of the team, you'll develop libraries, code generators,...SeniorFull timeRemote work$110k - $150k
...About us Topaz Labs is a full-stack AI company that develops, trains, and deploys... ...rely on. About the role As a Software Engineer supporting our AI Engine, you would report... ...NVIDIA, AMD, Intel, Apple) to optimize inference on their hardware. About you...Full timeWork experience placementRelocation$224k - $356.5k
...into the unlimited potential of AI to define the next era of... ...Local AI team is building the software stack that makes large language... ...in leading open-source LLM inference frameworks — identify performance... ...Computer Science, Computer Engineering, Electrical Engineering, or...SeniorFull timeLocal area$152k - $241.5k
...recently, GPU deep learning ignited modern AI - the new era of computing,... ...the digital landscape.Local AI seeks a Senior Systems Software Engineer interested in solving client-side AI... ...models, data processing pipelines, and inference runtime features.Identifying, evaluating...SeniorFull timeLocal area$184k - $287.5k
...for a motivated Deep Learning engineer to bring advanced CUDA... ...Distributed Runtime technologies into AI stacks, including PyTorch,... ...on scales up to 100K GPUs to inference down at microsecond latency.... ...systems principles (aka systems software fundamentals)Adaptability and...SeniorFull time$170.6k - $261.3k
Job DescriptionAs a Senior Software Engineer on the SimCore team, you will build and deploy applied AI/ML solutions that directly support simulation workflows, internal tools... ...excel at building robust, high-performance inference pipelines. This role is not focused on...SeniorFull timeLocal areaWork from homeFlexible hours$163.2k - $244.8k
...Time Employee /On-siteShield AI is a venture-backed defense-tech... ...include Hivemind autonomy software, V-BAT and X-BAT aircraft, and... ...research and production, our engineers build the data pipelines,... ...TensorRT, and hardware-accelerated inference frameworks. Perception &...SeniorFull timeTemporary workPart timeWork experience placementWorldwide$184k - $287.5k
Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing... ...pre-training, post-training, inference. Our objective is to deliver... ...seeking an AI infrastructure software engineer to join our team. You'll be... ...availability of AI systems.As a senior DGX Cloud AI Infrastructure...SeniorFull timeRemote work$152k - $241.5k
...solve some of the hardest problems in AI: storage, access, ingestion, governance... ...high-performance GPU-based training and inference jobs. Our work gives NVIDIA teams the foundational... ...and production environments.Use modern software engineering practices, including AI-assisted and...SeniorFull timeRemote work$184k - $287.5k
...experienced and highly motivated software professional to work on... ...hardware performance for emerging AI workloads. You will be a... ...bottlenecks in both training and inference pipelines.Collaborate closely... ...Computer Science, Computer Engineering, Electrical Engineering, or related...SeniorFull time- ...Senior Software Engineer: Applied AI (Voice Agents & ML Systems) The pitch We build and operate production AI voice agents that hold real phone conversations... ...: feature engineering, model training, and scheduled inference Imbalanced, messy real‑world data; calibration and...SeniorPermanent employmentCasual work
$190k - $215k
...relationships thrive, and intelligent software helps millions of people make meaningful connections every day.As a Senior Software Engineer on our AI & Intelligent Systems team, you'll... ...Engineers to productionize models, build inference services, and integrate AI...SeniorPermanent employmentLive in- ....About the RoleADT is looking for a Senior iOS Mobile Software Engineer to join our Product Engineering mobile... ...improve the app experience.Leverage AI-assisted development tools such as... ...Matter, Z-Wave Long Range, or local AI inference.Objective-C and UIKit experience....SeniorH1bLocal areaRemote work
- ...for building accurate and performant AI applications at scale in production. Pinecone... ...the Team and Role: We are hiring a senior software engineer to help design and build core... ...throughput and cost across large-scale inference and retrieval workloads Drive technical...SeniorLocal areaRemote workWork from homeFlexible hours
$130.7k - $205.2k
...low-level device platform software component installed on millions... ...telemetry, and more modern AI or subscription use cases.... ...for performing local inference and more. This role operates... ...priorities. We are looking for a Senior Software Engineering to work on the HP On Device...SeniorTemporary workWork experience placementLocal areaFlexible hours$115k - $160k
Join to apply for the Senior Software Engineer role at Spirent Communications Join to apply for the Senior... ...role at Spirent Communications Get AI-powered advice on this job and more... ...interviewing at Spirent Communications by 2x Inferred from the description for this job...SeniorFull timeWork experience placement- Get AI-powered advice on this job and more exclusive features. Direct message the... ...poster from Elios Talent Tech Lead / Senior Software Engineer - EHR Integrations (Epic & IKM) Why... ...of interviewing at Elios Talent by 2x Inferred from the description for this job Medical...SeniorFull timeFreelanceRemote work
$145k - $182k
Overview Make Your Mark: We’re looking for a Senior AI/ML Engineer to design, build, and optimize data... ..., and efficiency. Collaborate with software engineers to integrate machine... ...Implement and maintain APIs for model inference. Design and manage training infrastructure...Senior$200k
...Workforce Strategy Expert Senior Full Stack LLM Developer (MLOps & Prompt Engineering Focus) Location:... ...Time Reports to: Head of AI Team At Titan Intake,... ...Full-time Industries Software Development Hospitals... ...at Titan Intake by 2x Inferred from the description for...SeniorFull timeRemote work$114.6k - $234.6k
...Senior AI Software Engineer As a Senior AI Software Engineer in an AI Innovation organization within OCI, you will help build AI capabilities... ...OCI AI platform capabilities, including agent execution, inference systems, model serving, AI workflow orchestration, evaluation...SeniorTemporary workFlexible hours$160k - $200k
...Build, Deploy, and Maintain AI for an Unpredictable World Striveworks helps organizations... .... Founded by data scientists and engineers, Striveworks set out to make the journey... ...environments. The Role As a Senior Software Engineer at Striveworks, you'll be challenged...SeniorFull timeWork at officeRemote work- ...to help recruiters source candidates, and layering agentic AI capabilities on top of that foundation. It's an early, high... ...decisions shape the product. The Role: We're hiring one Senior Backend Engineer to own and drive meaningful initiatives end-to-end. This...Senior
$100k - $163k
About this role: Wells Fargo is seeking a Senior Software Engineer in our Platform Management Group as part of Consumer Technology. Learn more about... ...addition, you will help lead the adoption of Wells Fargo’s AI-enabled tools and capabilities, identifying innovative ways...SeniorFull timeWork experience placement$170k - $240.75k
...who they uniquely are.We are seeking a Senior Principal Software Engineer to lead the design and development of next-generation Agentic AI platforms and applications that power personalization... ..., safety, and business impact Scalable inference and orchestration infrastructure Ensure...SeniorRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Software Engineer - AI Inference. Be the first to apply!
- ngo software engineer Texas
- software data engineer Texas
- entry level software engineer remote Texas
- software developer positions Texas
- software development engineer aws Texas
- consulting software engineer Texas
- financial software developer Texas
- software developer Texas
- part time software developer Texas
- software engineer healthcare Texas




