Senior Inference Engineer, AGI
$193.3k - $261.5kAmazon
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI. This is a full-stack inference role: you will work across the entire path a modeltakes from research to production — shaping model architecture so it is servable, building thereal-time runtime that serves it within hard latency budgets, and building the offline systemsthat train and reinforce it.You will operate at the boundary of Science and Inference, taking frontier-scale speech andaudio models and making them run within real-time latency budgets on production hardware.You will co-design architectures with scientists to make them inference-friendly from inception,own the low-latency streaming serving path, and build the training and reinforcement-learninginfrastructure that closes the loop. You will have the compute, data, and runway to solveproblems that few teams in the world are positioned to tackle.As a Senior Engineer, you will own a significant area of the inference stack end to end, drive itstechnical execution, contribute to the team's roadmap, and work closely with scientists andhardware partners to ensure our models run fast enough to feel human in real time — and at acost that makes them viable at scale. You may go deep in one of the areas below whilecontributing across the others.Key job responsibilitiesModel Architecture & Inference Co-Design• Partner with research scientists to make model architectures servable from inception —surfacing the latency, memory, and cost implications of architecture choices before they arelocked in• Implement and optimize the inference path for large-scale multimodal models — attentionand KV-cache mechanisms, multimodal/autoregressive decoding, and the computeprimitives on the critical pathApply efficiency techniques across the stack — quantization (per-tensor/per-channel/per-group, INT8/FP8/BF16), speculative decoding, operator fusion, and paged KV-cache — andquantify their quality/latency trade-offs• Develop and tune high-performance kernels for critical operations where off-the-shelfimplementations leave performance on the table, integrating them into production servingwith minimal overhead• Profile end-to-end performance with tools such as Nsight Compute/Systems and rooflineanalysis to identify and eliminate bottlenecks in large-scale inference workloadsReal-Time & Interactive Runtime• Own the real-time serving path for streaming multimodal conversational AI, meeting sub-second, streaming latency budgets under concurrent session load• Build and tune continuous batching, scheduling, and preemption to balance throughputagainst per-request latency SLAs for interactive workloads• Customize production serving frameworks (e.g., vLLM, PyTorch) for real-time streaminggenerative models that fall outside standard LLM serving patterns — sustained low-latencyoutput under concurrent session load• Implement multi-GPU inference (tensor parallelism, collective communication) for latency-critical paths, and drive cost toward parity with existing production baselines• Establish latency, throughput, and cost benchmarking, and publish the operational metricsthat gate deploymentOffline Systems: Training, RL & Evaluation Infrastructure• Build and scale the offline inference systems behind post-training — high-throughput rolloutgeneration and reward-model serving for reinforcement learning (RL/RLHF/RLAIF)• Ensure train/serve consistency — that the inference path used in RL and evaluationfaithfully matches production online behavior (e.g., parity across sampling and logitprocessing)• Work with the evaluation team to enable offline inference that captures the qualitydimensions unique to real-time conversation — latency sensitivity, audio quality, andinteraction naturalnessBasic qualifications- 5+ years of non-internship professional software development experience- 5+ years of programming with at least one software programming language experience- 4+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience- Bachelor's degree in computer science or equivalent- Experience as a mentor, tech lead or leading an engineering team- 2+ years of hands-on experience optimizing inference for neural models — not just using inference frameworks, but profiling and improving them- Strong understanding of deep learning architectures (transformers, attention mechanisms, autoregressive decoding) and their application to speech/audio or other multimodal domains- Production track record delivering latency-constrained, real-time inference systems under concurrent load- Experience with GPU performance optimization — memory hierarchy, occupancy, KV-cache management, and the accelerator programming model- Demonstrated ownership of a technical area — driving execution for a workstream and collaborating effectively across scientists and engineersPreferred qualification - Experience with production LLM/multimodal serving internals (e.g., vLLM, TensorRT-LLM): scheduler, batching, block manager, sampler customization- Hands-on experience building real-time or streaming AI systems — speech, audio, or video — with hard latency budgets- Experience authoring custom GPU kernels (CUTLASS, Triton, raw CUDA/PTX), fused attention (FlashAttention-style), or quantized GEMM- Familiarity with model-compression and efficiency techniques — quantization, pruning, distillation, speculative decoding, long-context optimization- Experience building offline inference or rollout/reward-serving infrastructure for reinforcement learning or large-scale evaluation- Experience with distributed training and post-training pipelines (SFT through RL) — parallelism strategies, training stability, and multi-accelerator communication (NCCL, NVLink)- Familiarity with multiple hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how architecture choices affect inference latency, memory, and cost- Background in speech-to-speech or audio generative models (codec models, autoregressive audio generation), speech recognition, or speech synthesis- Experience shipping research to production at scale — models serving real users, not just benchmark results- Contributions to open-source inference/kernel projects (vLLM, CUTLASS, FlashAttention, TensorRT-LLM, Triton, or similar)Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Sunnyvale - 193,300.00 - 261,500.00 USD annuallyUSA, MA, Boston - 168,100.00 - 227,400.00 USD annuallyUSA, WA, Seattle - 168,100.00 - 227,400.00 USD annually
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SeniorFull time$184k - $287.5k
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful of performance analysis and optimization... ...will be doing:Implement language and multimodal model inference as part of NVIDIA Inference Microservices (NIMs).Contribute...SeniorFull time- ...Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and...Senior
- ...allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud... ...ultra high-speed inference.About The RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an expert on state-of-the...SeniorContract workShift work
- Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs...SeniorLocal area
- ...CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor...Senior
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry...SeniorFull time$184k - $356.5k
A leading technology company in California is seeking a Senior DL Algorithms Engineer to drive inference performance for Deep Learning workloads. The role involves implementing advanced model inference and collaborating with co-design teams to optimize performance across...Senior- NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal candidate... ...3+ years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks, and...Senior
- NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable inferencing software and contribute to high-performance GPU-accelerated deployments. Join a cross-disciplinary team to push...Senior
$152k - $241.5k
...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise...SeniorFull time- NVIDIA Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team that...Senior
- ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute...Senior
- ...NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines...Senior
- ...NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models...Senior
- ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware...Senior
- ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will collaborate... ..., while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The...SeniorLocal area
$195.2k - $361.2k
...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment,...SeniorFull timeInternshipLocal areaImmediate startShift work- ...infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will...SeniorFlexible hours
- ...future of AI should belong to the people it serves Role Summary Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H100 environment,...SeniorLocal areaShift work
- ...Adobe Firefly's Generative AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across... ..., Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency, and develop APIs...Senior
$184k - $287.5k
...understand the world.We are now looking for an extraordinary Senior Perception Engineer to develop and productize NVIDIA’s autonomous driving... ...ability to implement CUDA kernels as part of training or inference pipelines.Your base salary will be determined based on your...SeniorFull timeWork experience placement$184k - $287.5k
NVIDIA is seeking an NCX Senior Engineer to join our DSX team, collaborating closely with strategic customers to implement and enhance groundbreaking... ...NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed on NVIDIA...SeniorFull timeRemote work$184k - $287.5k
...known as “the AI computing company”.We're looking for a Senior Performance Compiler Engineer to join our team and work on the open-source Triton... ...impact AI applications, accelerating both training and inference. You will be immersed in a diverse, supportive environment...SeniorFull timeRemote work$152k - $241.5k
...NVIDIA Gruppe, located in Santa Clara, is seeking a talented engineer to design and optimize containerized inference for advanced AI models. You will validate and release production-grade software, collaborating closely with research and product teams. The ideal candidate...Senior$152k - $241.5k
...the AI computing company”.We are looking for versatile software engineers for our XLA team. NVIDIA is at the center for the AI... ...optimization algorithms for deep learning workloads. You will optimize inference and training performance for the JAX framework and the OpenXLA...SeniorFull timeRemote work$152k - $241.5k
NVIDIA is seeking a Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-latency AI services. You’ll lead architecture and performance work across Dynamo and open source frameworks. You’ll collaborate...SeniorFull time- ...highly skilled individual to develop and optimize containerized inference execution for their cutting-edge AI models. The role involves... ...have advanced degrees in Computer Science or Electrical Engineering, coupled with extensive experience in AI systems, software engineering...Senior
$184k - $287.5k
...perceive and understand the world.We are seeking an exceptional Senior Perception Engineer to help design and productize NVIDIA’s next-generation... ...with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated...SeniorFull time$152k - $241.5k
...impact on the worldWe are looking for a Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning... ...speech recognition, etc. Our DLC has been the backbone of NVIDIA inference engine, spanning across data centers, personal devices, automotive...SeniorFull timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Inference Engineer, AGI. Be the first to apply!
- senior lead project manager Sunnyvale, CA
- senior robotics software engineer Sunnyvale, CA
- senior devops engineer remote Sunnyvale, CA
- senior sas administrator Sunnyvale, CA
- senior IT manager Sunnyvale, CA
- senior contracts analyst Sunnyvale, CA
- sr project manager Sunnyvale, CA
- senior windows systems engineer Sunnyvale, CA
- senior manager data science Sunnyvale, CA
- senior ui ux designer Sunnyvale, CA

