Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Engineer, AGI

$193.3k - $261.5k

Amazon

We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational AI. This is a full-stack inference role: you will work across the entire path a modeltakes from research to production — shaping model architecture so it is servable, building thereal-time runtime that serves it within hard latency budgets, and building the offline systemsthat train and reinforce it.You will operate at the boundary of Science and Inference, taking frontier-scale speech andaudio models and making them run within real-time latency budgets on production hardware.You will co-design architectures with scientists to make them inference-friendly from inception,own the low-latency streaming serving path, and build the training and reinforcement-learninginfrastructure that closes the loop. You will have the compute, data, and runway to solveproblems that few teams in the world are positioned to tackle.As a Senior Engineer, you will own a significant area of the inference stack end to end, drive itstechnical execution, contribute to the team's roadmap, and work closely with scientists andhardware partners to ensure our models run fast enough to feel human in real time — and at acost that makes them viable at scale. You may go deep in one of the areas below whilecontributing across the others.Key job responsibilitiesModel Architecture & Inference Co-Design• Partner with research scientists to make model architectures servable from inception —surfacing the latency, memory, and cost implications of architecture choices before they arelocked in• Implement and optimize the inference path for large-scale multimodal models — attentionand KV-cache mechanisms, multimodal/autoregressive decoding, and the computeprimitives on the critical pathApply efficiency techniques across the stack — quantization (per-tensor/per-channel/per-group, INT8/FP8/BF16), speculative decoding, operator fusion, and paged KV-cache — andquantify their quality/latency trade-offs• Develop and tune high-performance kernels for critical operations where off-the-shelfimplementations leave performance on the table, integrating them into production servingwith minimal overhead• Profile end-to-end performance with tools such as Nsight Compute/Systems and rooflineanalysis to identify and eliminate bottlenecks in large-scale inference workloadsReal-Time & Interactive Runtime• Own the real-time serving path for streaming multimodal conversational AI, meeting sub-second, streaming latency budgets under concurrent session load• Build and tune continuous batching, scheduling, and preemption to balance throughputagainst per-request latency SLAs for interactive workloads• Customize production serving frameworks (e.g., vLLM, PyTorch) for real-time streaminggenerative models that fall outside standard LLM serving patterns — sustained low-latencyoutput under concurrent session load• Implement multi-GPU inference (tensor parallelism, collective communication) for latency-critical paths, and drive cost toward parity with existing production baselines• Establish latency, throughput, and cost benchmarking, and publish the operational metricsthat gate deploymentOffline Systems: Training, RL & Evaluation Infrastructure• Build and scale the offline inference systems behind post-training — high-throughput rolloutgeneration and reward-model serving for reinforcement learning (RL/RLHF/RLAIF)• Ensure train/serve consistency — that the inference path used in RL and evaluationfaithfully matches production online behavior (e.g., parity across sampling and logitprocessing)• Work with the evaluation team to enable offline inference that captures the qualitydimensions unique to real-time conversation — latency sensitivity, audio quality, andinteraction naturalnessBasic qualifications- 5+ years of non-internship professional software development experience- 5+ years of programming with at least one software programming language experience- 4+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience- Bachelor's degree in computer science or equivalent- Experience as a mentor, tech lead or leading an engineering team- 2+ years of hands-on experience optimizing inference for neural models — not just using inference frameworks, but profiling and improving them- Strong understanding of deep learning architectures (transformers, attention mechanisms, autoregressive decoding) and their application to speech/audio or other multimodal domains- Production track record delivering latency-constrained, real-time inference systems under concurrent load- Experience with GPU performance optimization — memory hierarchy, occupancy, KV-cache management, and the accelerator programming model- Demonstrated ownership of a technical area — driving execution for a workstream and collaborating effectively across scientists and engineersPreferred qualification - Experience with production LLM/multimodal serving internals (e.g., vLLM, TensorRT-LLM): scheduler, batching, block manager, sampler customization- Hands-on experience building real-time or streaming AI systems — speech, audio, or video — with hard latency budgets- Experience authoring custom GPU kernels (CUTLASS, Triton, raw CUDA/PTX), fused attention (FlashAttention-style), or quantized GEMM- Familiarity with model-compression and efficiency techniques — quantization, pruning, distillation, speculative decoding, long-context optimization- Experience building offline inference or rollout/reward-serving infrastructure for reinforcement learning or large-scale evaluation- Experience with distributed training and post-training pipelines (SFT through RL) — parallelism strategies, training stability, and multi-accelerator communication (NCCL, NVLink)- Familiarity with multiple hardware backends (NVIDIA GPU, AWS Neuron/Trainium, edge accelerators) and how architecture choices affect inference latency, memory, and cost- Background in speech-to-speech or audio generative models (codec models, autoregressive audio generation), speech recognition, or speech synthesis- Experience shipping research to production at scale — models serving real users, not just benchmark results- Contributions to open-source inference/kernel projects (vLLM, CUTLASS, FlashAttention, TensorRT-LLM, Triton, or similar)Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.Los Angeles County applicants: Job duties for this position include: work safely and cooperatively with other employees, supervisors, and staff; adhere to standards of excellence despite stressful conditions; communicate effectively and respectfully with employees, supervisors, and staff to ensure exceptional customer service; and follow all federal, state, and local laws and Company policies. Criminal history may have a direct, adverse, and negative relationship with some of the material job duties of this position. These include the duties and responsibilities listed above, as well as the abilities to adhere to company policies, exercise sound judgment, effectively manage stress and work safely and respectfully with others, exhibit trustworthiness and professionalism, and safeguard business operations and the Company’s reputation. Pursuant to the Los Angeles County Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at .USA, CA, Sunnyvale - 193,300.00 - 261,500.00 USD annuallyUSA, MA, Boston - 168,100.00 - 227,400.00 USD annuallyUSA, WA, Seattle - 168,100.00 - 227,400.00 USD annually

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Inference Engineer, AGI in Sunnyvale, CA vacancy
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

    We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful of performance analysis and optimization...  ...will be doing:Implement language and multimodal model inference as part of NVIDIA Inference Microservices (NIMs).Contribute... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...Join us as we shape the future of AI and beyond. Together, we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and... 
    Senior

    AMD

    Santa Clara, CA
    1 day ago
  •  ...allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud...  ...ultra high-speed inference.About The RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an expert on state-of-the... 
    Senior
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    1 day ago
  • Intel Corporation is seeking a software engineer to make models fast on hardware people own, optimizing inference engines for edge environments. You will work with llama.cpp and vLLM, tuning KV cache, batching, and quantization while reducing CPU overhead and startup costs... 
    Senior
    Local area

    Intel Corporation

    Santa Clara, CA
    3 days ago
  •  ...CoreWeave is seeking a Senior Engineer for its Benchmarking & Performance team to own kernel-level optimization for LLM inference and end-to-end model serving, focusing on CUDA kernels and throughput/latency improvements. You will lead kernel design reviews, mentor... 
    Senior

    Neura Market

    Sunnyvale, CA
    2 hours ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $356.5k

    A leading technology company in California is seeking a Senior DL Algorithms Engineer to drive inference performance for Deep Learning workloads. The role involves implementing advanced model inference and collaborating with co-design teams to optimize performance across... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    5 days ago
  • NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal candidate...  ...3+ years of experience in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks, and... 
    Senior

    NVIDIA

    Santa Clara, CA
    5 days ago
  • NVIDIA Corporation in Santa Clara, CA seeks a Senior Software Engineer to advance Deep Learning Inference within TensorRT. You will build scalable inferencing software and contribute to high-performance GPU-accelerated deployments. Join a cross-disciplinary team to push... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • NVIDIA Corporation in Santa Clara, CA is seeking outstanding AI systems engineers to develop groundbreaking inference technologies for the hardware-accelerated stack. You will create libraries, code generators, and GPU kernel innovations for LLM workloads. Join a team that... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    3 days ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer (Finance) to build agentic systems and advance AI inference workloads in real-world finance contexts. You will design and implement scalable software that drives experimental agents, optimize performance, and contribute... 
    Senior

    Nvidia Corporation in

    Santa Clara, CA
    2 hours ago
  •  ...NVIDIA is seeking a Senior Agentic AI Software Engineer to advance agentic AI systems and workloads from scalable research to production-grade solutions. You will build agentic components, analyze inference dynamics, and collaborate with teams owning evaluation pipelines... 
    Senior

    NVIDIA

    Santa Clara, CA
    2 hours ago
  •  ...NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions for LLM serving engines, and contribute to accelerators and runtimes that power large language models... 
    Senior

    NVIDIA

    Santa Clara, CA
    2 hours ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 
    Senior

    Jobleads-US

    Santa Clara, CA
    3 days ago
  •  ...NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will collaborate...  ..., while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The... 
    Senior
    Local area

    NVIDIA

    Santa Clara, CA
    2 hours ago
  • $195.2k - $361.2k

     ...future of AI should belong to the people it servesRole SummaryMake models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments — GPU/iGPUs, Vulkan backends — not datacenter H100 environment,... 
    Senior
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    1 day ago
  •  ...infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates will... 
    Senior
    Flexible hours

    RadixArk

    Palo Alto, CA
    4 days ago
  •  ...future of AI should belong to the people it serves Role Summary Make models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments - GPU/iGPUs, Vulkan backends - not datacenter H100 environment,... 
    Senior
    Local area
    Shift work

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    5 days ago
  •  ...Adobe Firefly's Generative AI Services team is seeking Senior Machine Learning Engineers to help build scalable GenAI systems powering features across...  ..., Express, Stock, and Premiere. You will design inference pipelines, optimize models for latency, and develop APIs... 
    Senior

    Adobe Inc.

    San Jose, CA
    2 hours ago
  • $184k - $287.5k

     ...understand the world.We are now looking for an extraordinary Senior Perception Engineer to develop and productize NVIDIA’s autonomous driving...  ...ability to implement CUDA kernels as part of training or inference pipelines.Your base salary will be determined based on your... 
    Senior
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

    NVIDIA is seeking an NCX Senior Engineer to join our DSX team, collaborating closely with strategic customers to implement and enhance groundbreaking...  ...NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed on NVIDIA... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...known as “the AI computing company”.We're looking for a Senior Performance Compiler Engineer to join our team and work on the open-source Triton...  ...impact AI applications, accelerating both training and inference. You will be immersed in a diverse, supportive environment... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

     ...NVIDIA Gruppe, located in Santa Clara, is seeking a talented engineer to design and optimize containerized inference for advanced AI models. You will validate and release production-grade software, collaborating closely with research and product teams. The ideal candidate... 
    Senior

    NVIDIA Gruppe

    Santa Clara, CA
    2 hours ago
  • $152k - $241.5k

     ...the AI computing company”.We are looking for versatile software engineers for our XLA team. NVIDIA is at the center for the AI...  ...optimization algorithms for deep learning workloads. You will optimize inference and training performance for the JAX framework and the OpenXLA... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    NVIDIA is seeking a Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-latency AI services. You’ll lead architecture and performance work across Dynamo and open source frameworks. You’ll collaborate... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...highly skilled individual to develop and optimize containerized inference execution for their cutting-edge AI models. The role involves...  ...have advanced degrees in Computer Science or Electrical Engineering, coupled with extensive experience in AI systems, software engineering... 
    Senior

    NVIDIA Corporation

    Santa Clara, CA
    10 hours ago
  • $184k - $287.5k

     ...perceive and understand the world.We are seeking an exceptional Senior Perception Engineer to help design and productize NVIDIA’s next-generation...  ...with CUDA development and optimizing training or inference pipelines through custom CUDA kernels or other GPU-accelerated... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    21 hours ago
  • $152k - $241.5k

     ...impact on the worldWe are looking for a Deep Learning Compiler Engineer. NVIDIA is hiring software engineers for its Deep Learning...  ...speech recognition, etc. Our DLC has been the backbone of NVIDIA inference engine, spanning across data centers, personal devices, automotive... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Engineer, AGI. Be the first to apply!