Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior SDE: Neuron Inference for LLM Acceleration

Jobleads-US

Amazon AWS Neuron is seeking a senior software engineer for the Machine Learning Inference Applications team in Seattle. You will develop and optimize core building blocks of LLM inference on Neuron chips, including attention, MLP, quantization, speculative decoding, and Mixture of Experts.

Responsibilities include applying the latest research in LLM optimization to extract top performance from open source and internal models, and working across teams to deliver scalable, high-performance

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior SDE: Neuron Inference for LLM Acceleration in Seattle, WA vacancy
  • $236k - $330k

     ...trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in...  ...Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next generation... 
    Senior

    Snowflake

    Bellevue, WA
    4 days ago
  • $168.1k - $227.4k

    AWS Neuron is the complete software stack for AWS Inferentia...  ..., AWS purpose-built accelerators for cloud-scale machine learning. This senior software engineering...  ...the Machine Learning Inference Applications team and focuses...  ...deliver best-in-class LLM serving performance.The... 
    Senior
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  •  ...Snowflake is building the next generation of AI-native inference systems. We seek experienced researchers and systems developers to advance LLM inference, optimization, and platform-wide performance across GPUs, CPUs, and distributed runtimes. You will collaborate... 
    Senior

    Jobleads-US

    Bellevue, WA
    2 days ago
  • Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product... 
    Senior

    Snowflake Computing

    Bellevue, WA
    5 days ago
  • $200k - $287.5k

     ...to how you solve problems and accelerate your impact. We look for low-...  ...future of how work gets done.Senior Software Engineer — Cortex...  ...Snowflake. Cortex Training is our LLM post-training platform: it...  ...orchestration, multi-node training and inference, fault tolerance, and... 
    Senior

    Snowflake

    Bellevue, WA
    2 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference...  ...three interconnected systems, all aimed at accelerating NVIDIA's LLM inference stack. The first is... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    4 days ago
  • $184k - $287.5k

     ...understanding performance aspects related to tasks like large scale LLM training and inference.Conducting regular technical customer meetings for...  ....Understanding of systems architecture including AI accelerators and networking as it relates to the performance of an overall... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    4 days ago
  •  ...notch technology products.As a Senior Lead Software Engineer at...  ...code review/refactoring, test acceleration, release readiness, incident/...  ...architecture, ML training, and inference.Experience with...  ...and scaling a high-performance LLM inference platform—leveraging... 
    Senior
    For contractors

    JP Morgan Chase

    Seattle, WA
    4 days ago
  • $168.1k - $227.4k

    Shape the Future of AI Accelerators at AWS NeuronWe build Amazon Neuron, the software development kit...  ...and Trainium.As a Senior Software Engineer on our...  ...in building distributed inference support for Pytorch in the...  ...of Machine Learning and LLM fundamentals, including transformer... 
    Senior
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $262.7k - $355.4k

    As a Principal Engineer on Arm’s AI Inference Cloud team, you will shape the technical direction...  ...AI infrastructure, model serving, or accelerator-backed workloads.Familiarity with...  ...PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and... 
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    5 days ago
  •  ...Amazon Inc. in Seattle, WA is seeking a Sr. Software Development Engineer for the Inference Team on AWS Neuron to deliver high-performance model inference for customer workloads on Inferentia- and Trainium-powered instances. You will lead customization of open-source... 
    Senior

    Jobleads-US

    Seattle, WA
    2 days ago
  •  ...About the Role We are looking for a Senior Data Scientist - Experimentation & Causal Inference to serve as the statistical brain behind our rapidly growing...  ...assistants to automate repetitive analytical tasks, accelerate code development, and build self-serve tooling... 
    Senior
    Full time

    ByLabs

    Seattle, WA
    4 days ago
  • $208k - $327.75k

    Running GPU-accelerated Kubernetes reliably is deceptively hard. A small change to a driver, kernel, operator, or Kubernetes version can...  ...across accelerators, clouds, and workloads, from training and inference to agentic AI.Keep the stack current and trustworthy, turning... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    4 days ago
  • $184k - $287.5k

     ...transforming computer graphics, PC gaming, and accelerated computing for more than 25 years, driven...  ...of lives. We’re searching for a Senior Systems Software Engineer with deep expertise...  ...operators, device plugins, distributed inference serving, and major cloud platforms. You’... 
    Senior
    Full time
    Remote work

    Nvidia

    Seattle, WA
    1 day ago
  • $152.2k - $205.9k

     ...largest foundation models to enterprises deploying inference at scale — rely on Amazon EC2 accelerated instances to deliver breakthrough performance for their...  ...to market at unprecedented scale.We are seeking a **Senior Product Manager - Technical** who will own and drive... 
    Senior
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $151.2k - $204.6k

     ...Books, Prime Video, and Amazon Music. As a Senior Product Manager, Technical, you will help...  ..., text-to-SQL, agentic workflows, LLM evaluation and guardrails) who is equally...  ...building per client.About the teamDigital Acceleration’s (DA) vision is “to be a complete e-commerce... 
    Senior
    Immediate start
    Flexible hours
    Shift work

    Amazon

    Bellevue, WA
    1 day ago
  • $168.1k - $227.4k

     ...solutions at incredible scale. As a Sr. SDE on this team, you will own defining, implementing...  ...Learning workloads across HPC and Accelerated Platforms in EC2 Nitro. We are seeking a...  ..., including architecture, training/inference lifecycles, and optimization of model execution... 
    Senior
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    16 hours ago
  • $135.2k - $306.4k

     ...large-scale customer and organizational impact.Personal Growth: Accelerate your career at the intersection of technical and strategic...  ...complex solutions for OCI’s networking and storage needs, mentor senior engineers, drive strategy, and be responsible for cross-team alignment... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  • $168.1k - $227.4k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...cloud-scale machine learning accelerators and the Trn1 and Inf1 servers...  ...executive leadership and other senior management and technical...  ...scale distributed training and inference solutions. This organization... 
    Senior
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $190k - $225k

     ...the next generation of voice applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio daily, and power 2...  ...Speech-to-text | Streaming speech-to-text | Speech Understanding | LLM Gateway Try the Playground Our $50M Series C fundraise Check us... 
    Senior

    AssemblyAI, Inc.

    Seattle, WA
    1 day ago
  • $262.7k - $355.4k

    As a Principal Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed...  ...execution paths.Experience optimizing kernels using accelerator programming tools, assembly, or intrinsics, including attention... 
    Work at office
    Local area

    ARM

    Seattle, WA
    5 days ago
  • $320k

     ...beneficial AI systems. About the role The Cloud Inference team scales and optimizes Claude to...  ...and cheap enough to run on the same accelerators that serve customers, trustworthy enough...  ...qualifications Have a strong interest in LLM serving; prior inference or ML... 
    Senior
    Visa sponsorship

    United States Digital Space LLC

    Seattle, WA
    3 days ago
  •  ...Adobe is seeking a Senior Staff Machine Learning Engineer for its GenAI Services team. You will design scalable inference pipelines, optimize models for latency, and build APIs that integrate diverse generative models into the Adobe product suite including Firefly, Photoshop... 
    Senior

    Jobleads-US

    Seattle, WA
    6 days ago
  • $188k - $303k

     ...deep technical expertise to accelerate breakthroughs and turn compute...  ...team for our next-generation Inference Platform. What You'll Do:...  ...team, you will lead a team of senior and staff engineers building...  ...such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    11 days ago
  •  ...networking at massive scale.Responsibilities:- Design, implementation and deployment of high-speed network technologies to support AI/LLM applications.- Design and development of platforms/systems for monitoring, analysis and diagnosis of large scale AI/LLM network.-... 
    Senior

    TikTok

    Seattle, WA
    5 days ago
  •  ...ideal candidate will have over 4 years of experience in product management, leading cross-functional teams, and leveraging advanced LLM concepts. Benefits include competitive compensation, medical premium coverage, and unlimited PTO. An AI-native environment encourages... 
    Senior

    Sunbound

    Seattle, WA
    1 day ago
  • NVIDIA Corporation is seeking a Senior Research Engineer with a passion for Generative AI inference. You will help shape how AI is infused into products by developing optimized inferencing technologies and collaborating across research, engineering, and open-source communities... 
    Senior

    NVIDIA Corporation

    Seattle, WA
    6 days ago
  • Metropolis Technologies is seeking a Senior AI Engineer to join our Applied AI organization. You will own end-to-end AI-powered systems...  ...scalable AI solutions. You will build robust data pipelines, maintain LLM-based architectures, and ensure reliable, high-impact AI outputs... 
    Senior

    Metropolis Corp

    Seattle, WA
    1 day ago
  •  ...of concept to robust, well-documented APIs serving production traffic at scale. You’ll design and ship customer-facing APIs, scale inference infrastructure for 1M+ users, and collaborate with researchers, applied engineers, and customers to #J-18808-Ljbffr AssemblyAI,... 
    Senior

    AssemblyAI, Inc.

    Seattle, WA
    3 days ago
  • $197.9k - $267.8k

     ...recommendation systems, real-time inference, large-scale distributed training, LLM infrastructure, and partner...  ...-facing outcomes* Influence senior leadership and cross-...  ...LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware- Familiarity with recommendation... 
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior SDE: Neuron Inference for LLM Acceleration. Be the first to apply!