Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior MTS - Real-Time LLM Inference & Optimization

Nuance Labs

Nuance Labs in Seattle is seeking an experienced Member of Technical Staff to optimize and accelerate real-time AI inference across a multimodal stack, from model weights to serving infrastructure. Own end-to-end latency improvements, tune KV caches, and evaluate frameworks like vLLM, SGLang, and TensorRT-LLM for our high-throughput workloads; apply quantization and diffusion optimizations. This senior role emphasizes collaboration with research and infrastructure teams, visa sponsorship on day #J-18808-Ljbffr Nuance Labs

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior MTS - Real-Time LLM Inference & Optimization in Seattle, WA vacancy
  • $236k - $309.75k

     ...Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization. Our mission is to build the next generation of high-...  ...health savings account; at least 12 paid holidays; paid time off; parental leave; employee assistance program; and... 
    Senior
    Flexible hours
    Shift work

    Streamlit

    Bellevue, WA
    2 hours ago
  • $300k

     ...significant experience in RL and post-training methods, ensuring the development of systems that achieve real-time, multimodal interaction. The candidate will build and optimize training systems, implementing modern methods and aligning research with production needs. A focus... 
    Senior

    Nuance Labs

    Seattle, WA
    2 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance...  ...kernel implementations at real-silicon fidelity across the...  ..., TX, Austin; US, NY, New York; US, WA, SeattleType: Full time
    Senior
    Full time

    Nvidia

    Seattle, WA
    1 day ago
  •  ...system intelligence. As a Model Optimization & Deployment Engineer, you will...  ...kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices...  ...technologies (e.g., TensorRT-LLM). Base Salary Range   There are... 
    Senior
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    20 days ago
  •  ...with category-leaders in real estate (CBRE), energy (...  .... The first is our inference control plane — open-weight...  ...operating open-weight LLM serving infrastructure...  ...you measure before you optimize, and you can tell the story...  ...Flexible paid time off Equity Fully... 
    Senior
    Full time
    Remote work
    Work visa
    Flexible hours
    Day shift

    Azx Inc

    Seattle, WA
    25 days ago
  • $202.16k - $368.22k

    Senior Research Engineer / Scientist - Storage for LLM Location: Seattle Team: Infrastructure...  ..., infrastructure optimization through machine...  ...value stores and LLM inference KV caches. The...  ...systems problems with real‑world impact,...  ...costs and response time for inference workloads... 
    Senior
    Temporary work
    Local area

    ByteDance

    Seattle, WA
    2 days ago
  • Blue Origin Engines seeks an Aerospace Software Apps Engineer III to develop and verify real-time embedded software that controls rocket engines for human-capable spacecraft. You will work with multidisciplinary teams across the full software lifecycle, from requirements... 
    Senior

    Blue Origin LLC

    Seattle, WA
    4 days ago
  • Disney Entertainment and ESPN Product & Technology is hiring a Senior Data Engineer to design, build, and scale data foundations powering...  ...systems, partnering with AI Core Engineering for robust, real‑time data delivery. You will mentor engineers, drive end‑to‑end data... 
    Senior

    The Walt Disney Company

    Seattle, WA
    2 days ago
  •  ...software and tools for spaceflight systems, spanning requirements, design, implementation, integration, and testing of safety-critical real-time software. The role emphasizes collaboration with avionics and ground systems teams, experience in C/C++, C/C++, Python, and... 
    Senior

    Blue Origin LLC

    Seattle, WA
    1 day ago
  • Adobe is seeking a senior, hands‑on engineer to own and evolve the cross‑platform GPU rendering platform at the heart of Premiere and...  ...with CUDA, DirectX, Metal, Vulkan, and shader tooling to advance real‑time video processing and rendering pipelines, partnering with... 
    Senior

    Adobe

    Seattle, WA
    1 day ago
  • $182k - $242k

     ...visual effects, rendering, and real-time inference. Our stack is engineered for...  .... We're looking for a Senior Engineer for CoreWeave's Benchmarking...  ...on kernel authoring and optimization. You will write, profile,...  ...epilogues-on the critical path of LLM inference. Optimize for the... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    12 hours ago
  • $148.01k - $207.22k

     ...critical software with a focus on safety and quality. The ideal candidate has a B.S. in a related field and 5+ years of experience with real-time embedded systems. Benefits include medical, dental, wellness programs, and up to 14 paid holidays. The compensation range for WA... 
    Senior

    jobs.frontdoordefense.com - Jobboard

    Seattle, WA
    12 hours ago
  • Studio Wildcard in Seattle, WA is seeking a Senior VFX Artist to create and implement high-quality real-time VFX for ARK Survival Ascended. You will work with designers, engineers, and artists to shape visuals and performance for next-gen gameplay. The role requires strong... 
    Senior

    Studio Wildcard

    Bellevue, WA
    3 days ago
  •  ...Platform is seeking an engineer to design and develop scalable software applications, data pipelines, and database solutions in a real-time, streaming environment. You will apply scientific principles to deliver data solutions that support Nordstrom business users and leadership... 
    Senior

    Nordstrom

    Seattle, WA
    2 days ago
  • $165k - $310k

    Senior Research Engineer, LLM Training & Post-Training New York, New York...  ..., and production inference, with security, observability...  ...built, trained, and optimized modern transformer-...  ...AI’s platform and real‑world customer workloads...  ...evolve over time. Master’s degree, PhD... 
    Senior
    For contractors
    For subcontractor
    Work at office
    Remote work
    Work from home
    Flexible hours
    2 days per week

    Lightning AI

    Seattle, WA
    2 days ago
  • ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-...  ...performance KV cache layer for LLM inference across GPUs and nodes. This role...  ...inference and serving teams, optimize eviction policies, memory management... 
    Senior

    ByteDance

    Seattle, WA
    3 days ago
  •  ...strong proficiency in Python, Java, or Scala, and extensive hands-on experience with Apache Spark. Key responsibilities include optimizing data workflows and supporting data scientists' needs. This is an excellent opportunity for those passionate about large-scale data... 
    Senior

    3M Consultancy

    Seattle, WA
    2 hours ago
  • $190k - $225k

     ...voice applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio...  ...Figure AI, and CallRail, the company has real scale - processing over 2 million hours...  ...speech-to-text | Speech Understanding | LLM Gateway Try the Playground Our $50M... 
    Senior

    AssemblyAI

    Seattle, WA
    9 days ago
  • $205.9k - $407.5k

     ...! We are looking to bring on a Senior Principal ML GPU Architect to lead the ML GPU optimization team in Adobe Firefly, reporting...  ...changes in training and inference speed/scale for all our ML workloads...  ...will not only enable you to make real world impact by optimizing ML... 
    Senior

    Adobe

    Seattle, WA
    2 days ago
  •  ...robotaxis on public roads, and it is a great time to join Zoox and have a significant...  ...you excited to drive our ML Performance Optimization initiatives and make our ML models that...  ...operation of cutting-edge ML Training OR Inference performance optimization techniques to scale... 
    Senior

    Zoox

    Seattle, WA
    a month ago
  • Job Overview Job Title: Senior Consultant (Full-Time) Location: Seattle, WA Work Type: Hybrid: In office...  ...Control, and Production Orders, optimized for automation and scalability. Lead...  ...and hypercare, resolving SCM issues in real-time across global teams and time zones... 
    Senior
    Full time
    Work at office
    Remote work

    ADN Group

    Seattle, WA
    4 days ago
  • $175k - $200k

     ...Senior Machine Learning Engineer Truveta...  ...reason, and accelerate real-world impact....  ...applied AI, model optimization, and agentic intelligence...  ...understanding of LLM fundamentals —...  ..., and inference performance....  ...Now is the perfect time to join Truveta. We... 
    Senior
    Full time
    For contractors
    Visa sponsorship
    Work visa
    Flexible hours

    Truveta

    Seattle, WA
    24 days ago
  • Senior Lead Software Engineer Be an integral part of an agile team...  ..., scalable cloud platforms optimized for AI/ML workloads. Partner...  ...architecture, ML training, and inference. Experience with Infrastructure...  ...scaling a high-performance LLM inference platform—leveraging... 
    Senior
    For contractors

    Hackajob

    Seattle, WA
    12 hours ago
  • $61k - $101k

     ...transformer architecture, training, and inference. We require experience with...  ...We value experience coaching senior engineers and leads on...  ...and scaling a high-performance LLM inference platform using vLLM and GPU/CPU serving optimization for low-latency, high-throughput... 
    Senior
    Full time
    For contractors

    J.P. Morgan

    Seattle, WA
    4 days ago
  •  ...Engineer to serve as the senior technical authority...  ...large-scale training and inference systems, own model...  ...production at scale, optimizing for latency, cost, and...  ...in deep learning and LLM systems, and translate...  ...in feature stores and real-time inference systems Track... 
    Full time

    CONFIDENTIAL Scovai

    Seattle, WA
    6 days ago
  • $146.88k - $220.32k

     ...operational decisions. By integrating real-time data from hardware sensors,...  ...; distributed training and inference; and a viewer so the outputs...  ...Hands-on experience shipping LLM agents to production,...  ...when their work/life balance is optimized. While we value powerful research... 
    Senior
    Full time
    Contract work
    Work at office
    Local area
    Flexible hours
    Weekend work

    The Allen Institute For Artificial Intelligence

    Seattle, WA
    a month ago
  • $135.2k - $306.4k

     ...Job Description The Senior Principal AI Agent / ML...  ...autonomous workflows, scalable inference infrastructure, and...  ...distributed services optimized for low latency, high...  ...Deep understanding of LLM application patterns, including...  ...match 8. Paid time off: Flexible Vacation... 
    Senior
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  • $199.2k - $269.5k

     ...tokens, model choice, inference patterns, and increasingly...  ...the cost, usage, and optimization insights that power financial...  ...and storage today. As Senior Manager, Product...  ...to understand where the real cost drivers and governance...  ...401(k) matching, paid time off, and parental leave... 
    Senior
    Flexible hours

    Socket.dev

    Seattle, WA
    3 days ago
  • BlackSky seeks a Senior Image Processing Engineer to design, build, and maintain the image processing pipeline that transforms raw pixels into image products. You will work on image-to-image registration, photogrammetry, feature extraction, and image-quality enhancement... 
    Senior

    Blacksky Holdings LLC

    Seattle, WA
    1 day ago
  • $152k - $204k

     ...at  What You'll Do: Senior engineers are area...  ...our Kubernetes-native inference platform and meet strict...  ...Implement advanced optimizations (e.g., micro-batch schedulers...  ...vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe)...  ...based offerings for full-time employees; for roles in... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    21 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior MTS - Real-Time LLM Inference & Optimization. Be the first to apply!