Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior/Staff System Research Engineer - LLM Inference Optimization

$236k - $330k

Snowflake

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next generation of high-performance and intelligent inference systems. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads.Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost.Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development.This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI.ResponsibilitiesDesign and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.RequirementsBachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus.Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams.Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.How do you want to make your impact?For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.comCompensation Range: $236K - $330KLocationUS-WA-BellevueEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensation$236K – $330K

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Senior/Staff System Research Engineer - LLM Inference Optimization in Bellevue, WA vacancy
  •  ...technologies to support AI/LLM applications.- Design...  ...of platforms/systems for monitoring, analysis...  ...scale AI/LLM network.- Research and development of high...  ...stacks, and codesign optimization of host-network-application...  ...science, electronic engineering, network engineering... 
    Senior

    TikTok

    Seattle, WA
    1 day ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference...  ...tooling, and agentic optimization systems that improve GPU kernels at the assembly layer... 
    Senior
    Full time

    Nvidia

    Seattle, WA
    1 day ago
  • $254k - $350k

     ...next generation of autonomous system intelligence.As a Machine Learning and System Optimization Engineer, you will orchestrate and allocate...  ...that allow for more efficient inference by sharing various parts of...  ...technologies (e.g., TensorRT-LLM).$254,000 - $350,000 a yearBase... 
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    2 days ago
  •  ...software stack for Inferentia and Trainium, delivering high-performance model inference for customer workloads on AWS. This senior software engineering role focuses on building and optimizing large-scale inference solutions within the ML Inference Applications team. The... 
    Senior

    Annapurna Labs (U.S.) Inc.

    Seattle, WA
    5 days ago
  •  ...Experiences. As a senior machine learning engineer on our team, you...  ...design software systems and algorithms...  ...scalable training and inference for Apple's AI-...  ...learning research engineers to help...  ...understanding of LLM architectures and...  ...on AI/ML-optimized runtime stacks.... 
    Senior
    Flexible hours

    Apple

    Seattle, WA
    2 days ago
  • Google is seeking a senior XR System Performance Engineer to optimize the performance of XR hardware and software across the Glasses OS stack. You will identify bottlenecks and implement end-to-end optimizations under tight latency, power, and thermal constraints. You will... 
    Senior

    Google

    Seattle, WA
    6 days ago
  •  ...TeamThe Vision-Applied Research team focuses on...  ...looking for a Research Engineer / Scientist who can take...  ...enabling scalable training, optimization, and deployment....  ...acceleration, hardware-efficient inference, and their...  ...training generative AI or LLM models using widely adopted... 
    Senior

    TikTok

    Seattle, WA
    1 day ago
  • $125.5k - $169.8k

     ...customer-obsessed Sr. Innovation Engineer to lead the application of...  ...Robotics and Automation systems in the Delivery Stations network...  ...of design & innovation, research & development work experience...  ...concepts like system architecture, optimization, system dynamics, system... 
    Senior
    Work experience placement
    Flexible hours

    Amazon

    Bellevue, WA
    2 days ago
  • $132.1k - $178.8k

     ...of complex mechatronic systems? Do you thrive on...  ...performance insights and drive engineering excellence? Does the...  ..., analytically-minded Senior Performance Engineer...  ...design andperformance optimization will have direct...  ...design & innovation, research & development work experience... 
    Senior
    Work experience placement
    Worldwide
    Flexible hours

    Amazon

    Bellevue, WA
    5 hours ago
  • Junior Imaging Systems Research Engineer Join to apply for the Junior Imaging Systems...  ...and pipelines, ensuring optimal content reproduction across...  ...currently Public Company. Seniority level Seniority level Entry...  ...at Jobright.ai by 2x Inferred from the description for this... 
    Full time

    Jobright.ai

    Redmond, WA
    2 days ago
  • $125.5k - $169.8k

     ...the notion of leading engineering projects amongst the world...  ...last mile logistics systems in the multi-national...  ...and work with senior leaders across multiple...  ...design & innovation, research & development work experience...  ...system architecture, optimization, system dynamics, system... 
    Senior
    Work experience placement
    Immediate start
    Worldwide
    Flexible hours

    Amazon

    Bellevue, WA
    5 hours ago
  • $241.68k

     ...The Vision-Applied Research team focuses on...  ...looking for a Research Engineer / Scientist who...  ...scalable training, optimization, and deployment....  ...hardware-efficient inference, and their applications...  ...generative AI or LLM models using...  ...information technology systems; Exercising sound... 
    Senior
    Temporary work
    Local area

    ByteDance

    Seattle, WA
    6 days ago
  •  ...infrastructure, memory systems, and routing...  .... Our applied research function sits at...  ...research and production engineering — investigating...  ....You Are:As a Senior Advanced AI Research...  ...selection and inference routing strategies...  ...their digital core, optimize their operations,... 
    Senior
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Relocation

    Accenture

    Seattle, WA
    2 days ago
  • $200.8k - $251k

     ...technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This... 
    Full time

    Scale AI

    Seattle, WA
    5 days ago
  •  ...modernizing transmission systems, our industry-recognized engineers and scientists have been...  ...the world. In the role of Senior Distribution Project Engineer...  ...act as a mentor to other staff members as neededPerform...  ...preparation, line optimization, preparing specifications... 
    Senior
    Contract work
    Work at office

    HDR

    Bellevue, WA
    5 hours ago
  • $200k - $287.5k

     ...of how work gets done.Senior Software Engineer — Cortex TrainingThe...  ...Cortex Training is our LLM post-training...  ...the hard distributed-systems parts, including scheduling...  ...multi-node training and inference, fault tolerance, and...  ...reliability and the researchers behind DeepSpeed. We'... 
    Senior

    Snowflake

    Bellevue, WA
    3 days ago
  • $153k - $204k

     ...more at    What You'll Do: The Systems Engineering team owns the host software stack that...  ...the stack. About the role: As a Senior Software Engineer on the Systems Engineering...  .... Explore AI-native testing — LLM-driven log triage, failure classification... 
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Live in
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    23 days ago
  • $182k - $242k

     ...ll Do: CoreWeave is seeking a highly skilled and motivated Senior Systems Engineer, Legal Systems to build and scale our contract lifecycle...  ...GTM, Procurement, Security, and IT to design, implement, and optimize legal systems and workflows that support a fast-growing, complex... 
    Senior
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    19 days ago
  • $160k - $220k

     ...mission. If you are too, let's talk.Senior Database Reliability Engineer (DBRE) Experience Level: Mid-Senior...  ...you will design, operationalize, and optimize the data persistence layer that powers...  ...our large-scale, mission-critical systems. You will work closely with SRE, Platform... 
    Senior
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    1 day ago
  • $115.4k - $173k

     ...stream Unity content into other game engines and 3D environments — in-process,...  ...and over the network.We’re building systems that allow real-time 3D experiences...  ...with MCP-style tool interfaces, LLM agents, multimodal models, inference pipelines, embeddings, vector search... 
    Senior
    Full time
    Work at office
    Remote work
    Worldwide

    Unity Technologies

    Bellevue, WA
    3 days ago
  • $170k - $190k

     ...industry. We are looking for a skilled and customer focused Senior Systems Engineer to provide frontline technical support and troubleshooting...  ...support teams to ensure the successful implementation and optimization of our products and services. Job Details Work Days and... 
    Senior
    Full time
    Temporary work
    Night shift
    Weekend work

    Heads Up Technologies, Inc.

    Kirkland, WA
    3 days ago
  •  ...climate risk, energy systems, and global...  ...The first is our inference control plane — open...  ...building the fork engine, the guest agent,...  ...operating open-weight LLM serving infrastructure...  ...before you optimize, and you can tell...  ...genuine willingness to research your way into the... 
    Senior
    Full time
    Remote work
    Work visa
    Flexible hours
    Day shift

    Azx Inc

    Seattle, WA
    1 day ago
  • $132.1k - $178.8k

     ...fulfillment centers and logistics systems across the globe? Do you want...  ...around the globe, to develop optimal solutions for the fulfillment...  ...leadership for large-scale engineering projects- Lead ergonomic...  ...years of design & innovation, research & development work experience... 
    Senior
    For contractors
    Work experience placement
    Flexible hours

    Amazon

    Bellevue, WA
    4 days ago
  • $196k - $245k

     ...developing intelligent systems that prevent fraud. These...  ...here takes both strong engineering and scientific rigor.As a Senior Machine Learning...  ...bringing relevant external research into the team's practiceCollaborate...  ...detection, causal inference, or LLM applications to risk.... 
    Senior
    Full time
    Work at office
    Worldwide
    Flexible hours
    3 days per week

    Remitly

    Seattle, WA
    4 days ago
  • $132.1k - $178.8k

     ...and solutions oriented engineers to be a part of our...  ...fulfillment and transportation systems. If the idea of using...  ...keep reading. As a Senior Innovation and Design...  ...efforts to develop optimal solutions for the transportation...  ...design & innovation, research & development work... 
    Senior
    Full time
    Temporary work
    For contractors
    Work experience placement
    Seasonal work
    Worldwide
    Flexible hours

    Amazon

    Bellevue, WA
    2 days ago
  • $132.1k - $178.8k

     ...of designing advanced system integrations that leverage...  ...Innovation & Design Engineer to our Advanced...  ...verticals, with regular senior leadership engagement....  ...simulation and modeling to optimize space utilization,...  ...design & innovation, research & development work experience... 
    Senior
    Full time
    Temporary work
    Work experience placement
    Seasonal work
    Worldwide
    Flexible hours
    Day shift

    Amazon

    Bellevue, WA
    2 days ago
  • $132.1k - $178.8k

     ...Logistics - Join World Wide Design Engineering!Imagine designing the next...  ...complex fulfillment center systems, from concept development...  ...rigorous technical analysis to optimize facility layouts and material...  ...years of design & innovation, research & development work experience... 
    Senior
    Full time
    Temporary work
    Work experience placement
    Seasonal work
    Work at office
    Worldwide
    Flexible hours

    Amazon

    Bellevue, WA
    5 hours ago
  • $168.1k - $227.4k

     ...This role is for a senior software engineer in the Machine Learning Inference Applications team....  ...and performance optimization of core building blocks of LLM Inference - Attention...  ...adapting latest research in LLM optimization...  ...new and existing systems experience- Experience... 
    Senior
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    5 hours ago
  • $174.24k - $261.36k

     ...a platform that is evolving quickly. We are looking for a Senior Research Engineer who can collaborate with our partners to tailor the OlmoEarth...  ...acquisition, annotation, distributed model training and inference, and visualization. Our partners span some of the most... 
    Senior

    The Allen Institute for Artificial Intelligence

    Seattle, WA
    2 days ago
  •  ...mission is to enable every engineering organization to build...  ...production software systems. We work closely with...  ...to turn advanced research into technology that improves...  ...learning, optimization, GPU-accelerated algorithms...  ...~ Experience with LLM training, post-training... 
    Full time

    International Recruiting LLC

    Bellevue, WA
    21 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior/Staff System Research Engineer - LLM Inference Optimization. Be the first to apply!