Senior/Staff System Research Engineer - LLM Inference Optimization
$236k - $309.75kStreamlit
Snowflake AI Research Team Position At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don't just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization. Our mission is to build the next generation of high-performance and intelligent inference systems. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads. Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost. Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development. Recent innovations from Snowflake AI Research include Arctic Inference, our open-source inference system, and technologies such as Shift Parallelism, which dynamically adapts parallelism to workload characteristics; SwiftKV, which reduces redundant prefill computation; Arctic Speculator and SuffixDecoding for fast speculative decoding; Jacobi Forcing for causal parallel decoding; and Semi-Persistence for fast model swapping and dynamic multi-model serving. This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI. Responsibilities Design and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels. Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost. Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization. Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments. Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning. Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production. Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism. Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling. Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components. Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference. Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution. Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production. Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences. Requirements Bachelor's degree in Computer Science, Electrical Engineering, or a related field. A Master's degree or PhD is preferred. 5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing. Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models. Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems. Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving. Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments. Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies. Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools. Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity. Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs. Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus. Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams. Snowflake is growing fast, and we're scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake. How do you want to make your impact? For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com The following represents the expected range of compensation for this role: The estimated base salary range for this role is $236,000 - $309,750. Additionally, this role is eligible to participate in Snowflake's bonus and equity plan. The successful candidate's starting salary will be determined based on permissible, non-discriminatory factors such as skills, experience, and geographic location. This role is also eligible for a competitive benefits package that includes: medical, dental, vision, life, and disability insurance; 401(k) retirement plan; flexible spending & health savings account; at least 12 paid holidays; paid time off; parental leave; employee assistance program; and other company benefits. To comply with pay transparency requirements and other statutes, you can notify us if you believe that a job posting is not compliant by completing this form. Streamlit
$232.56k - $427.5k
Research Engineer - LLM/VLM Inference Optimization (Seed Infra) Location: Seattle Team: Technology Employment Type: Regular Job Code: A236224 Responsibilities... ..., develop, and optimize high-performance inference systems for large-scale LLMs and VLMs, covering inference...SuggestedTemporary workLocal area- ByteDance in Seattle is seeking a Senior Research Engineer/Scientist for Storage for LLM to design and maintain a high-... ...performance KV cache layer for LLM inference across GPUs and nodes. This role... ...inference and serving teams, optimize eviction policies, memory management...Senior
- ...experienced Member of Technical Staff to optimize and accelerate real-time AI inference across a multimodal stack,... ..., SGLang, and TensorRT-LLM for our high-throughput... ...diffusion optimizations. This senior role emphasizes collaboration with research and infrastructure teams,...SeniorVisa sponsorship
$165k - $310k
Senior Research Engineer, LLM Training & Post-Training New York, New York, United... ...training, and deploying AI systems—designed to take ideas from... ...training, and production inference, with security, observability... ...has built, trained, and optimized modern transformer-based...SeniorFor contractorsFor subcontractorWork at officeRemote workWork from homeFlexible hours2 days per week$202.16k - $368.22k
Senior Research Engineer / Scientist - Storage for LLM Location: Seattle Team: Infrastructure Employment... ...The Infrastructure System Lab is a hybrid research... ...databases, infrastructure optimization through machine... ...key‑value stores and LLM inference KV caches. The team thrives...SeniorTemporary workLocal area- Google is seeking a senior XR System Performance Engineer to optimize the performance of XR hardware and software across the Glasses OS stack. You will identify bottlenecks and implement end-to-end optimizations under tight latency, power, and thermal constraints. You will...Senior
- ...technologies to support AI/LLM applications.- Design... ...of platforms/systems for monitoring, analysis... ...scale AI/LLM network.- Research and development of high... ...stacks, and codesign optimization of host-network-application... ...science, electronic engineering, network engineering...Senior
- Junior Imaging Systems Research Engineer Join to apply for the Junior Imaging Systems... ...and pipelines, ensuring optimal content reproduction across... ...currently Public Company. Seniority level Seniority level Entry... ...at Jobright.ai by 2x Inferred from the description for this...Full time
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference... ...tooling, and agentic optimization systems that improve GPU kernels at the assembly layer...SeniorFull time$241.68k
...The Vision-Applied Research team focuses on... ...looking for a Research Engineer / Scientist who... ...scalable training, optimization, and deployment.... ...hardware-efficient inference, and their applications... ...generative AI or LLM models using... ...information technology systems; Exercising sound...SeniorTemporary workLocal area$202.16k - $368.22k
Senior Research Engineer / Scientist - AI for Databases Location: Seattle Team... ...in databases, large‑scale systems, and artificial intelligence... ...Your work will span query optimization, indexing strategies, workload... ...skills. Familiarity with LLM, reinforcement learning,...SeniorTemporary workLocal area- ...next generation of autonomous system intelligence. As a Machine Learning and System Optimization Engineer, you will orchestrate and... ...that allow for more efficient inference by sharing various parts of the... ...technologies (e.g., TensorRT-LLM). Base Salary Range...SeniorTemporary workRelocation package
$106.9k - $200.6k
...We are seeking an AI Systems Engineer to own the delivery, model... .... Works with senior engineers to test and... ...operating high-performance inference (GPUs, model servers,... ...dashboards (Grafana), LLM debugging and evaluation... ...attribute, bound, and optimize AI consumption cost per...SeniorFull timeSummer holidayFlexible hours$200.8k - $251k
...technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This...Full time$56 - $95 per hour
...Job Description Position Title: Senior Systems Engineer - Applications Position Description:... ...resolve application issues, and mentor staff on support for DI authoring, conversion... ...platform conversion tooling, then optimize and enrich the content to fully leverage...SeniorPermanent employmentContract workWork experience placementRemote work- ...Senior Database Reliability Engineer (DBRE)Identity is the key to unlocking the potential of AI. Okta secures... ...you will design, operationalize, and optimize the data persistence layer that... ...powers our large-scale, mission-critical systems. You will work closely with SRE,...Senior
- VAST Data is looking for a Senior Systems Engineer to join our growing team! This is a great opportunity... ...data analysis and AI training and inference. Designed from the ground up to make... ...IT organizations Positive energy, optimism and vision. Ability to perform in an...SeniorTraineeship
- Heads Up Technologies seeks a Senior Systems Engineer to provide frontline technical support, troubleshoot, and engage with customers to tailor... ..., and support, enabling successful implementation and optimization of our products. The ideal candidate has extensive aerospace...Senior
$182k - $242k
...rendering, and real-time inference. Our stack is engineered for speed, scale,... ...We're looking for a Senior Engineer for... ...kernel authoring and optimization. You will write, profile... ...the critical path of LLM inference. Optimize... ...performance-critical systems. Hands-on CUDA experience...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$170k - $190k
...industry. We are looking for a skilled and customer focused Senior Systems Engineer to provide frontline technical support and troubleshooting... ...support teams to ensure the successful implementation and optimization of our products and services. Job Details Work Days and...SeniorFull timeTemporary workNight shiftWeekend work- ...continual learning capabilities across production tasks. This role demands rapid learning, strong engineering taste, and the ability to ship. You will work with a GPU-rich stack and collaborate with teams to push forward cutting-edge RL and LLM #J-18808-Ljbffr CoreWeaveSenior
$153k - $204k
...more at What You'll Do: The Systems Engineering team owns the host software stack that... ...the stack. About the role: As a Senior Software Engineer on the Systems Engineering... .... Explore AI-native testing — LLM-driven log triage, failure classification...SeniorPermanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$182k - $242k
...ll Do: CoreWeave is seeking a highly skilled and motivated Senior Systems Engineer, Legal Systems to build and scale our contract lifecycle... ...GTM, Procurement, Security, and IT to design, implement, and optimize legal systems and workflows that support a fast-growing, complex...SeniorPermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours- ...climate risk, energy systems, and global... ...The first is our inference control plane — open... ...building the fork engine, the guest agent,... ...operating open-weight LLM serving infrastructure... ...before you optimize, and you can tell... ...genuine willingness to research your way into the...SeniorFull timeRemote workWork visaFlexible hoursDay shift
- ...around foundation models for Earth observation, is seeking a Senior Research Engineer to collaborate with partners and tailor the OlmoEarth... ...handling and annotation to distributed model training and inference, within the AI for the Planet group at the Allen Institute...Senior
$174.24k - $261.36k
...a platform that is evolving quickly. We are looking for a Senior Research Engineer who can collaborate with our partners to tailor the OlmoEarth... ...acquisition, annotation, distributed model training and inference, and visualization. Our partners span some of the most...Senior- ...mission is to enable every engineering organization to build... ...production software systems. We work closely with... ...to turn advanced research into technology that improves... ...learning, optimization, GPU-accelerated algorithms... ...~ Experience with LLM training, post-training...Full time
$100k - $150k
...Machine Learning Research Engineer - Remote Bright Vision Technologies... ...advanced machine learning systems that solve high-impact... ...production-quality training and inference pipelines using modern ML... ...and accelerator resources. Optimize models for accuracy, latency...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...TeamThe Vision-Applied Research team focuses on... ...looking for a Research Engineer / Scientist who can take... ...enabling scalable training, optimization, and deployment.... ...acceleration, hardware-efficient inference, and their... ...training generative AI or LLM models using widely adopted...Senior
- TikTok’s Vision-Applied Research team seeks a Research Engineer/Scientist to design and implement efficient large-scale generative AI models, emphasizing... ...frameworks, model acceleration, and hardware-efficient inference across image/video generation and multimodal...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior/Staff System Research Engineer - LLM Inference Optimization. Be the first to apply!
- senior staff systems engineer Bellevue, WA
- engineering aide Bellevue, WA
- assistant engineer Bellevue, WA
- technology administrator Bellevue, WA
- staff engineer Bellevue, WA
- healthcare systems engineer Bellevue, WA
- mission system engineer Bellevue, WA
- senior linux systems engineer Bellevue, WA
- system engineer remote Bellevue, WA
- systems engineer Bellevue, WA




