Senior/Staff System Research Engineer - LLM Inference Optimization
$236k - $330kSnowflake
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next generation of high-performance and intelligent inference systems. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads.Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost.Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development.This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI.ResponsibilitiesDesign and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.RequirementsBachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus.Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams.Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.How do you want to make your impact?For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.comCompensation Range: $236K - $330KLocationUS-WA-Bellevue; US-CA-Menlo ParkEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensation$236K – $330K
- ...We’re looking for a motivated LLM Systems Engineer willing to explore new and unconventional inference systems based on emerging... ...role is part engineering, part research – you’ll be responsible for searching... ...• Prototype and optimize emerging ML inference systems...SuggestedVisa sponsorshipRelocation package
- ...TeamThe Vision-Applied Research team focuses on... ...looking for a Research Engineer / Scientist who can take... ...enabling scalable training, optimization, and deployment.... ...acceleration, hardware-efficient inference, and their... ...training generative AI or LLM models using widely adopted...Senior
$125.5k - $169.8k
...customer-obsessed Sr. Innovation Engineer to lead the application of... ...Robotics and Automation systems in the Delivery Stations network... ...of design & innovation, research & development work experience... ...concepts like system architecture, optimization, system dynamics, system...SeniorWork experience placementFlexible hours- ...next generation of autonomous system intelligence. As a Machine Learning and System Optimization Engineer, you will orchestrate and... ...that allow for more efficient inference by sharing various parts of the... ...technologies (e.g., TensorRT-LLM). Base Salary Range...SeniorTemporary workRelocation package
$182k - $242k
...rendering, and real-time inference. Our stack is engineered for speed, scale,... ...We're looking for a Senior Engineer for... ...kernel authoring and optimization. You will write, profile... ...the critical path of LLM inference. Optimize... ...-critical systems. ~ Hands-on CUDA experience...SeniorPermanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$157.1k - $212.6k
...innovative, hands-on, and customer-obsessed Senior Hardware Development Engineer to lead the development of new... ...new ideas and concepts to optimize current workflows and automate manual... ...validation of components, subsystems and systems experience- Experience using one or...SeniorFlexible hours$132.1k - $178.8k
...of complex mechatronic systems? Do you thrive on... ...performance insights and drive engineering excellence? Does the... ..., analytically-minded Senior Performance Engineer... ...design andperformance optimization will have direct... ...design & innovation, research & development work experience...SeniorWork experience placementWorldwideFlexible hours$117k - $161k
...all in on this mission. If you are too, let's talk. The Senior Systems Engineer Opportunity As a Senior Systems Engineer at Okta, you will... ...’ll Be Doing Infrastructure Management: Administer and optimize our core productivity suite, including Google Workspace, Gemini...SeniorPermanent employmentLocal areaWorldwideFlexible hours$153k - $204k
...Senior Systems Engineer, Test Frameworks & Validation Platform Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA CoreWeave is The... ...run into a root cause quickly. Explore AI-native testing — LLM-driven log triage, failure classification, and regression...SeniorPermanent employmentFull timeTemporary workCasual workLive inWork at officeFlexible hours$119.8k - $234.7k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind... .... Microsoft's Hardware Systems organization is developing AI... ...industry-leading AI training and inference. The Platform Systems... ..., performance tuning, or optimization. Experience with high-speed...SeniorOngoing contractWork at officeLocal areaFlexible hours$170k - $190k
...industry. We are looking for a skilled and customer focused Senior Systems Engineer to provide frontline technical support and troubleshooting... ...support teams to ensure the successful implementation and optimization of our products and services. Key Details • Work...SeniorFull timeTemporary workNight shiftWeekend work$56 - $95 per hour
...Job Description Position Title: Senior Systems Engineer - Applications Position Description:... ...resolve application issues, and mentor staff on support for DI authoring, conversion... ...platform conversion tooling, then optimize and enrich the content to fully leverage...SeniorPermanent employmentContract workWork experience placementRemote work- ...Research Engineer Sesame believes in a future where computers are lifelike - with the ability... ...Squeeze silicon — scale training and inference for LLM-class workloads; chase latency,... ...production—especially user-facing, online ML systems—despite shifting requirements and...Full timeContract workFlexible hoursShift work
$200k - $287.5k
...of how work gets done.Senior Software Engineer — Cortex TrainingThe... ...Cortex Training is our LLM post-training... ...the hard distributed-systems parts, including scheduling... ...multi-node training and inference, fault tolerance, and... ...reliability and the researchers behind DeepSpeed. We'...Senior$182k - $242k
...ll Do: CoreWeave is seeking a highly skilled and motivated Senior Systems Engineer, Legal Systems to build and scale our contract lifecycle... ...GTM, Procurement, Security, and IT to design, implement, and optimize legal systems and workflows that support a fast-growing, complex...SeniorPermanent employmentFull timeContract workTemporary workCasual workWork at officeFlexible hours$160k - $220k
...mission. If you are too, let's talk. Senior Database Reliability Engineer (DBRE) Experience Level:... ...will design, operationalize, and optimize the data persistence layer that powers... ...powers our large-scale, mission-critical systems. You will work closely with SRE,...SeniorPermanent employmentWork at officeLocal areaWorldwideFlexible hours- ...mission is to enable every engineering organization to build... ...production software systems. We work closely with... ...to turn advanced research into technology that improves... ...learning, optimization, GPU-accelerated algorithms... ...~ Experience with LLM training, post-training...Full time
$196k - $245k
...developing intelligent systems that prevent fraud. These... ...here takes both strong engineering and scientific rigor.As a Senior Machine Learning... ...bringing relevant external research into the team's practiceCollaborate... ...detection, causal inference, or LLM applications to risk....SeniorFull timeWork at officeWorldwideFlexible hours$61k - $101k
...certification in software engineering concepts, along... ...-on experience in system design,... ...architecture, training, and inference. We require... ...experience coaching senior engineers and... ...high-performance LLM inference platform... ...and GPU/CPU serving optimization for low-latency, high...SeniorFull timeFor contractors- ...technology products. As a Senior Lead Software Engineer at JPMorgan Chase... ...cloud platforms optimized for AI/ML workloads.... ...-on experience in system design, application... ...architecture, ML training, and inference. Experience with... ...a high-performance LLM inference platform -...SeniorFor contractors
$132.1k - $178.8k
...fulfillment centers and logistics systems across the globe? Do you want... ...around the globe, to develop optimal solutions for the fulfillment... ...leadership for large-scale engineering projects- Lead ergonomic... ...years of design & innovation, research & development work experience...SeniorFor contractorsWork experience placementFlexible hours$50 per hour
...Local hosts to ensure high-availability and optimal performance. Execute and monitor... ...and access controls. Collaborate with Engineering teams to implement new features and functionality... ...of complex environments. Seniority level ~ Mid-Senior level...SeniorFull timeLocal area$132.1k - $178.8k
...of designing advanced system integrations that leverage... ...Innovation & Design Engineer to our Advanced... ...verticals, with regular senior leadership engagement.... ...simulation and modeling to optimize space utilization,... ...design & innovation, research & development work experience...SeniorFull timeTemporary workWork experience placementSeasonal workWorldwideFlexible hoursDay shift$132.1k - $178.8k
...and solutions oriented engineers to be a part of our... ...fulfillment and transportation systems. If the idea of using... ...keep reading. As a Senior Innovation and Design... ...efforts to develop optimal solutions for the transportation... ...of Design/Innovation, research & development,...SeniorFull timeTemporary workFor contractorsSeasonal workWorldwideFlexible hours$52 - $85 per hour
...Senior Fuel Handling & Dry Storage System Engineer Company: System One Location: (Remote/Offsite) Position Type: Long-Term Contract Pay Rate: $52.00 - $85.00/hour Schedule: 40 hours/week Position Overview System One is seeking an experienced...SeniorLong term contractTemporary workFor contractorsLocal areaRemote work- ...developing medical imaging systems from concept to launch through... ...needs Support sustaining engineering initiatives for strategic products... ...and written reports to Senior Management Facilitate engineering... ...14971) Team Leadership ATS Optimization Keywords Hard Skills R&D...SeniorContract work
$152k - $241.5k
...looking for an outstanding Senior Compiler Engineer to help build the next generation... ...of compilers, agentic systems, numerical correctness, and... ...reason about, generate, optimize, and validate code transformations... ...stack, spanning models and inference, compiler pipelines,...SeniorFull timeRemote work$104.8k - $149.3k
...working with? Wabtec's Train Handling solutions include Trip Optimizer® and Locotrol®, which work together to optimize train... ...the ability to operate longer and heavier trains. As a Senior Systems Engineer , you will serve as a technical leader across engineering,...SeniorFull timeWork experience placementWork at officeWorldwideRelocation package$180k - $240k
...Senior Systems Engineer We are looking for a Senior Systems Developer to lead the design, development, and operation of large-scale distributed... ...Deep expertise in Linux systems and performance optimization Strong background in distributed systems architecture,...SeniorFlexible hours$55 - $75 per hour
...Prime Team Partners Senior IT Recruiter @ Prime... ...network, and storage systems with high availability... ...teams and vendors. Optimize system performance and... ...experience in systems engineering and IT infrastructure.... ...Team Partners by 2x Inferred from the description for...SeniorContract workLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior/Staff System Research Engineer - LLM Inference Optimization. Be the first to apply!
- technology administrator Bellevue, WA
- assistant engineer Bellevue, WA
- staff engineer Bellevue, WA
- senior staff systems engineer Bellevue, WA
- engineering aide Bellevue, WA
- distributed systems engineer Bellevue, WA
- sr systems engineer Bellevue, WA
- system engineer contract Bellevue, WA
- senior linux systems engineer Bellevue, WA
- mission system engineer Bellevue, WA



