Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior/Staff Software Engineer - LLM Inference & Reinforcement Learning Platform

$236k - $330k

Snowflake

At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done.We are looking for talented systems developers and researchers to join the Snowflake AI Research team and advance the state of the art in LLM inference systems and optimization.Our mission is to build the next generation of high-performance and intelligent inference systems. We optimize not only how fast and efficiently models run, but also how quickly inference systems can adapt to new models, architectures, hardware, and workloads.Our work spans the full inference stack—from distributed serving and runtime systems to GPU kernels and model-system co-design. We explore techniques such as adaptive parallelism, speculative and parallel decoding, disaggregated inference, scheduling and batching, KV-cache optimization, model swapping, quantization, and GPU kernel optimization to push the frontier of latency, throughput, scalability, and cost.Beyond optimizing individual models, we are building intelligent and adaptive inference systems that can automate performance optimization—rapidly profiling new models and workloads, identifying bottlenecks, selecting effective execution strategies, and adapting system configurations with minimal manual tuning. We embrace AI-native engineering, using AI not only as the workload we optimize, but also as a tool to accelerate system development, experimentation, debugging, optimization, and adaptation to new models. Our goal is to accelerate both the speed of inference and the agility of inference development.Recent innovations from Snowflake AI Research include Arctic Inference, our open-source inference system, and technologies such as Shift Parallelism, which dynamically adapts parallelism to workload characteristics; SwiftKV, which reduces redundant prefill computation; Arctic Speculator and SuffixDecoding for fast speculative decoding; Jacobi Forcing for causal parallel decoding; and Semi-Persistence for fast model swapping and dynamic multi-model serving.This is an exciting opportunity to collaborate with a world-class team, including founding members of DeepSpeed, vLLM, and TensorFlow. Together, we will push the boundaries of AI systems and bring cutting-edge research into production-scale AI.ResponsibilitiesDesign and develop high-performance LLM inference systems, spanning distributed serving, runtime systems, GPU execution, and performance-critical kernels.Develop novel techniques to improve inference latency, generation speed, throughput, memory efficiency, scalability, and cost.Explore advanced inference techniques including speculative and parallel decoding, prefill/decode disaggregation, adaptive parallelism, continuous batching and scheduling, KV-cache management, quantization, and communication optimization.Develop adaptive and intelligent inference systems that automatically optimize execution for new model architectures, hardware platforms, workload characteristics, and deployment environments.Apply AI-driven and AI-native approaches to systems engineering, including automated profiling, bottleneck identification, configuration search, code generation, experimentation, runtime strategy selection, debugging, and performance tuning.Independently identify high-impact performance and systems problems, formulate hypotheses, prototype solutions, and drive promising ideas from research through production.Design distributed inference strategies across GPUs and nodes, including tensor, sequence, pipeline, data, and expert parallelism.Develop efficient approaches for multi-model serving, dynamic resource management, model loading and swapping, and workload-aware scheduling.Analyze and optimize GPU kernels and operators for attention, MoE, communication, and other performance-critical model components.Explore model-system co-design, including model or post-training techniques that unlock substantially more efficient inference.Profile and benchmark end-to-end workloads to identify bottlenecks across compute, memory, communication, networking, scheduling, and model execution.Collaborate closely with model researchers, infrastructure teams, and product teams to deploy research innovations in production.Open-source and publish innovations through technical blogs and top-tier systems and machine learning conferences.RequirementsBachelor’s degree in Computer Science, Electrical Engineering, or a related field. A Master’s degree or PhD is preferred.5+ years of experience in one or more of the following areas: LLM inference systems, distributed AI systems, GPU systems, or high-performance computing.Strong understanding of modern LLM inference architectures and the performance tradeoffs involved in serving large-scale models.Hands-on experience with modern LLM inference and serving frameworks, such as vLLM, SGLang, TensorRT-LLM, or similar systems.Experience designing, extending, or optimizing inference runtimes, including areas such as scheduling, batching, KV-cache management, distributed execution, parallelism, speculative decoding, or disaggregated serving.Strong understanding of GPU architectures and experience with CUDA, Triton, or similar GPU programming environments.Experience with performance-oriented libraries and frameworks such as CUTLASS, cuBLAS, cuDNN, or related technologies.Experience profiling and diagnosing end-to-end system performance using Nsight Systems, Nsight Compute, or equivalent tools.Demonstrated ability to operate as an independent problem identifier and solver—recognizing important problems with limited direction, defining the right technical questions, and driving solutions through ambiguity.Strong ability to work across model, runtime, distributed system, and hardware layers and reason about end-to-end performance tradeoffs.Experience using AI-native engineering approaches to accelerate software development, experimentation, debugging, optimization, or system adaptation is a strong plus.Excellent communication skills and the ability to collaborate effectively across research, engineering, and product teams.Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.How do you want to make your impact?For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.comCompensation Range: $236K - $330KLocationUS-WA-BellevueEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensation$236K – $330K

Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Senior/Staff Software Engineer - LLM Inference & Reinforcement Learning Platform in Bellevue, WA vacancy
  •  ...Foundation Model Inference team within Cloud...  ...efficiency, hardware/software codesign, systems...  ...direction for the engineers around you. This...  ...working closely with platform and security teams...  ...experience with LLM inference stacks....  ...Science, Machine Learning, Artificial Intelligence... 
    Platform
    Senior
    Worldwide

    Apple

    Seattle, WA
    5 days ago
  • $207.48k - $368.22k

     ...believe TikTok is an ideal platform to deliver a brand new and better...  ...- Responsible for exploring reinforcement learning and operational research...  ...multi-agent tools, models, engines, and platforms to enhance automatic...  ...multimodal, search, graph, LLM, Agent etc. to provide... 
    Platform
    Senior
    Temporary work
    Local area
    Overseas

    Tik Tok

    Seattle, WA
    5 days ago
  •  ...Build an E-commerce ecosystem and improve platform service capabilitiesResponsibilities:-...  ...domain.- Responsible for exploring reinforcement learning and operational research algorithms in...  ...NLP, vision, multimodal, search, graph, LLM, etc. to provide support for governance... 
    Platform
    Senior
    Overseas

    TikTok

    Seattle, WA
    4 days ago
  •  ...the future of how work gets done.Senior Software Engineer — Cortex TrainingThe Snowflake ML Platform team's mission is to let...  ...Snowflake. Cortex Training is our LLM post-training platform: it turns...  ...orchestration, multi-node training and inference, fault tolerance, and... 
    Platform
    Senior

    Snowflake

    Bellevue, WA
    a month ago
  • $168.1k - $227.4k

     ...Alexa+, Amazon's LLM-powered...  ...dozens of chained inferences, coupled to real...  ...looking for a Senior Machine Learning Engineer to build and own...  ...in this agentic platform. You will take...  ...infrastructure, reinforcement learning training...  ...professional software development experience... 
    Platform
    Senior
    Internship
    Flexible hours
    Day shift

    Amazon

    Bellevue, WA
    3 hours ago
  • $152k - $204k

     ...pioneers, CoreWeave delivers a platform of technology, tools,...  ...: CRWV) in March 2025. Learn more at  What You'll Do: Senior engineers are area owners who lead...  ...our Kubernetes-native inference platform and meet strict...  ...(vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe).... 
    Platform
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    14 days ago
  •  ...high-performance inference infrastructure for machine-learning research and support...  ...AI accelerator platforms. Analyze observability...  ...Significant software engineering experience, particularly...  ...Familiarity with LLM inference...  ...policy requiring staff to work from an Anthropic... 
    Platform
    Senior
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    26 days ago
  •  ...TeamDoorDash’s GenAI Platform team sits...  ...Machine Learning Platform and...  ...batch inference, and fine-tuning...  ...including the LLM Gateway,...  ...and inference engines, fine-tuning...  ...ideal for a senior engineer who...  ...directions such as reinforcement learning (...  ...in software engineeringDeep... 
    Platform
    Senior
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    Seattle, WA
    4 days ago
  • $115.4k - $173k

     ...content into other game engines and 3D environments —...  ...engines and emerging software platforms. We’re also developing...  ...evaluation harnesses, reinforcement learning, or agentic workflows....  ...style tool interfaces, LLM agents, multimodal models, inference pipelines, embeddings,... 
    Platform
    Senior
    Full time
    Work at office
    Remote work
    Worldwide

    Unity Technologies

    Bellevue, WA
    a month ago
  • $175k - $308.5k

    Senior Machine Learning Engineer, Search & Knowledge Platforms Do you want to make Siri and Apple products smarter...  ...upon the latest LLM/ML techniques in order...  ...latest advances for online inference optimization. As a...  ...experience and strong software engineering skills. Responsibilities... 
    Platform
    Senior
    Relocation

    Apple

    Seattle, WA
    5 days ago
  •  ...executing this mission. The ML Platform team at Zoox plays a crucial...  ...alongside a team of strong software engineers and act as a force...  ...ML domains. If you want to learn more about our stack behind...  ...cutting-edge ML Training OR Inference performance optimization techniques... 
    Platform
    Senior

    Zoox

    Seattle, WA
    23 days ago
  •  ...Search & Knowledge Platforms team builds...  ...MLE, SWE, and data engineers responsible for delivering...  ...on-device LLM models for personal...  ...device and on-server software frameworks for context...  ...LLM-based model inference. Integrate the...  ...the full Machine Learning life cycle MS or... 
    Platform
    Senior
    Temporary work
    Worldwide

    Apple

    Seattle, WA
    5 days ago
  • $200k - $250k

     ...Metropolis is seeking a Senior Manager of Machine Learning Engineering within the Advanced Technologies...  ...annotation workflows (LLM-in-the-loop) to reduce...  ...registries, and low-latency inference services. Ensure high...  ...external vendors and annotation platform providers to ensure high-... 
    Platform
    Senior
    Temporary work
    Work at office
    Local area

    Metropolis Corp

    Seattle, WA
    1 day ago
  •  ...team within Apple's software organization is a multidisciplinary...  ...Experiences. As a senior machine learning engineer on our team, you...  ...training and inference for Apple's AI-...  ...Apple's centralized ML platform. Candidates should bring...  ...understanding of LLM architectures and... 
    Platform
    Senior
    Flexible hours

    Apple

    Seattle, WA
    5 days ago
  • $200.1k - $270.6k

     ...Alexa+, Amazon's LLM-powered...  ...dozens of chained inferences, coupled to real...  ...a Principal Engineer to lead the engineering...  ...this agentic platform. You will own...  ...concurrency), reinforcement learning training...  ...on; influence senior leadership on...  ...professional software development experienceKnowledge... 
    Platform
    Permanent employment
    Internship
    Flexible hours
    Day shift

    Amazon

    Bellevue, WA
    10 days ago
  •  ...training, evaluation, inference, and serving...  ...reasoning. Ensure platforms are secure,...  ...direction, raise engineering standards, and mentor...  ...field. ~7+ years of software engineering experience...  ...agentic systems, LLM tool-use...  ...wellness support. Learning and development programs... 
    Platform
    Senior
    Full time
    Work at office
    Remote work

    Axon

    Seattle, WA
    25 days ago
  •  ...simulation to develop our driving software, validate our safety, and...  .... ~ Best for: Architects/engineers who care about robust and elegant...  ...MA) ~ Role: Develop Python platform for large scale generation of...  ...of robotics, machine learning, and design, Zoox aims to provide... 
    Platform
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    1 day ago
  • $185k - $235k

     ...NewsBreak is the Content Intelligence platform shaping the future content economy....  ...We are looking for an experienced Senior Machine Learning Engineer to design and build the data and ML...  ...vector stores, and low-latency, high-QPS inference for real-time bidding. Experience... 
    Platform
    Senior
    Full time
    Local area
    Work from home

    NewsBreak

    Bellevue, WA
    4 days ago
  • $92k - $135k

     ...pioneers, CoreWeave delivers a platform of technology, tools,...  ...: CRWV) in March 2025. Learn more at  What You'll Do: Join the Inference team to ship production...  ...from experienced engineers. About the role: Implement...  ...Triton, vLLM, TensorRT-LLM, Ray Serve). Write... 
    Platform
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    17 days ago
  • $188k - $303k

     ...CoreWeave delivers a platform of technology,...  ...) in March 2025. Learn more at...  ...looking for a Sr. Engineering Manager to lead the...  ...next-generation Inference Platform. What...  ...will lead a team of senior and staff engineers building...  ...Triton, TensorRT-LLM, Ray Serve, or TorchServe... 
    Platform
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    17 days ago
  •  ...system intelligence. As a Machine Learning and System Optimization Engineer, you will orchestrate and allocate...  ...that allow for more efficient inference by sharing various parts of the perception...  ...technologies (e.g., TensorRT-LLM). Base Salary Range    There... 
    Senior
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    22 days ago
  •  ...simulation to develop our driving software, validate our safety, and...  ...?" \n Simulation Data Platform (Foster City, CA) ~ Role :...  ...queues ~ More Info:   Software Engineer - Simulation Data Platform...  ...of robotics, machine learning, and design, Zoox aims to provide... 
    Platform
    Senior
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    1 day ago
  • $209.1k - $282.9k

    As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction...  ...and usability of Arm’s AI platform.Responsibilities:Define and...  ...management.Strong software and production engineering...  ...vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators... 
    Platform
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    6 days ago
  • $182k - $242k

     ...CoreWeave delivers a platform of technology,...  ...CRWV) in March 2025. Learn more at  Our...  ...fraction of the total inference market, which is...  ...for strong engineers with great taste....  ...high-performance LLM tracing dashboards...  ...supervised fine-tuning, reinforcement learning, and on-policy... 
    Platform
    Senior
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    17 days ago
  •  ...infrastructure platform for...  ...do machine learning. While being...  ...that enable ML engineers and data scientists...  .... As a Staff Engineer, you...  ...product, and senior leadership to...  ...training, LLM fine-tuning,...  ...latency model inference, large-scale...  ...great software systems.Who... 
    Platform
    Flexible hours

    Stripe

    Seattle, WA
    4 days ago
  •  ...System Optimization Engineer System optimization team at Huawei...  ...analyze and optimize system software using new technologies including deep learning and reinforcement learning. Those technologies...  ...algorithms on a variety of Huawei's platforms Collaborate with academia... 
    Platform
    Work experience placement

    Netpace

    Bellevue, WA
    1 day ago
  • $228.7k - $306.7k

    Senior Principal Machine Learning Engineer, Ad Platforms Technology is at the heart of Disney's past, present, and future...  ...experience, deep technical knowledge of software and systems including Machine...  ...Model optimization and inference (TensorRT, ONNX, DeepSpeed) Ad Tech... 
    Platform
    Senior
    Work experience placement

    Disney

    Seattle, WA
    5 days ago
  • $152k - $241.5k

     ...outstanding Senior High Performance AI Engineers to build the...  ...internal NVIDIA software, model, and...  ...models and inference through...  ...systems, machine learning, compilers,...  ...agents, reinforcement learning, or...  ...accelerator platforms, and the ability...  ...with TRT‑LLM, SGLang, vLLM... 
    Platform
    Senior

    NVIDIA Corporation

    Redmond, WA
    6 days ago
  • Role: Senior AI Engineer - Privacy Location: Bellevue WA The Senior...  ...Client data privacy platform at scale. Embedded...  ...regulations. AI Agent & LLM Engineering: Design...  ...training, fine-tuning, and inference. Apply prompt engineering, few-shot learning, and fine-tuning... 
    Platform
    Senior

    Axelon

    Bellevue, WA
    5 days ago
  • $228.7k - $306.7k

    Sr Principal Machine Learning Engineer Technology is at the heart of Disney...  ...around the world. Ad Platforms is responsible for Disney's...  ...deep technical knowledge of software and systems including Machine...  ...libraries Model optimization and inference (TensorRT, ONNX, DeepSpeed)... 
    Platform
    Senior
    Work experience placement
    Local area

    The Walt Disney Studios

    Seattle, WA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior/Staff Software Engineer - LLM Inference & Reinforcement Learning Platform. Be the first to apply!