Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer- Inference Performance

Jobleads-US

ABOUT BASETEN

Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting‑edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.

THE ROLE

We're looking for inference performance engineers who want to make the world's most demanding AI workloads run faster and more efficiently. You'll work across the stack, from the inference engine and runtime through scheduling, serving, and routing. Along the way you'll apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management. You'll reason from first principles about where time and memory go, find what's holding performance back, and close the gap. Your work directly impacts how fast our customers' models run and how efficiently we serve them. This role is ideal for someone who thrives in a fast‑paced startup environment and is eager to make significant contributions to the exciting field of LLM inference.

EXAMPLE INITIATIVES

You'll get to work on these types of projects as an Inference Performance engineer:

  • Agentic Kernels in Production
  • How we built the new fastest API for GLM-5.2
  • Live draft model training for speculative decoding
  • Making Kimi K3 Tokenization 18x faster
  • The Baseten Inference Stack
  • Driving model performance optimization

RESPONSIBILITIES

  • Implement and productionize cutting-edge inference techniques, working deep in runtime internals. That includes quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation for structured outputs, and custom scheduling and routing algorithms.
  • Profile and optimize inference end to end, from kernel launch overhead and memory layout up to request scheduling, prefill/decode disaggregation, and cache‑aware routing. Run cross‑layer investigations, such as tracing a tail‑latency regression from request timing through routing and batching down to a kernel.
  • Turn performance into cost savings. Improve tokens per GPU‑hour, raise utilization, and give customers and internal teams clear latency/throughput/cost tradeoffs.
  • Bring up and tune new model architectures on new hardware quickly, often in the same week they're released.
  • Build benchmarking frameworks that measure real‑world performance across model architectures, batch sizes, sequence lengths, and hardware configurations.
  • Contribute upstream to open‑source inference engines (vLLM, SGLang, TensorRT‑LLM), and partner closely with model, infrastructure, and customer‑facing teams to ship wins.

REQUIREMENTS

  • Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.
  • Experience with one or more general‑purpose programming languages, such as Python or C++.
  • Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).
  • Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT‑LLM.
  • Demonstrated interest and experience in LLMs.
  • Deep understanding of GPU architecture.

NICE TO HAVE

  • Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs)
  • Contributed to vLLM, SGLang, TensorRT‑LLM, or another inference engine.
  • Worked on large‑scale distributed serving: autoscaling, load balancing, multi‑region or multi‑cloud capacity.
  • Written or optimized GPU kernels (CUDA, Triton, CUTLASS, or similar)
  • Worked on quantization (FP8/FP4, AWQ, GPTQ) or speculative decoding in production.
  • Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.

BENEFITS

  • Competitive compensation, including meaningful equity
  • (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
  • Paid parental leave
  • Fertility and family‑building stipend through Carrot
  • (U.S. only) Company‑facilitated 401(k)
  • Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.

At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.

We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).

#J-18808-Ljbffr Jobleads-US
Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer- Inference Performance in Seattle, WA vacancy
  • $168.1k - $227.4k

    AWS Neuron is the complete software stack for AWS Inferentia and Trainium, AWS...  ...machine learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance model inference solutions for... 
    Performance
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    4 days ago
  • $143.7k - $194.4k

    AWS Neuron is the complete software stack for AWS Inferentia and...  .... Join the Machine Learning Inference Applications team to build the...  ...on Neuron chips.As an engineer on this team, you'll work on...  ...SGLang, optimizing model serving performance on Neuron and broadening the... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $209.1k - $282.9k

    As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for...  ...execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence... 
    Performance
    Work at office
    Local area

    ARM

    Seattle, WA
    1 day ago
  • $209.1k - $282.9k

    As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly...  ..., and product teams to enhance the performance and usability of Arm’s AI platform....  ...lifecycle management.Strong software and production engineering skills, including... 
    Performance
    Work at office
    Local area
    Visa sponsorship
    Relocation package

    ARM

    Seattle, WA
    1 day ago
  • $152k - $204k

     ...CoreWeave combines superior infrastructure performance with deep technical expertise to...  ...Learn more at  What You'll Do: Senior engineers are area owners who lead designs, raise...  ...teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours
    Shift work

    CoreWeave

    Bellevue, WA
    5 days ago
  •  ...technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate...  ...manage, and optimize cloud resources for performance and cost efficiency.Build CI/CD,...  ...transformer architecture, ML training, and inference.Experience with Infrastructure as Code... 
    Performance
    For contractors

    JP Morgan Chase

    Seattle, WA
    5 days ago
  • $180k

     ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...knowledge with their teammates. We are building the high-performance inference platform that serves Grok to millions of users every day with... 
    Performance
    Temporary work

    SpaceXAI

    Seattle, WA
    2 days ago
  • SpaceXAI is seeking a Member of Technical Staff - Inference in Seattle to design and optimize large-scale model serving systems end-to...  ...emphasizes hands-on development, CI/CD for deployment, and continued performance improvements across the full stack, from orchestration to GPU... 
    Performance

    SpaceXAI

    Seattle, WA
    2 days ago
  •  ...individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized...  ...inference routing and orchestration. Ensure software solutions are optimized for peak... 
    Performance
    Full time
    Local area
    Immediate start

    F5 Networks

    Seattle, WA
    4 days ago
  • Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product... 
    Performance

    Snowflake Computing

    Bellevue, WA
    1 day ago
  •  ...United States Digital Space LLC is seeking an engineer to join the Cloud Inference team and scale Claude across AWS, GCP, Azure, and future CSPs....  ...You will focus on fast, cost-efficient validation, performance improvements, and reliability to ensure consistent behavior... 
    Performance

    Jobleads-US

    Seattle, WA
    4 days ago
  • $143.7k - $194.4k

     ...The right candidate will possess proven software engineering skills, with experience creating and...  ...powered components that are scalable, performant and easy to maintain.* Provide guidance...  ..., including architecture, training/inference lifecycles, and optimization of model... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    5 days ago
  • $168.1k - $227.4k

     ...customized stack of hardware, firmware, and software to deliver unparalleled virtualization at...  ...EC2 Supercomputers, optimized for high-performance training and inference workloads.We are looking for an experienced software engineer to drive development for new EC2 machine... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    1 day ago
  • $169k - $338k

     ...OfficeRole summary:As a Distinguished Software Engineer, you will lead the design and development...  ...mentorship, you will enhance system performance, scalability, and maintainability,...  ...transformer-based model training and real-time inference on GPU backed infrastructure.... 
    Performance
    Full time
    Temporary work
    Part time

    Walmart

    Bellevue, WA
    1 day ago
  • $143.7k - $194.4k

     ...team at Amazon builds AWS Neuron, the software development kit used to accelerate...  ...PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement and...  ...the hardware-software boundary, our engineers build systematic infrastructure, innovate... 
    Performance
    Work experience placement
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    5 days ago
  • $143.7k - $194.4k

     ...scale systems, solving critical engineering problems, and delivering...  ...able to design and write high-performance, reliable, maintainable code...  ...non-internship professional software development experience- 2+...  ...including architecture, training/inference lifecycles, and optimization... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    19 hours ago
  • $155k - $190k

    Senior Software Engineer - ML OpsTruveta is the world’s first health provider led data platform...  ...to optimize model throughput, performance, and efficiency at scale.Are familiar...  ...maximize compute utilization and improve inference performance.Have experience working with... 
    Performance
    For contractors
    Visa sponsorship
    Work visa
    Flexible hours

    Truveta

    Seattle, WA
    5 days ago
  • $168.1k - $227.4k

    We are seeking a Senior Software Development Engineer (Sr. SDE) to join the Music Catalog team and drive...  ...and evolution of scalable, high-performance catalog systems. In this role, you will...  ..., including architecture, training/inference lifecycles, and optimization of model... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    5 days ago
  • $143.7k - $194.4k

    Marketing Measurement and Performance Science (MAPS) measures the incremental...  ...joins that causal inference demands.Key job responsibilities...  ...delivery of production software on complex, ambiguous problems...  ...across science, product, and engineering teams to scope solutions,... 
    Performance
    Contract work
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $168.1k - $227.4k

    AWS Neuron is the complete software stack for the AWS Inferentia...  ...the Sr. Software Development Engineer for the Neuron Foundation Tools...  ...develop and maintain high-performance monitoring and profiling...  ...scale distributed training and inference solutions. This organization... 
    Performance
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $143.7k - $194.4k

     ...constantly improving on the quality and performance of these recommendations. Our mission...  ...technologies.About you:You are a software engineer with an interest in modern user interfaces...  ...including transformer architecture, training/inference lifecycles, and optimization... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    19 hours ago
  • $143.7k - $194.4k

     ...with talented scientists and engineers to innovate on behalf of our...  ...non-internship professional software development experience- 2+...  ...transformer architecture, training/inference lifecycles, and optimization...  ...- Knowledge of system performance, memory management, and parallel... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $168.1k - $227.4k

     ...industry bar for security and performance across our product line.We...  ...hardware, firmware, application software and services to deliver new...  ...- Collaborate with hardware engineering teams to influence future...  ...high-performance training and inference workloads. Basic... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $165.2k - $223.6k

     ...) is building a central pipeline of Software Development Engineer (SDE) talent for anticipated roles in...  ...new features- Identify and resolve performance bottlenecks and bugs- Participate in...  ...transformer architecture, training/inference lifecycles, and optimization techniques... 
    Performance
    Internship
    Local area
    Flexible hours
    Day shift

    Amazon

    Seattle, WA
    1 day ago
  • $168.1k - $227.4k

     ...NeuronWe build Amazon Neuron, the software development kit used to...  ....As a Senior Software Engineer on our Machine Learning Applications...  ...of machine learning, high-performance computing, and distributed...  ...efforts in building distributed inference support for Pytorch in the... 
    Performance
    Internship
    Work from home
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $168.1k - $227.4k

     ...generation of AI? We are seeking a Sr. Software Development Engineer to join the AWS Mantle team and...  ...technical vision for our distributed inference engine that serves millions of customers...  ...for a globally distributed, high-performance ML inference platform serving models... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $168.1k - $227.4k

     ...dashboard summary can provide.As a Senior Software Development Engineer on the Log-Level Data team, you will...  ...and methodologies to improve system performance and capabilitiesBasic qualifications...  ..., including architecture, training/inference lifecycles, and optimization of... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $143.7k - $194.4k

     ...selection. This is AI-first engineering: you'll build the systems that...  ...and optimization to performance analysis and customer insights...  ...non-internship professional software development experience- 2+ years...  ...transformer architecture, training/inference lifecycles, and optimization... 
    Performance
    Internship
    Flexible hours

    Amazon

    Seattle, WA
    2 days ago
  • $193.3k - $261.5k

     ...closely with the hardware and software teams to ensure the right tools are available for performance profiling of large ML...  ...provides ability for performance engineers to develop and improve custom...  ...other teams including training, inference and runtime.* Collaborate with... 
    Performance
    Internship
    Local area
    Flexible hours

    Amazon

    Seattle, WA
    3 days ago
  • $184.03k - $275.4k

     ...this positionWhat you can expectAs a Software Engineer, you will play a key role in designing...  ...), debugging failures, and addressing performance regressions; andDocument code, models,...  ...data pipelines for model training and inference, including data validation, transformation... 
    Performance
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer- Inference Performance. Be the first to apply!