Software Engineer- Inference Performance
Jobleads-US
ABOUT BASETEN
Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting‑edge models into production. We're growing quickly and recently raised our $1.5B Series F, led by Altimeter Capital, Conviction Partners, and Spark Capital. Join us and help build the platform engineers turn to ship AI products.
THE ROLE
We're looking for inference performance engineers who want to make the world's most demanding AI workloads run faster and more efficiently. You'll work across the stack, from the inference engine and runtime through scheduling, serving, and routing. Along the way you'll apply techniques like prefill/decode disaggregation, speculative decoding, and KV-cache management. You'll reason from first principles about where time and memory go, find what's holding performance back, and close the gap. Your work directly impacts how fast our customers' models run and how efficiently we serve them. This role is ideal for someone who thrives in a fast‑paced startup environment and is eager to make significant contributions to the exciting field of LLM inference.
EXAMPLE INITIATIVES
You'll get to work on these types of projects as an Inference Performance engineer:
- Agentic Kernels in Production
- How we built the new fastest API for GLM-5.2
- Live draft model training for speculative decoding
- Making Kimi K3 Tokenization 18x faster
- The Baseten Inference Stack
- Driving model performance optimization
RESPONSIBILITIES
- Implement and productionize cutting-edge inference techniques, working deep in runtime internals. That includes quantization, speculative decoding, KV-cache reuse, chunked prefill, LoRA, guided generation for structured outputs, and custom scheduling and routing algorithms.
- Profile and optimize inference end to end, from kernel launch overhead and memory layout up to request scheduling, prefill/decode disaggregation, and cache‑aware routing. Run cross‑layer investigations, such as tracing a tail‑latency regression from request timing through routing and batching down to a kernel.
- Turn performance into cost savings. Improve tokens per GPU‑hour, raise utilization, and give customers and internal teams clear latency/throughput/cost tradeoffs.
- Bring up and tune new model architectures on new hardware quickly, often in the same week they're released.
- Build benchmarking frameworks that measure real‑world performance across model architectures, batch sizes, sequence lengths, and hardware configurations.
- Contribute upstream to open‑source inference engines (vLLM, SGLang, TensorRT‑LLM), and partner closely with model, infrastructure, and customer‑facing teams to ship wins.
REQUIREMENTS
- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.
- Experience with one or more general‑purpose programming languages, such as Python or C++.
- Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).
- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT‑LLM.
- Demonstrated interest and experience in LLMs.
- Deep understanding of GPU architecture.
NICE TO HAVE
- Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs)
- Contributed to vLLM, SGLang, TensorRT‑LLM, or another inference engine.
- Worked on large‑scale distributed serving: autoscaling, load balancing, multi‑region or multi‑cloud capacity.
- Written or optimized GPU kernels (CUDA, Triton, CUTLASS, or similar)
- Worked on quantization (FP8/FP4, AWQ, GPTQ) or speculative decoding in production.
- Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.
BENEFITS
- Competitive compensation, including meaningful equity
- (U.S. only) 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company wide Winter Break (our offices are closed from Christmas Eve to New Year's Day!)
- Paid parental leave
- Fertility and family‑building stipend through Carrot
- (U.S. only) Company‑facilitated 401(k)
- Exposure to a variety of ML startups, offering unparalleled learning and networking opportunities.
At Baseten, we are committed to fostering a diverse and inclusive workplace. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, or veteran status.
We are an Equal Opportunity Employer and will consider qualified applicants with criminal histories in a manner consistent with applicable law (by example, the requirements of the San Francisco Fair Chance Ordinance, where applicable).
#J-18808-Ljbffr Jobleads-US$168.1k - $227.4k
AWS Neuron is the complete software stack for AWS Inferentia and Trainium, AWS... ...machine learning. This senior software engineering role is part of the Machine Learning Inference Applications team and focuses on delivering high-performance model inference solutions for...PerformanceWork experience placementInternshipLocal areaFlexible hours$143.7k - $194.4k
AWS Neuron is the complete software stack for AWS Inferentia and... .... Join the Machine Learning Inference Applications team to build the... ...on Neuron chips.As an engineer on this team, you'll work on... ...SGLang, optimizing model serving performance on Neuron and broadening the...PerformanceInternshipFlexible hours$209.1k - $282.9k
As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference runtime for... ...execution, kernel development and optimization, and performance benchmarking and analysis. Your work will directly influence...PerformanceWork at officeLocal area$209.1k - $282.9k
As an engineer on Arm’s AI Inference Cloud team, you will shape the technical direction and develop highly... ..., and product teams to enhance the performance and usability of Arm’s AI platform.... ...lifecycle management.Strong software and production engineering skills, including...PerformanceWork at officeLocal areaVisa sponsorshipRelocation package$152k - $204k
...CoreWeave combines superior infrastructure performance with deep technical expertise to... ...Learn more at What You'll Do: Senior engineers are area owners who lead designs, raise... ...teams to evolve our Kubernetes-native inference platform and meet strict P99 SLAs at scale...PerformancePermanent employmentFull timeTemporary workCasual workWork at officeFlexible hoursShift work- ...technology products.As a Senior Lead Software Engineer at JPMorgan Chase within the Corporate... ...manage, and optimize cloud resources for performance and cost efficiency.Build CI/CD,... ...transformer architecture, ML training, and inference.Experience with Infrastructure as Code...PerformanceFor contractors
$180k
...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who... ...knowledge with their teammates. We are building the high-performance inference platform that serves Grok to millions of users every day with...PerformanceTemporary work- SpaceXAI is seeking a Member of Technical Staff - Inference in Seattle to design and optimize large-scale model serving systems end-to... ...emphasizes hands-on development, CI/CD for deployment, and continued performance improvements across the full stack, from orchestration to GPU...Performance
- ...individual can thrive.Job DescriptionThe AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized... ...inference routing and orchestration. Ensure software solutions are optimized for peak...PerformanceFull timeLocal areaImmediate start
- Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product...Performance
- ...United States Digital Space LLC is seeking an engineer to join the Cloud Inference team and scale Claude across AWS, GCP, Azure, and future CSPs.... ...You will focus on fast, cost-efficient validation, performance improvements, and reliability to ensure consistent behavior...Performance
$143.7k - $194.4k
...The right candidate will possess proven software engineering skills, with experience creating and... ...powered components that are scalable, performant and easy to maintain.* Provide guidance... ..., including architecture, training/inference lifecycles, and optimization of model...PerformanceInternshipFlexible hours$168.1k - $227.4k
...customized stack of hardware, firmware, and software to deliver unparalleled virtualization at... ...EC2 Supercomputers, optimized for high-performance training and inference workloads.We are looking for an experienced software engineer to drive development for new EC2 machine...PerformanceInternshipFlexible hours$169k - $338k
...OfficeRole summary:As a Distinguished Software Engineer, you will lead the design and development... ...mentorship, you will enhance system performance, scalability, and maintainability,... ...transformer-based model training and real-time inference on GPU backed infrastructure....PerformanceFull timeTemporary workPart time$143.7k - $194.4k
...team at Amazon builds AWS Neuron, the software development kit used to accelerate... ...PyTorch and JAX enabling unparalleled ML inference and training performance.The Inference Enablement and... ...the hardware-software boundary, our engineers build systematic infrastructure, innovate...PerformanceWork experience placementInternshipFlexible hours$143.7k - $194.4k
...scale systems, solving critical engineering problems, and delivering... ...able to design and write high-performance, reliable, maintainable code... ...non-internship professional software development experience- 2+... ...including architecture, training/inference lifecycles, and optimization...PerformanceInternshipFlexible hours$155k - $190k
Senior Software Engineer - ML OpsTruveta is the world’s first health provider led data platform... ...to optimize model throughput, performance, and efficiency at scale.Are familiar... ...maximize compute utilization and improve inference performance.Have experience working with...PerformanceFor contractorsVisa sponsorshipWork visaFlexible hours$168.1k - $227.4k
We are seeking a Senior Software Development Engineer (Sr. SDE) to join the Music Catalog team and drive... ...and evolution of scalable, high-performance catalog systems. In this role, you will... ..., including architecture, training/inference lifecycles, and optimization of model...PerformanceInternshipFlexible hours$143.7k - $194.4k
Marketing Measurement and Performance Science (MAPS) measures the incremental... ...joins that causal inference demands.Key job responsibilities... ...delivery of production software on complex, ambiguous problems... ...across science, product, and engineering teams to scope solutions,...PerformanceContract workInternshipFlexible hours$168.1k - $227.4k
AWS Neuron is the complete software stack for the AWS Inferentia... ...the Sr. Software Development Engineer for the Neuron Foundation Tools... ...develop and maintain high-performance monitoring and profiling... ...scale distributed training and inference solutions. This organization...PerformanceInternshipWork from homeFlexible hours$143.7k - $194.4k
...constantly improving on the quality and performance of these recommendations. Our mission... ...technologies.About you:You are a software engineer with an interest in modern user interfaces... ...including transformer architecture, training/inference lifecycles, and optimization...PerformanceInternshipFlexible hours$143.7k - $194.4k
...with talented scientists and engineers to innovate on behalf of our... ...non-internship professional software development experience- 2+... ...transformer architecture, training/inference lifecycles, and optimization... ...- Knowledge of system performance, memory management, and parallel...PerformanceInternshipFlexible hours$168.1k - $227.4k
...industry bar for security and performance across our product line.We... ...hardware, firmware, application software and services to deliver new... ...- Collaborate with hardware engineering teams to influence future... ...high-performance training and inference workloads. Basic...PerformanceInternshipFlexible hours$165.2k - $223.6k
...) is building a central pipeline of Software Development Engineer (SDE) talent for anticipated roles in... ...new features- Identify and resolve performance bottlenecks and bugs- Participate in... ...transformer architecture, training/inference lifecycles, and optimization techniques...PerformanceInternshipLocal areaFlexible hoursDay shift$168.1k - $227.4k
...NeuronWe build Amazon Neuron, the software development kit used to... ....As a Senior Software Engineer on our Machine Learning Applications... ...of machine learning, high-performance computing, and distributed... ...efforts in building distributed inference support for Pytorch in the...PerformanceInternshipWork from homeFlexible hours$168.1k - $227.4k
...generation of AI? We are seeking a Sr. Software Development Engineer to join the AWS Mantle team and... ...technical vision for our distributed inference engine that serves millions of customers... ...for a globally distributed, high-performance ML inference platform serving models...PerformanceInternshipFlexible hours$168.1k - $227.4k
...dashboard summary can provide.As a Senior Software Development Engineer on the Log-Level Data team, you will... ...and methodologies to improve system performance and capabilitiesBasic qualifications... ..., including architecture, training/inference lifecycles, and optimization of...PerformanceInternshipFlexible hours$143.7k - $194.4k
...selection. This is AI-first engineering: you'll build the systems that... ...and optimization to performance analysis and customer insights... ...non-internship professional software development experience- 2+ years... ...transformer architecture, training/inference lifecycles, and optimization...PerformanceInternshipFlexible hours$193.3k - $261.5k
...closely with the hardware and software teams to ensure the right tools are available for performance profiling of large ML... ...provides ability for performance engineers to develop and improve custom... ...other teams including training, inference and runtime.* Collaborate with...PerformanceInternshipLocal areaFlexible hours$184.03k - $275.4k
...this positionWhat you can expectAs a Software Engineer, you will play a key role in designing... ...), debugging failures, and addressing performance regressions; andDocument code, models,... ...data pipelines for model training and inference, including data validation, transformation...PerformanceFull timeWork at officeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Software Engineer- Inference Performance. Be the first to apply!
- software engineer full time Seattle, WA
- graduate software developer no experience Seattle, WA
- software engineer healthcare Seattle, WA
- network software engineer Seattle, WA
- software engineer internship Seattle, WA
- senior software engineer Seattle, WA
- software system engineer Seattle, WA
- software developer Seattle, WA
- ngo software engineer Seattle, WA
- startup software engineer Seattle, WA

