Staff Software Engineer, AI Inference
$190k - $230kTLA Tech
Role Description
As we continue to scale our AI platform, we're investing in our own inference stack to deliver best-in-class performance, reliability, cost efficiency, and flexibility across the latest generation of open-source language models. We're looking for a Staff Software Engineer to spearhead this effort.
You'll define the architecture, evaluate emerging technologies, and build the systems that power model serving. You'll partner closely with machine learning, infrastructure, and product engineering to establish the foundation for how AI models are deployed, optimized, monitored, and operated in production. This is a highly hands-on technical role. You'll spend the majority of your time designing, building, and optimizing production systems while helping shape our long-term AI infrastructure strategy.
Responsibilities
- Lead the design and development of our production inference platform.
- Define the technical roadmap for inference infrastructure, model serving, and runtime optimization.
- Build and operate scalable, cost-effective systems for serving large language models in production.
- Evaluate and integrate modern inference technologies, frameworks, and serving runtimes.
- Optimize latency, throughput, GPU utilization, memory efficiency, and infrastructure cost.
- Develop systems for model deployment, traffic routing, autoscaling, scheduling, observability, and operational excellence.
- Partner with ML engineers to productionize new models and inference techniques.
- Establish benchmarking methodologies to evaluate new models, runtimes, and hardware.
- Make key architectural decisions around when to build internally versus leverage open-source or commercial solutions.
- Mentor engineers as the team grows and help establish engineering best practices for AI infrastructure.
Qualifications
- Significant experience designing and operating production AI inference systems.
- Experience building or leading production LLM serving infrastructure.
- Deep experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or comparable technologies.
- Strong background in distributed systems, backend infrastructure, or high-performance platform engineering.
- Experience optimizing inference performance across GPU workloads, including latency, throughput, batching, memory utilization, and serving efficiency.
- Experience operating GPU infrastructure in production.
- Strong proficiency in Python and at least one systems programming language (such as Go, Rust, or C++).
- Proven ability to lead technical architecture for complex infrastructure initiatives.
- Excellent communication skills and the ability to influence technical direction across engineering teams.
Requirements
- Salary Range ($190- $230K) plus health insurance and equity.
- United States - Remote Pay Range: $190,000 - $230,000 USD
$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...Our mandate is to make inference deployment boring and... ...and unattended. As a Software Engineer on the Launch... ...Currently, we expect all staff to be in one of our...SuggestedFull timeWork at officeVisa sponsorshipFlexible hoursShift work- ...leading security-first enterprise AI company. We build cutting-edge... ...is a team of researchers, engineers, designers, and more, who are... ...looking for Members of Technical Staff to join the Model Serving team... ...latency and throughput of inference.Strong understanding or working...SuggestedFull timeWork experience placementWork at officeLocal areaRemote workHome office
$190k - $230k
Role Description As we continue to scale our AI platform, we're investing in our own inference stack to deliver best-in-class performance, reliability... ...open-source language models. We're looking for a Staff Software Engineer to spearhead this effort. You'll define the...SuggestedFull timeRemote work$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...role The Cloud Inference team scales and optimizes... ...Have significant software engineering experience,... ...Currently, we expect all staff to be in one of our offices...SuggestedFull timeWork at officeVisa sponsorshipFlexible hours- ...of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the... ...systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between...SuggestedFull time
- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...the future. We are seeking a Staff Engineer to help our development of... ...of AI training and inference at scale. As a Staff Engineer... ...10+ years of experience in software engineering, platform engineering...Full timeWork at officeLocal areaImmediate startWork from homeFlexible hours
- ...Business Area: Engineering Seniority Level: Mid-Senior level Job... ...: Cloudera is looking for a Staff Software Engineer to join the Enterprise AI Platform team and help drive development... ...building and deploying AI Inference and Generative AI applications....Work from homeFlexible hours
$152k - $241.5k
NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in...Full timeRemote work$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference...Temporary workFor contractorsWork experience placement$152k - $241.5k
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize... ...software that powers today’s most sophisticated AI applications. Our team is responsible for developing...Full timeRemote work$2,000 per month
...Etched is building AI chips that are hard-coded... ...design of the Sohu host software stack Implement high... ...batching and real time inference Implement inference-... ...Cupertino, and greatly value engineering skills. We do not have... ...all of our technical staff to contribute to both...Full timeWork at officeRelocation package$92k - $135k
...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services... .... What You’ll Do: Join the Inference team to ship production features that improve... ...quickly with mentorship from experienced engineers. About the role: Implement well...Permanent employmentFull timeTemporary workCasual workInternshipWork at officeRemote workFlexible hours$229k - $343k
...digital services.We’re looking for a Staff Software Engineer to join Snap Inc on our Feature Store... ...features reliably for large-scale batch inference and low-latency online inference, maintaining... ...with responsible use of emerging AI technologies to improve engineering...Full timeLive inWork at officeLocal area$193.3k - $261.5k
...builds AWS Neuron, the software development kit used to... ...enabling unparalleled ML inference and training... ...software boundary, our engineers build systematic infrastructure... ...of what's possible in AI acceleration.As part of... ...employees, supervisors, and staff; adhere to standards of...Work experience placementInternshipLocal areaFlexible hours- ...Description:DataRobot delivers AI that maximizes impact and... ...DataRobot’s Fleet team is the engine behind how our platform runs... ...That’s where you come in.As a Staff Software Engineer, you’ll be responsible... ...for training and inference.Why Join the Fleet Management...Full timeLocal areaRemote workWorldwideFlexible hours
- ...reliable on-demand, logistics engine for last-mile retail... ...to join our team. As a Staff Machine Learning... .... Proficiency in using AI coding tools (e.g., Claude... ...Codex, Cursor) in the full software development lifecycle,... ...applied ML for Causal Inference and Recommendation...Hourly payWork at officeLocal areaRemote workFlexible hours
$242k - $389k
..., WA / Remote (United States)Software – Software Systems /Full-time... ...alongside a team of strong software engineers and act as a force multiplier... ...cutting-edge ML Training OR Inference performance optimization... ...use artificial intelligence (AI) tools to support parts of...Full timeRemote work$135k - $160k
...the ultimate goal of enabling human life on Mars.APPLICATION SOFTWARE ENGINEER, INFERENCEThe application software team is the central... ...respect, and support.Our team maintains a high-performance AI inference platform that serves the best models internally at SpaceX to...Permanent employmentTemporary workRemote workWorldwideWeekend work$207k - $301k
...machine earning pipelines for model training, inference, and integration with high-throughput Ad... ...experience.8 years of experience in software development.5 years of experience... ...or recommender systems.Google's software engineers develop the next-generation technologies...- ...foundation of a neurosymbolic AI agent that learns more about... ...Join Onton as a Founding Engineer and set the strategic foundation... ...out a performant and scalable inference engine to support more... ...and is passionate about making software tools accessible to all, we want...Full timeWork at officeLocal areaRemote workRelocation3 days per week
- ...entertainment company and the leader in AI music. We are backed by... ...team at the Senior and Staff levels. These aren’t maintenance... ...features end-to-end, drive engineering quality, and contribute to architectural... ...features, and real-time AI inference all live on this platform. If...Full timeWork at officeLocal areaWorldwide
$216k - $258k
...full potential of their data for AI applications. The platform... ...infrastructure that power AI inference at 100K+ QPS with millisecond... ...Evolve Tecton’s query execution engine to support complex, multi-stage... ...distributed and/or highly concurrent software systems ~ Degree in Computer...Full timeRemote workFlexible hours$190k - $210k
Role Description Vanilla is seeking a Staff Software Engineer - AI Applications with a strong background in software development, data science,... ...can build tooling to support model training, evaluation, inference serving, monitoring, and alerting. ~You want to use the...Full timeWork experience placementWork at officeHome officeFlexible hours$225k - $265k
...telemetry infrastructure for the AI era. At Cribl, we partner... ...and grounded in a simple idea: software is a people business. Cribl... ...and a group of highly-skilled engineers to shape the future of search... ..., Prompt Engineering, and Inference Platforms ~This position will...Full timeTemporary workRemote work- ...acted on the same day. We're hiring a staff-level engineer to own that path end to end. It's a... ...the orchestration layer for ingestion, inference, and downstream jobs. ~Prove the numbers... ..., and Ruby. ~Set the bar on AI-assisted engineering. Who You Are...Full timeTemporary workRemote workFlexible hours
- ...group of high-performance, high-octane engineers, so direction gets shaped together, and... ...Work directly with founders, design, and inference teams to turn big ideas into interfaces... ...and mentorship. Qualifications ~Staff-level depth in TypeScript and React....Full timeLocal area
$265k
Role Description We are seeking a Senior Staff Software Engineer to lead the architecture and evolution... ...on which every product, workflow, and AI capability is built. You will define... ...infrastructure for AI/LLM workloads: inference serving, agent runtimes, GPU scheduling...Full timeRemote workWork from homeFlexible hours$235.7k - $277k
...Description You'll help build Confluent Cloud's AI capabilities — the layer that lets... ...moving data out to a separate system to run inference or build an agent, our customers do it in... ...and process events at scale. As an engineer, you'll own delivery of significant pieces...Full timeLive in$151.8k - $332.2k
What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on...Full timeWork at officeRemote work- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...groundbreaking AI training and inference possible.The Lambda Infrastructure Engineering organization forges the... ...Role:We are seeking a seasoned Staff Storage Software Engineer with deep experience designing...Work at officeLocal areaWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference. Be the first to apply!
- software product owner Remote
- software engineer - cloud services Remote
- id software Remote
- mid-level software developer Remote
- healthcare software sales Remote
- software technical support Remote
- software asset management analyst Remote
- android software developer Remote
- software implementation project manager Remote
- software trainer Remote









