Staff Software Engineer, AI Inference Gateway
$180k - $220kJobleads-US
Staff Software Engineer, AI Inference Gateway
by utilizing latent consumer resources to power a sustainable, affordable, environmentally friendly cloud for everyone.
salad.com/about
Staff Software Engineer, AI Inference Gateway
The role
SaladCloud runs inference on thousands of geodistributed workstation and consumer GPUs — NVIDIA and AMD — that nobody else can use. Sitting in front of that fleet is a unique AI Gateway written in Rust, built on Pingora and WireGuard, that serves real-time and batch inference and embeddings to customers as a single, reliable API. You'll own it.
This is an infrastructure role with real autonomy. You'll design and ship the systems that decide how requests are fulfilled, when to add or drop capacity, and how to bill for every token — and you'll be the person who knows whether a new model, quantization, or GPU is actually worth running.
What you'll do
- Own the AI Gateway system end to end: request routing, streaming, batch/async job handling, and horizontal scaling across multiple gateway servers
- Build and tune the fleet-efficiency algorithms — autoscaling nodes, scoring performance, and evicting underperformers so the network converges on peak throughput per dollar
- Operate and scale the underlying inference nodes running vLLM, llama.cpp, and similar servers
- Design and maintain OpenTelemetry-based observability across the gateway and the fleet; pay-per-token and subscription billing already rides on this telemetry, so you'll keep that integration accurate
- Spend roughly 25% of your time benchmarking new models, quantizations, and GPU hardware, and work with the product and marketing teams to make the call on what goes into production
- Collaborate on cross-system design and integrations with the rest of Salad engineering; explain trade-offs clearly to both technical and non-technical teammates
The team
You'll join a small team — five engineers plus the CTO, who still writes code, and two deeply technical product managers who ship their own fixes rather than queue them for you. You'll report directly to the CTO. Everyone on the engineering team has 10+ years of experience and has built systems that already run at massive scale — with minimal incidents in a typical year, because we ship things that are production-worthy the first time.
You'll be the first hire with inference-gateway expertise and you'll be asked to own the whole surface. You won't do it alone: your teammates are eager to learn the domain, will pair on design and integrations, and share after-hours support. The bar is high; so is the support.
We've listed a lot below. If you're strong in two of the first three — Rust, LLM inference servers, distributed systems — apply, and expect to learn the rest fast with people who'll help.
What we're looking for
- Strong production Rust experience, ideally on high-throughput networked services
- Hands-on experience running LLM inference servers (vLLM, llama.cpp, TGI, TensorRT-LLM, or similar) — you know what a KV cache is and why it fills up
- Comfort operating what you build: metrics, tracing, on-call instincts
- Willing to be on-call for a system you helped make quiet — we treat after-hours pages as a design bug to fix, not a lifestyle
- Clear written and verbal communication; you can run a design without hand-holding and bring others along
Nice to have
- Pingora, Tokio, or proxy/gateway internals
- Experience with heterogeneous or consumer-grade GPU fleets, CUDA, or ROCm
- Quantization formats (GGUF, AWQ, GPTQ, FP8) and their performance characteristics
Why Salad
You'll work on a hard, unusual problem — coaxing datacenter-grade reliability out of hardware in people's homes — with the freedom to make major decisions and see them hit production quickly.
We're scaling rapidly: demand for inference on SaladCloud currently exceeds the capacity we can bring online, and this role exists to keep up with it. There is no shortage of work and no prospect of it ramping down.
The platform handles workload security and isolation, so you can focus on performance. We don't log or train on our customers' prompts or data.
- Unlimited PTO
- 75% of health insurance premiums covered for you and your dependents
- Dental and vision coverage
- 401(k) plan
- Stock options
- Company-provided computer
- Fully remote, with flexible hours
Compensation
$180,000-$220,000/year
The pay range for this role is:
180,000 - 220,000 USD per year (Remote (United States))
#J-18808-Ljbffr Jobleads-US- ...Salad is seeking a Staff Software Engineer for the AI Inference Gateway to own the system end-to-end, from request routing to scaling across gateway servers. You will optimize fleet efficiency, operate inference nodes (vLLM, llama.cpp, etc.), and maintain observability...SuggestedRemote job
- Anthropic is seeking an experienced software engineer to design, build, and maintain scalable inference systems powering Claude for millions of users. You will implement... ...teams to ensure reliable, scalable operations from SF offices. #J-18808-Ljbffr AI Chopping BlockSuggested
- Staff Software Development EngineerWe use cookies to make our... ...Software Development Engineer** to be the hands-on... ...infrastructure, and brings AI capabilities into... ...monitor ML and AI model inference on AWS, including... ...ACH processing, payment gateways, underwriting systems,...SuggestedFull timeWork at office
- ...You'll help build Confluent Cloud's AI capabilities — the layer that lets customers... ...moving data out to a separate system to run inference or build an agent, our customers do it in... ...and process events at scale. As an engineer, you'll own delivery of significant pieces...SuggestedLive in
$141.13k - $211.69k
...Explicitly requires Vibe Coding and AI-assisted engineering practices, using agentic workflows and... ...of AI outputs. About the Role Staff Software Engineer based at John Deere World Headquarters... ...applications (AWS Lambda and API Gateway or equivalent). ~5+ years designing...SuggestedFlexible hours- # Senior / Staff Software EngineerTrack to Lead Engineer← All open positionsAustinOnsite preferred (hybrid in rare cases... ...they run on.* Build and scale our AI employees platform: the application... ...models, ML pipelines, real-time inference) who wants to do it full time.* Engineer...Full time
$225k - $265k
...B2B SAAS data observability software. Join the company that's building... ...infrastructure for the AI era. At Cribl, we partner with... ...and a group of highly-skilled engineers to shape the future of search... ...models, Prompt Engineering, and Inference PlatformsThis position will...Temporary workRemote work£110k - £170k per year
...right place.The roleWe're looking for a Staff Software Engineer to join our engineering team and play a... ...Run, Cloud SQL, Cloud Storage + more)AI / agentic platformOpus/GPT/Gemini/and many... ...engineMCP servers as the tool layerAI gateway (Bifrost) for logging, routing,...Remote workWorldwideFlexible hoursShift work$138k - $163k
Software Engineer - AI Infrastructure We’re seeking a software engineer to support our AI infrastructure team in Columbia, MD. In this role,... ...applications. Responsibilities: Procure, configure, and test new inference models, preparing them for release to our user base....Temporary work$118.4k - $219.8k
Staff Software Engineer — Content Tools*Content Tools • CoCounsel Agent Context* Overview of the RoleThe... ...— the purpose-built APIs that give AI agents precise, governed access to Thomson... ...-as-code (Terraform) and API gateways (such as Apigee) — enough to operate, secure...Contract workWork at officeLocal areaFlexible hours$165k - $247.5k
* /* /* Senior Staff Software Engineer - AI Governance# Senior Staff Software Engineer - AI GovernanceAtlanta, Georgia | Engineering###### Strength... ...analysis.* Experience designing SDK telemetry, model gateways, runtime instrumentation, or post-market monitoring.* Proficiency...Work experience placementWork at officeWorldwide3 days per week1 day per week$255k - $310k
# Senior Staff Software Engineer, ServingUnited States9 hours agoID 1185189Price on request## DetailsEmployment... ...DescriptionLiftoff is a leading AI-powered performance marketing platform... .... • Develop and optimize GPU-powered inference services that execute neural network...Full timeWork experience placementRemote work$224k - $260k
...Staff Software Engineer, Artificial Intelligence/LLM Lead the end-to-end development of LLM-powered... ...year About The Role About Beacon AI We're a fast-moving team of... ...series or video. Familiarity with GPU inference, Triton, or TensorRT-LLM. Aviation...Permanent employmentFull timeLocal areaRemote work3 days per week$190.9k - $334.1k
...Senior Staff Software Engineer_Voice Connectivity Full-time Employee Type: Regular Region: AMS... ...work. Today, ServiceNow is the AI control tower for business reinvention.... ...differs — different SIP providers, media gateways, authentication schemes, and CCaaS vendor...Full timeWork at officeImmediate startRemote workFlexible hours- As a Staff Engineer at Capital One, you will be a part of a community of... ...redefine the speed and quality of software development. We are a high-... ...a team to become elite AI-native builders. You will own... ...architectures (GraphQL, REST, API gateways, gRPC, caching, and streaming...
- ...States Full time JR0037924 Job Title: Staff Software Development Engineer About Trellix Trellix is a global... ...security standards, and modern AI development tooling. This role is hybrid... ...environments leveraging VPC architecture, NAT Gateways, Security Groups, EC2 fleets, Load...Full timeWork at officeRemote workFlexible hours3 days per week
- ...Thomson Reuters is seeking a Senior Inference Engineer, AI to productionize, optimize, and scale AI/LLM workloads powering TR’s AI-driven products. The role focuses on deploying across multi-cloud environments (AWS, Azure, GCP) and on‑prem Kubernetes clusters, reducing...
- ...Thomson Reuters is seeking a Senior Inference Engineer, AI to collaborate with platform teams to productionize and scale AI and LLM workloads across AWS, Azure, GCP and internal Kubernetes clusters. You will optimize models for low latency, develop containerized inference...
- Black Sesame seeks an experienced AI/ML engineer to deploy AI inference and end-to-end enablement, AI framework integration, model accuracy, and performance tuning. You will develop high-quality software that enables state-of-the-art AI inference on BST Intelligence processors...
- ...INFRO is seeking a senior backend engineer to own large parts of the gateway path between customer applications and AI provider accounts. You will design and ship end-to-end gateway features in TypeScript running on Cloudflare Workers, backed by Postgres, focusing on...
$130k - $170k
...reflecting our world. We are seeking a Staff Software Engineer to lead the development of innovative... ...rapidly emerging landscape of frontier AI capabilities. This engineer will help... ...logic that leverages API standards, NBCU gateways, and established best practices and...Remote jobLocal area- ...Machines Lab Inc. in San Francisco, California is seeking an infrastructure research engineer to design, optimize, and scale the systems that power large AI models. Your work will make inference faster, more cost-effective, more reliable, and more reproducible to enable our...
- OpenAI is seeking an AI Systems Engineer to scale infrastructure behind training and evaluation workflows. You will own projects from bottleneck... ...and collaboration with researchers. You will build shared inference and grading platforms, improve scheduling, and develop...
- ...Observable Intuition, Inc. seeks a founding Infrastructure Engineer to define and own the production inference platform behind our data-driven AI layer. You will collaborate with the founding team to build core systems from the ground up, make foundational architectural...
- ...Collective Intuition, Inc. seeks a Founding Infrastructure Engineer to define and own the production inference platform behind a new layer of AI intelligence. You will build core systems, set foundational architecture decisions, and influence engineering culture from day...
- About Beacon AI We’re a fast-moving team of aviators, engineers, and operators building an AI platform to make flying... ...Overview We are seeking skilled Staff Software Engineers, Cloud Infrastructure... ...ECS/EKS, Lambda, or Batch for inference jobs. Build and maintain application...Permanent employmentFull timeLocal areaRemote work3 days per week
- SpaceX in Palo Alto seeks a Software Engineer, Inference (AI Data Engineering) to design and optimize high-throughput AI model serving systems. You will own distributed infrastructure from routing to batching and work with SpaceX AI teams to deliver reliable, scalable...
- ...About the Role The Grading Operations software engineering team builds and operates the internal software... ...and shipping what you design, and AI tooling is central to how we work... ...ECS), serverless compute behind an API gateway, and managed messaging and data stores...Remote workWorldwide
$253.9k - $298.7k
...about working at Coinbase. As a Senior Staff Software Engineer on the Data Platform team within... ...distributed systems, data engineering, and AI-readiness, reporting to the Senior... ...ML training, feature stores, real-time inference, and multi-agent AI architectures at scale...Local area- ...sensing, computer vision, AI-driven analytics, and... .... ABOUT THE ENGINEERING TEAM The Engineering... ...hardware and embedded software to computer vision pipelines... ...vision and ML inference Own our Infrastructure... ...advanced ingress / API gateway patterns Experience...Remote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference Gateway. Be the first to apply!

