Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, AI Inference

$190k - $230k
Full-time

TLA Tech

Role Description

As we continue to scale our AI platform, we're investing in our own inference stack to deliver best-in-class performance, reliability, cost efficiency, and flexibility across the latest generation of open-source language models. We're looking for a Staff Software Engineer to spearhead this effort.

You'll define the architecture, evaluate emerging technologies, and build the systems that power model serving. You'll partner closely with machine learning, infrastructure, and product engineering to establish the foundation for how AI models are deployed, optimized, monitored, and operated in production. This is a highly hands-on technical role. You'll spend the majority of your time designing, building, and optimizing production systems while helping shape our long-term AI infrastructure strategy.

Responsibilities

  • Lead the design and development of our production inference platform.
  • Define the technical roadmap for inference infrastructure, model serving, and runtime optimization.
  • Build and operate scalable, cost-effective systems for serving large language models in production.
  • Evaluate and integrate modern inference technologies, frameworks, and serving runtimes.
  • Optimize latency, throughput, GPU utilization, memory efficiency, and infrastructure cost.
  • Develop systems for model deployment, traffic routing, autoscaling, scheduling, observability, and operational excellence.
  • Partner with ML engineers to productionize new models and inference techniques.
  • Establish benchmarking methodologies to evaluate new models, runtimes, and hardware.
  • Make key architectural decisions around when to build internally versus leverage open-source or commercial solutions.
  • Mentor engineers as the team grows and help establish engineering best practices for AI infrastructure.

Qualifications

  • Significant experience designing and operating production AI inference systems.
  • Experience building or leading production LLM serving infrastructure.
  • Deep experience with one or more modern inference runtimes and frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Hugging Face TGI, NVIDIA Dynamo, or comparable technologies.
  • Strong background in distributed systems, backend infrastructure, or high-performance platform engineering.
  • Experience optimizing inference performance across GPU workloads, including latency, throughput, batching, memory utilization, and serving efficiency.
  • Experience operating GPU infrastructure in production.
  • Strong proficiency in Python and at least one systems programming language (such as Go, Rust, or C++).
  • Proven ability to lead technical architecture for complex infrastructure initiatives.
  • Excellent communication skills and the ability to influence technical direction across engineering teams.

Requirements

  • Salary Range ($190- $230K) plus health insurance and equity.
  • United States - Remote Pay Range: $190,000 - $230,000 USD
Vacancy posted 2 hours ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, AI Inference in Remote vacancy
  • $320k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...Our mandate is to make inference deployment boring and...  ...and unattended. As a Software Engineer on the Launch...  ...Currently, we expect all staff to be in one of our... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    Remote
    1 day ago
  •  ...leading security-first enterprise AI company. We build cutting-edge...  ...is a team of researchers, engineers, designers, and more, who are...  ...looking for Members of Technical Staff to join the Model Serving team...  ...latency and throughput of inference.Strong understanding or working... 
    Suggested
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    New York, NY
    5 days ago
  • $190k - $230k

    Role Description As we continue to scale our AI platform, we're investing in our own inference stack to deliver best-in-class performance, reliability...  ...open-source language models. We're looking for a Staff Software Engineer to spearhead this effort. You'll define the... 
    Suggested
    Full time
    Remote work

    Syllo

    Remote
    5 days ago
  • $320k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...role The Cloud Inference team scales and optimizes...  ...Have significant software engineering experience,...  ...Currently, we expect all staff to be in one of our offices... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Remote
    1 day ago
  •  ...of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the...  ...systems and/or platforms. Experience in serving LLMs using inference engines like vLLM, TensorRT-LLM, TEI, SGLang, and knowing tradeoffs between... 
    Suggested
    Full time

    Snowflake

    Remote
    1 day ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...the future. We are seeking a Staff Engineer to help our development of...  ...of AI training and inference at scale. As a Staff Engineer...  ...10+ years of experience in software engineering, platform engineering... 
    Full time
    Work at office
    Local area
    Immediate start
    Work from home
    Flexible hours

    Lambda

    Bellevue, WA
    1 day ago
  •  ...Business Area: Engineering Seniority Level: Mid-Senior level Job...  ...:   Cloudera is looking for a Staff Software Engineer to join the Enterprise AI Platform team and help drive development...  ...building and deploying AI Inference and Generative AI applications.... 
    Work from home
    Flexible hours

    Cloudera

    Austin, TX
    22 days ago
  • $152k - $241.5k

    NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by contributing directly to upstream inference engines like vLLM and SGLang-ensuring they run best‑in... 
    Full time
    Remote work

    Nvidia

    New York, NY
    2 days ago
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779 Description & Requirements Our team: Join the team that is building the core infrastructure for AI at Bloomberg. The Bloomberg AI Inference... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    5 days ago
  • $152k - $241.5k

    NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize...  ...software that powers today’s most sophisticated AI applications. Our team is responsible for developing... 
    Full time
    Remote work

    Nvidia

    Texas
    22 hours ago
  • $2,000 per month

     ...Etched is building AI chips that are hard-coded...  ...design of the Sohu host software stack Implement high...  ...batching and real time inference Implement inference-...  ...Cupertino, and greatly value engineering skills. We do not have...  ...all of our technical staff to contribute to both... 
    Full time
    Work at office
    Relocation package

    Etched

    Remote
    1 day ago
  • $92k - $135k

     ...CoreWeave is the AI Hyperscaler™, delivering a cloud platform of cutting edge services...  ....  What You’ll Do: Join the Inference team to ship production features that improve...  ...quickly with mentorship from experienced engineers. About the role: Implement well... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Internship
    Work at office
    Remote work
    Flexible hours

    Coreweave

    Washington DC
    1 day ago
  • $229k - $343k

     ...digital services.We’re looking for a Staff Software Engineer to join Snap Inc on our Feature Store...  ...features reliably for large-scale batch inference and low-latency online inference, maintaining...  ...with responsible use of emerging AI technologies to improve engineering... 
    Full time
    Live in
    Work at office
    Local area

    Snap

    Palo Alto, CA
    2 days ago
  • $193.3k - $261.5k

     ...builds AWS Neuron, the software development kit used to...  ...enabling unparalleled ML inference and training...  ...software boundary, our engineers build systematic infrastructure...  ...of what's possible in AI acceleration.As part of...  ...employees, supervisors, and staff; adhere to standards of... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    5 days ago
  •  ...Description:DataRobot delivers AI that maximizes impact and...  ...DataRobot’s Fleet team is the engine behind how our platform runs...  ...That’s where you come in.As a Staff Software Engineer, you’ll be responsible...  ...for training and inference.Why Join the Fleet Management... 
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    DataRobot

    Boston, MA
    2 days ago
  •  ...reliable on-demand, logistics engine for last-mile retail...  ...to join our team. As a Staff Machine Learning...  .... Proficiency in using AI coding tools (e.g., Claude...  ...Codex, Cursor) in the full software development lifecycle,...  ...applied ML for Causal Inference and Recommendation... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    Doordash

    San Francisco, CA
    5 days ago
  • $242k - $389k

     ..., WA / Remote (United States)Software – Software Systems /Full-time...  ...alongside a team of strong software engineers and act as a force multiplier...  ...cutting-edge ML Training OR Inference performance optimization...  ...use artificial intelligence (AI) tools to support parts of... 
    Full time
    Remote work

    Zoox

    Foster, CA
    5 days ago
  • $135k - $160k

     ...the ultimate goal of enabling human life on Mars.APPLICATION SOFTWARE ENGINEER, INFERENCEThe application software team is the central...  ...respect, and support.Our team maintains a high-performance AI inference platform that serves the best models internally at SpaceX to... 
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    22 hours ago
  • $207k - $301k

     ...machine earning pipelines for model training, inference, and integration with high-throughput Ad...  ...experience.8 years of experience in software development.5 years of experience...  ...or recommender systems.Google's software engineers develop the next-generation technologies... 

    Google

    Mountain View, CA
    3 days ago
  •  ...foundation of a neurosymbolic AI agent that learns more about...  ...Join Onton as a Founding Engineer and set the strategic foundation...  ...out a performant and scalable inference engine to support more...  ...and is passionate about making software tools accessible to all, we want... 
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    3 days per week

    Onton

    Remote
    1 day ago
  •  ...entertainment company and the leader in AI music. We are backed by...  ...team at the Senior and Staff levels. These aren’t maintenance...  ...features end-to-end, drive engineering quality, and contribute to architectural...  ...features, and real-time AI inference all live on this platform. If... 
    Full time
    Work at office
    Local area
    Worldwide

    Suno

    Remote
    1 day ago
  • $216k - $258k

     ...full potential of their data for AI applications. The platform...  ...infrastructure that power AI inference at 100K+ QPS with millisecond...  ...Evolve Tecton’s query execution engine to support complex, multi-stage...  ...distributed and/or highly concurrent software systems ~ Degree in Computer... 
    Full time
    Remote work
    Flexible hours

    Tecton

    Remote
    1 day ago
  • $190k - $210k

    Role Description Vanilla is seeking a Staff Software Engineer - AI Applications with a strong background in software development, data science,...  ...can build tooling to support model training, evaluation, inference serving, monitoring, and alerting. ~You want to use the... 
    Full time
    Work experience placement
    Work at office
    Home office
    Flexible hours

    Vanilla Technologies

    Remote
    1 day ago
  • $225k - $265k

     ...telemetry infrastructure for the AI era. At Cribl, we partner...  ...and grounded in a simple idea: software is a people business. Cribl...  ...and a group of highly-skilled engineers to shape the future of search...  ..., Prompt Engineering, and Inference Platforms ~This position will... 
    Full time
    Temporary work
    Remote work

    Cribl

    Remote
    2 hours ago
  •  ...acted on the same day. We're hiring a staff-level engineer to own that path end to end. It's a...  ...the orchestration layer for ingestion, inference, and downstream jobs. ~Prove the numbers...  ..., and Ruby. ~Set the bar on AI-assisted engineering. Who You Are... 
    Full time
    Temporary work
    Remote work
    Flexible hours

    ApartmentIQ

    Remote
    7 hours ago
  •  ...group of high-performance, high-octane engineers, so direction gets shaped together, and...  ...Work directly with founders, design, and inference teams to turn big ideas into interfaces...  ...and mentorship. Qualifications ~Staff-level depth in TypeScript and React.... 
    Full time
    Local area

    Venice.ai

    Remote
    2 hours ago
  • $265k

    Role Description We are seeking a Senior Staff Software Engineer to lead the architecture and evolution...  ...on which every product, workflow, and AI capability is built. You will define...  ...infrastructure for AI/LLM workloads: inference serving, agent runtimes, GPU scheduling... 
    Full time
    Remote work
    Work from home
    Flexible hours

    Juniper Square

    Remote
    7 days ago
  • $235.7k - $277k

     ...Description You'll help build Confluent Cloud's AI capabilities — the layer that lets...  ...moving data out to a separate system to run inference or build an agent, our customers do it in...  ...and process events at scale. As an engineer, you'll own delivery of significant pieces... 
    Full time
    Live in

    Confluent

    Remote
    3 days ago
  • $151.8k - $332.2k

    What you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition and model inference. In this role, you will develop state-of-the-art automatic speech recognition system and ship it to various Zoom products. You will work on... 
    Full time
    Work at office
    Remote work

    Zoom

    Seattle, WA
    5 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...groundbreaking AI training and inference possible.The Lambda Infrastructure Engineering organization forges the...  ...Role:We are seeking a seasoned Staff Storage Software Engineer with deep experience designing... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, AI Inference. Be the first to apply!