Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Inference Runtime Engineer for LLMs & Diffusion

Inferact

Inferact is seeking an inference runtime engineer to enhance the performance and capabilities of LLM and diffusion model serving. This role requires expertise in optimizing model execution on various hardware architectures and has significant implications for AI inference. The ideal candidate must possess a bachelor's degree in computer science or related fields, strong programming skills in Python, and experience with LLM inference systems.Remote work options are available for exceptional candidates. #J-18808-Ljbffr Inferact

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Inference Runtime Engineer for LLMs & Diffusion in San Francisco, CA vacancy
  • $160k - $230k

     ...to enable efficient and scalable inference for large language models (LLMs). Our mission is to optimize inference...  ...Frameworks and Optimization Engineer to design, develop, and optimize distributed...  ...architectures and LLM/VLM/Diffusion model optimization.Knowledge of inference... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  • Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in production. You will build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability... 
    Suggested

    Inception

    San Francisco, CA
    4 days ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ...improve GPU efficiency via profiling, runtime tuning, and server-level optimizations.... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  •  ...unicorn founders and senior engineers with deep expertise in 3D, generative...  ...for a Founding Engineer, ML Inference with deep expertise in high-...  ...-time model performance for diffusion models Design and implement...  ...in-house inference runtime Implement optimizations using... 
    Suggested
    Relocation
    Visa sponsorship
    Relocation package

    Reactor

    San Francisco, CA
    14 hours ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more cost-efficient, and more reliable. You will extend orchestration frameworks (Kubernetes... 

    Inception

    San Francisco, CA
    4 days ago
  • OpenAI in San Francisco is seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime contexts. You will design the OpenAI-hosted control plane and partner-side software, focusing on reliable long... 

    Slope

    San Francisco, CA
    1 day ago
  •  ...Jobtailor in San Francisco, CA is seeking a senior software engineer to design and build scalable AI-enabled runtime systems. You will own features end-to-end, from...  ...in distributed systems, AI applications with LLMs, and strong backend fundamentals, with a bias toward... 

    Jobtailor

    San Francisco, CA
    18 hours ago
  • $206.3k - $388k

     ...re looking for a Principal ML Engineer to architect and scale the...  ...training-ready data Scale up inference throughput across the pipeline...  ...inference optimization for VLMs and LLMs and data curation for...  ...across distributed and ML-centric runtime environments. Deep knowledge... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 days ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative. You will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm... 

    Kindredventures

    San Francisco, CA
    4 days ago
  • $225k

    Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate... 

    Dormont Manufacturing Co

    San Francisco, CA
    3 days ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control...  ..., and explore KV caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep observability, tracing... 

    Sail

    San Francisco, CA
    3 days ago
  • Sail Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible for building high-performance schedulers and optimizing global routing while focusing... 

    Sail Research

    San Francisco, CA
    4 days ago
  • $190.9k - $232.8k

     ...85About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development...  ...full GenAI inference stack: kernels, runtimes, orchestration, memory, and...  ...serving stack optimized for large-scale LLMs inferencePartner closely with researchers... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...ABOUT THE ROLE You build and operate the inference systems that serve our models in...  ...The work spans serving infrastructure, runtime optimization, and the long tail of production...  ...with running real workloads. This is an engineering role, not a research role. You'll measure... 

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • $300 per month

     ...We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware Systems...  ...characterization studies across training and inference - dense, MoE, long-context, and...  ...frameworks, training frameworks, or ML compiler/runtime stacks.Familiarity with both x86 and ARM... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...AI workloads, covering training, deployment, observation, and inference. You will perform hands-on inference research, selecting high-...  ...workloads. You will work with customers alongside Forward Deployed Engineers to deploy and tune models, while expanding collaborations with... 

    Mixpeek

    San Francisco, CA
    1 day ago
  • Together AI in San Francisco is seeking a Staff ML Systems Engineer to design and prototype algorithms, architectures, and scheduling for low-latency, high-throughput inference. You will implement changes in production-grade inference engines, including kernel backends... 

    Together

    San Francisco, CA
    3 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates... 
    Full time

    Vast.ai

    San Francisco, CA
    2 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 

    Causal Labs

    San Francisco, CA
    2 days ago
  •  ...infrastructure, influencing latency, throughput, and reliability of RL and training loops. You will own the infrastructure enabling fast inference and scalable RL iteration, balancing KV-cache strategies, batching, and long-context workloads while collaborating with #J-18808-... 

    Magic AI, Inc

    San Francisco, CA
    4 days ago
  • $170k - $245k

     ...to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly... 
    Work at office

    Anyscale

    San Francisco, CA
    14 hours ago
  • $250k - $300k

     ...reliably in production. That means owning the inference stack end to end: profiling where time...  ...will also work directly with customer engineering teams to tailor deployments to their...  ...Familiarity with methods for optimizing LLMs for high throughput / low latency inference... 
    Temporary work

    Crusoe

    San Francisco, CA
    14 hours ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models...  ...world. We're a small, fast-moving team of engineers focused on delivering a world-class...  ...building and scaling inference systems for LLMs or multimodal models. Have worked with... 
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ...Stack team builds the distributed runtime that powers large-scale LLM inference across... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  •  ...models at massive scale. As part of the inference team, you’ll be responsible for unlocking...  ...We are looking for a kernel-focused engineer to lead efforts in writing, porting, and...  ...to and extend internal GPU libraries and runtime tools. Work closely with hardware-specific... 
    Full time

    OpenAI

    San Francisco, CA
    14 hours ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. At Baseten...  ...Serverless-Grade Startup Speeds for LLMs: You will work deeply with checkpointing... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    14 hours ago
  • $300k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the Role The Cloud Inference team scales and optimizes Claude to serve...  ...across all platforms, and ensure our LLMs meet rigorous safety, performance, and security... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    14 hours ago
  •  ...California, who is eager to contribute to developer ecosystems and work with cutting-edge technologies like LLMs and AI tooling. This role is suited for early‑career engineers with a builder mindset. You will collaborate with the engineering team, maintain open-source projects... 

    Stealth Startup

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Inference Runtime Engineer for LLMs & Diffusion. Be the first to apply!