Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Library Engineer - GPU-Optimized

LeoForce

LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures, squeezing maximum performance while integrating with model-serving infrastructure. The role emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building cutting-edge AI software. #J-18808-Ljbffr LeoForce

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Library Engineer - GPU-Optimized in San Francisco, CA vacancy
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Senior

    inference.net

    San Francisco, CA
    2 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be responsible...  ...into performance optimized, robust, validated and scalable...  ...models to real-time onboard inference—while serving as a core... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    21 hours ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal...  ...have strong knowledge in GPU-accelerated inference. Excellent... 
    Senior

    MakerMaker.AI

    San Francisco, CA
    4 days ago
  • $167.2k - $209k

     .... DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be...  ...at the inference engine and GPU kernel layers, ensuring our infrastructure...  ...using AMD's AITER library for AMD MI355X - identify and... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    3 days ago
  • $250k

     ...Join a rapidly scaling AI cloud infrastructure...  ...a next-generation GPU platform designed...  ...experimentation, and inference at scale. The company...  ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale...  ...Support and optimize Slurm-based GPU cluster... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist...  ...to design and operate large-scale GPU infrastructure. This role requires...  ...GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands... 
    Senior

    Reflection AI

    San Francisco, CA
    2 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    5 days ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation... 
    Senior
    Local area

    Inference

    San Francisco, CA
    1 day ago
  •  ...San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting...  ...design techniques to improve latency and throughput, optimize the inference stack to exhaust hardware, and extend Kubernetes... 
    Senior

    Causal Labs

    San Francisco, CA
    5 days ago
  • A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation...  ...design distributed training systems and optimize GPU utilization while collaborating with cross-... 
    Senior

    Baseten

    San Francisco, CA
    2 days ago
  • A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and...  ...Ideal candidates should have strong software engineering skills and experience with ML inference... 
    Senior

    Gimlet Labs

    San Francisco, CA
    5 days ago
  • $220k

     ...Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 
    Senior

    Perplexity

    San Francisco, CA
    1 day ago
  • OpenAI in San Francisco is seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime contexts. You will design the OpenAI-hosted control plane and partner-side software, focusing on reliable long... 
    Senior

    Slope

    San Francisco, CA
    4 days ago
  • $175k - $250k

    Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year...  ...designed to support modern AI models across a variety...  ...ROCm, Triton, or similar GPU/accelerator programming technologies...  ..., integrating, or optimizing performance-critical... 

    LeoForce

    San Francisco, CA
    2 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking...  .... Responsibilities include developing optimizations, collaborating with teams, and... 

    Anthropic

    San Francisco, CA
    1 day ago
  • A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal... 

    Baseten

    San Francisco, CA
    2 days ago
  • $175k - $250k

     ...us We're a well-funded AI infrastructure startup...  ...We're looking for an engineer to help build and maintain...  ...a high-performance inference library designed to support...  ...ROCm, Triton, or similar GPU/accelerator...  ...developing, integrating, or optimizing performance-critical compute... 
    Local area

    Jobot

    San Francisco, CA
    2 days ago
  • DigitalOcean is seeking a Senior Director of Engineering to lead a high‑performing team building and scaling our LLM inference products across control plane, optimization, and architecture. You will own Serverless Inference, Dedicated Inference, Inference Router, Batch... 
    Senior
    Worldwide

    DigitalOcean

    San Francisco, CA
    1 day ago
  • $325k

    A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience... 
    Senior

    Jobleads-US

    San Francisco, CA
    2 days ago
  •  ...API. You will build global, low-latency GPU ML inference systems in the critical path of customer...  ...and cost-efficiency for a revolutionary AI product. Join a 5-person team; work with...  ...Terraform, and Docker to deploy, monitor, and optimize the entire infrastructure stack. #J-188... 
    Senior

    Jack & Jill

    San Francisco, CA
    4 days ago
  • $160k - $230k

    About the RoleAt Together.ai, we are building state-...  ...efficient and scalable inference for large language...  ...LLMs). Our mission is to optimize inference frameworks, algorithms...  ...and Optimization Engineer to design, develop, and...  ...-throughput inference, GPU/accelerator... 
    Full time

    Together AI

    San Francisco, CA
    5 days ago
  • MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and optimizations to ensure high throughput and low latency in production. You will collaborate with... 
    Senior

    MakerMaker

    San Francisco, CA
    5 days ago
  • $179k - $218k

     ...energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we...  ..."Silicon Reality" must be bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • nineDots.io is hiring a Senior Software Engineer to build secure GPU sandbox environments and scalable GPU compute platforms. You will help define architecture...  ..., ensuring isolation, security, and reliability for AI agents. You’ll work across GPU virtualization, Linux... 
    Senior

    nineDots.io

    San Francisco, CA
    3 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure...  ...researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If... 

    Jobleads-US

    San Francisco, CA
    2 days ago
  • Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from... 

    Jobot

    San Francisco, CA
    2 days ago
  • Together AI is building state-of-the-art infrastructure to...  ...enable efficient and scalable inference for large language models (...  ...an Inference Frameworks and Optimization Engineer to design, develop, and optimize...  ...high-throughput inference, GPU/accelerator optimizations,... 

    Together AI

    San Francisco, CA
    2 days ago
  • $161.3k - $241.9k

     ...combining frontier agentic AI, an enterprise-grade...  ..., every model inference, and every production workload...  ...looking for a Production Engineer to help build and...  ...rightsizing, workload optimization, and utilization monitoring...  ...Experience operating GPU fleets, high-performance... 
    Senior

    Neura Market

    San Francisco, CA
    2 days ago
  • TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing... 

    TensorScale AI

    San Francisco, CA
    2 days ago
  • STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking... 
    Senior

    STN Inc

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Library Engineer - GPU-Optimized. Be the first to apply!