Senior AI Inference Library Engineer - GPU-Optimized
LeoForce
LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures, squeezing maximum performance while integrating with model-serving infrastructure. The role emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building cutting-edge AI software. #J-18808-Ljbffr LeoForce
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...Senior- ...worldwide.We’re a team of engineers, clinicians, and... ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be responsible... ...into performance optimized, robust, validated and scalable... ...models to real-time onboard inference—while serving as a core...SeniorLocal areaWorldwideFlexible hours
- MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this role, you will build and operate production inference systems, optimizing for performance and reliability. The ideal... ...have strong knowledge in GPU-accelerated inference. Excellent...Senior
$167.2k - $209k
.... DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team. DigitalOcean aims to be... ...at the inference engine and GPU kernel layers, ensuring our infrastructure... ...using AMD's AITER library for AMD MI355X - identify and...SeniorLocal areaRemote workWorldwideFlexible hours$250k
...Join a rapidly scaling AI cloud infrastructure... ...a next-generation GPU platform designed... ...experimentation, and inference at scale. The company... ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale... ...Support and optimize Slurm-based GPU cluster...SeniorFull timeRemote work- A cutting-edge AI technology company based in San Francisco is seeking a specialist... ...to design and operate large-scale GPU infrastructure. This role requires... ...GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands...Senior
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
$220k - $320k
A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation...SeniorLocal area- ...San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting... ...design techniques to improve latency and throughput, optimize the inference stack to exhaust hardware, and extend Kubernetes...Senior
- A leading AI technology company in San Francisco is looking for a Senior Software Engineer to build scalable infrastructure for large‑scale training and fine-tuning of foundation... ...design distributed training systems and optimize GPU utilization while collaborating with cross-...Senior
- A tech startup focused on AI workloads is seeking a Member of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and... ...Ideal candidates should have strong software engineering skills and experience with ML inference...Senior
$220k
...Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...Senior- OpenAI in San Francisco is seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime contexts. You will design the OpenAI-hosted control plane and partner-side software, focusing on reliable long...Senior
$175k - $250k
Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year... ...designed to support modern AI models across a variety... ...ROCm, Triton, or similar GPU/accelerator programming technologies... ..., integrating, or optimizing performance-critical...$315k
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking... .... Responsibilities include developing optimizations, collaborating with teams, and...- A leading AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for designing high-performance GPU kernels and using advanced techniques to boost computation efficiency. Ideal...
$175k - $250k
...us We're a well-funded AI infrastructure startup... ...We're looking for an engineer to help build and maintain... ...a high-performance inference library designed to support... ...ROCm, Triton, or similar GPU/accelerator... ...developing, integrating, or optimizing performance-critical compute...Local area- DigitalOcean is seeking a Senior Director of Engineering to lead a high‑performing team building and scaling our LLM inference products across control plane, optimization, and architecture. You will own Serverless Inference, Dedicated Inference, Inference Router, Batch...SeniorWorldwide
$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate has over 5 years of software engineering experience, strong familiarity with ML architectures, and experience...Senior- ...API. You will build global, low-latency GPU ML inference systems in the critical path of customer... ...and cost-efficiency for a revolutionary AI product. Join a 5-person team; work with... ...Terraform, and Docker to deploy, monitor, and optimize the entire infrastructure stack. #J-188...Senior
$160k - $230k
About the RoleAt Together.ai, we are building state-... ...efficient and scalable inference for large language... ...LLMs). Our mission is to optimize inference frameworks, algorithms... ...and Optimization Engineer to design, develop, and... ...-throughput inference, GPU/accelerator...Full time- MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and optimizations to ensure high throughput and low latency in production. You will collaborate with...Senior
$179k - $218k
...energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we... ..."Silicon Reality" must be bridged.We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority...SeniorTemporary work- nineDots.io is hiring a Senior Software Engineer to build secure GPU sandbox environments and scalable GPU compute platforms. You will help define architecture... ..., ensuring isolation, security, and reliability for AI agents. You’ll work across GPU virtualization, Linux...Senior
- An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure... ...researchers and product teams to push the boundaries of AI technology, ensuring reliable production services. If...
- Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from...
- Together AI is building state-of-the-art infrastructure to... ...enable efficient and scalable inference for large language models (... ...an Inference Frameworks and Optimization Engineer to design, develop, and optimize... ...high-throughput inference, GPU/accelerator optimizations,...
$161.3k - $241.9k
...combining frontier agentic AI, an enterprise-grade... ..., every model inference, and every production workload... ...looking for a Production Engineer to help build and... ...rightsizing, workload optimization, and utilization monitoring... ...Experience operating GPU fleets, high-performance...Senior- TensorScale AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing...
- STN Inc in San Francisco is seeking an experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will optimize GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work across storage, networking...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Library Engineer - GPU-Optimized. Be the first to apply!
- ai ml engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai developer San Francisco, CA
- ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- ai prompt engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- senior maintenance supervisor San Francisco, CA
- senior lead project manager San Francisco, CA


