Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Performance Engineer — GPU & CUDA

$220k - $320k

Inference

A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation of $220,000 to $320,000 plus equity and benefits, focusing on accelerating AI inference. Join an innovative team in downtown San Francisco, hybrid options available for local candidates. #J-18808-Ljbffr Inference

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Inference Performance Engineer — GPU & CUDA in San Francisco, CA vacancy
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Senior
    Performance

    inference.net

    San Francisco, CA
    4 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...work helps care teams perform with greater...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics...  ...to real-time onboard inference—while serving as a core...  ...in GPU Compute API - CUDA, OpenCL• Proficiency... 
    Senior
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  • $167.2k - $209k

     ...DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team....  ...the industry-leading performance for our inference services...  ...inference engine and GPU kernel layers,...  ...their software stacks (CUDA, ROCm, TensorRT, OpenAI... 
    Senior
    Performance
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    19 hours ago
  •  ...Inc. is seeking a systems engineer with HPC or parallel programming...  ...to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks...  ...push the bleeding edge of AI performance. This role is based on-site... 
    Performance

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  •  ...San Francisco is seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the...  ...requires deep hands-on experience with inference engines such as SGLang, vLLM or TensorRT, GPU architectures, and software stacks using #J... 
    Senior
    Performance

    NEAR AI

    San Francisco, CA
    3 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will leverage your HPC background...  ...to push the bleeding edge of AI, working with CUDA/C++ and a modern tech stack. Ideal candidates have... 
    Performance
    Full time

    Vast.ai

    San Francisco, CA
    2 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this...  ...build and operate production inference systems, optimizing for performance and reliability. The ideal candidate...  ...and have strong knowledge in GPU-accelerated inference.... 
    Senior
    Performance

    MakerMaker.AI

    San Francisco, CA
    19 hours ago
  •  ...San Francisco is seeking a senior ML infrastructure engineer to design and optimize...  ...distributed training systems and performance-critical components. You...  ...implement low‑level code (CUDA, Triton) and ensure...  ...utilization across multi‑node GPU clusters. Join a team focused... 
    Senior
    Performance

    GenesisAI

    San Francisco, CA
    5 days ago
  • $120k - $170k

     ...vertically integrated AI cloud engineered for AI. We own and...  ...energy, data centres, GPU superclusters,...  ...services — delivering high-performance infrastructure to AI-...  ...(Job Purpose) Senior Infrastructure Support...  ...stacks on AI training and inference clusters. Confident... 
    Senior
    Performance
    Full time
    Remote work
    Flexible hours

    Nscale

    San Francisco, CA
    2 days ago
  • nineDots.io is hiring a Senior Software Engineer to build secure GPU sandbox environments and scalable GPU compute...  ...drive the stack from KVM/QEMU to CUDA lifecycle, ensuring isolation, security...  ...software foundations and performance-critical paths in a fast-growing AI... 
    Senior
    Performance

    nineDots.io

    San Francisco, CA
    5 days ago
  • $250k

     ...building a next-generation GPU platform designed for AI...  ..., experimentation, and inference at scale. The company is...  ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Senior
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern... 
    Senior
    Performance

    Reflection AI

    San Francisco, CA
    4 days ago
  •  ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of... 
    Performance

    Baseten

    San Francisco, CA
    3 days ago
  • $190k - $235k

     ...Francisco, CaliforniaSoftware Engineering /Full time Exempt /...  ...you. ABOUT THE ROLEAs a senior Robot Perception...  ...learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT, and related acceleration...  ...)C/C++ experience for performance-critical... 
    Senior
    Performance
    Full time

    Bright Machines

    San Francisco, CA
    4 days ago
  •  ...building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You...  ...and drive reliability, observability, and performance across workloads. You’ll manage the... 
    Senior
    Performance

    Neura Market

    San Francisco, CA
    5 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack...  ...You will primarily work on uzu, our inference engine, and focus on supporting new...  ...work, and experience in writing high-performance GPU kernels or Rust systems programming.... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    3 days ago
  • $200k - $300k

     ...from AI research to systems engineering to product design —...  ...We are looking for a Performance Engineer to make World Labs...  ...Will Do: Optimize inference and serving end to end — latency...  ...scale. Write and tune GPU kernels (CUDA, Triton) for hot paths; drive... 
    Performance
    Full time

    World Labs

    San Francisco, CA
    2 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Senior
    Performance

    Causal Labs

    San Francisco, CA
    1 day ago
  • MakerMaker in San Francisco is seeking a Senior ML systems engineer to build and operate production inference systems for large models. You will own performance, profiling, and optimizations to ensure high throughput and low latency in production. You will collaborate with... 
    Senior
    Performance

    MakerMaker

    San Francisco, CA
    2 days ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 
    Senior

    Perplexity

    San Francisco, CA
    3 days ago
  •  ...of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems... 
    Senior
    Performance

    Gimlet Labs

    San Francisco, CA
    1 day ago
  • $160k - $230k

     ...efficient and scalable inference for large language models...  ...the boundaries of performance, scalability, and cost-...  ...Frameworks and Optimization Engineer to design, develop, and...  ...-throughput inference, GPU/accelerator...  ...performance serving.Apply CUDA graph optimizations, TensorRT... 
    Performance
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  •  ...seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime...  ...software, focusing on reliable long-running workflows, reproducible performance results, and secure handling of model and hardware data. This... 
    Senior
    Performance

    Slope

    San Francisco, CA
    1 day ago
  •  ...seeking an expert in high‑performance LLM serving systems and inference optimization. In this...  ...optimizing major inference engines such as SGLang, vLLM, or...  ...knowledge of state‑of‑the‑art GPU architectures, and...  ...using PyTorch, Triton, CuTe, CUDA, etc. Proven track record... 
    Performance

    NEAR.AI

    San Francisco, CA
    2 days ago
  •  ...company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure)...  ...role involves optimizing APIs, managing GPU workloads, and collaborating with...  ...with TypeScript/Node, strong skills in performance tuning and distributed systems. This position... 
    Senior
    Performance

    Vizcom

    San Francisco, CA
    2 days ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal...  ..., and scheduling Writing and tuning custom CUDA / Triton kernels for performance-critical paths... 
    Performance
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    4 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure...  ...for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate... 
    Performance

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 
    Performance

    Anthropic

    San Francisco, CA
    3 days ago
  • Luminal is hiring a Founding Compiler Engineer for an on-site role in downtown San Francisco. You will help shape the core compiler, implement CUDA kernels, and review model performance to accelerate AI pipelines. As a founding team member, you will contribute to production... 
    Performance

    Slope

    San Francisco, CA
    19 hours ago
  • $250k

     ...building a serverless inference platform, beginning...  ...chance to join as a Senior Inference Platform Engineer at an early stage...  ...systems to maximise GPU utilisation and minimise...  ...best practices in performance and efficiency. Implement...  ...software stacks (CUDA, Triton, NCCL) and... 
    Senior
    Performance
    Full time
    San Francisco, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Performance Engineer — GPU & CUDA. Be the first to apply!