Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Performance Engineer - GPU & CUDA

$220k - $320k

inference.net

inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions. #J-18808-Ljbffr inference.net

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Inference Performance Engineer - GPU & CUDA in San Francisco, CA vacancy
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...work helps care teams perform with greater...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics...  ...to real-time onboard inference—while serving as a core...  ...in GPU Compute API - CUDA, OpenCL• Proficiency... 
    Senior
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  •  ...San Francisco is seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the...  ...requires deep hands-on experience with inference engines such as SGLang, vLLM or TensorRT, GPU architectures, and software stacks using #J... 
    Senior
    Performance

    NEAR AI

    San Francisco, CA
    2 days ago
  • $250k

     ...building a next-generation GPU platform designed for AI...  ..., experimentation, and inference at scale. The company is...  ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Senior
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern... 
    Senior
    Performance

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of... 
    Performance

    Baseten

    San Francisco, CA
    2 days ago
  • $190k - $235k

     ...Francisco, CaliforniaSoftware Engineering /Full time Exempt /...  ...you. ABOUT THE ROLEAs a senior Robot Perception...  ...learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT, and related acceleration...  ...)C/C++ experience for performance-critical... 
    Senior
    Performance
    Full time

    Bright Machines

    San Francisco, CA
    3 days ago
  •  ...building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You...  ...and drive reliability, observability, and performance across workloads. You’ll manage the... 
    Senior
    Performance

    Neura Market

    San Francisco, CA
    4 days ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack...  ...You will primarily work on uzu, our inference engine, and focus on supporting new...  ...work, and experience in writing high-performance GPU kernels or Rust systems programming.... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    2 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Senior
    Performance

    Causal Labs

    San Francisco, CA
    14 hours ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 
    Senior

    Perplexity

    San Francisco, CA
    2 days ago
  •  ...of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems... 
    Senior
    Performance

    Gimlet Labs

    San Francisco, CA
    14 hours ago
  • $160k - $230k

     ...efficient and scalable inference for large language models...  ...the boundaries of performance, scalability, and cost-...  ...Frameworks and Optimization Engineer to design, develop, and...  ...-throughput inference, GPU/accelerator...  ...performance serving.Apply CUDA graph optimizations, TensorRT... 
    Performance
    Full time

    Together AI

    San Francisco, CA
    1 day ago
  •  ...seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime...  ...software, focusing on reliable long-running workflows, reproducible performance results, and secure handling of model and hardware data. This... 
    Senior
    Performance

    Slope

    San Francisco, CA
    5 days ago
  •  ...seeking an expert in high‑performance LLM serving systems and inference optimization. In this...  ...optimizing major inference engines such as SGLang, vLLM, or...  ...knowledge of state‑of‑the‑art GPU architectures, and...  ...using PyTorch, Triton, CuTe, CUDA, etc. Proven track record... 
    Performance

    NEAR.AI

    San Francisco, CA
    1 day ago
  • $300k

     ...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal...  ..., and scheduling Writing and tuning custom CUDA / Triton kernels for performance-critical paths... 
    Performance
    Relocation
    Visa sponsorship
    Free visa

    techire ai

    San Francisco, CA
    3 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure...  ...for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate... 
    Performance

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $100k - $150k

     ...vertically integrated AI cloud engineered for AI. We own and...  ...energy, data centres, GPU superclusters,...  ...services — delivering high-performance infrastructure to AI-...  ...(Job Purpose) Senior Infrastructure Support...  ...stacks on AI training and inference clusters. Confident... 
    Senior
    Performance
    Full time
    Remote work
    Flexible hours

    Nscale

    San Francisco, CA
    2 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 
    Performance

    Anthropic

    San Francisco, CA
    2 days ago
  • $250k

     ...building a serverless inference platform, beginning...  ...chance to join as a Senior Inference Platform Engineer at an early stage...  ...systems to maximise GPU utilisation and minimise...  ...best practices in performance and efficiency. Implement...  ...software stacks (CUDA, Triton, NCCL) and... 
    Senior
    Performance
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...build and operate the inference systems that serve...  .... This is an engineering role, not a research...  ...Own the performance characteristics of...  ...WE'RE LOOKING FOR Senior ML systems engineer...  ...Experience with GPU‑accelerated inference...  ...following languages: C++, CUDA, ROCm or Triton... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    5 days ago
  • $100k - $120k

     ...foundation models. As training and inference workloads grow, we need kernel‑level...  ...Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement, and...  ...kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators... 
    Performance

    Coda Robotics

    San Francisco, CA
    5 days ago
  • $250k

    A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines....  ...infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include strong skills in Go and Python,... 
    Senior
    Performance

    Acceler8 Talent

    San Francisco, CA
    2 days ago
  •  ...and unicorn founders and senior engineers with deep expertise in 3...  ...a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This is...  ...using torch.compile, custom CUDA kernels, and specialized...  ...Working knowledge of GPU hardware (NVIDIA) and... 
    Performance
    Relocation
    Visa sponsorship
    Relocation package

    Reactor.am

    San Francisco, CA
    4 days ago
  • $325k

    A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments...  ...role involves collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $490K. #J-1880... 
    Senior
    Performance

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our...  ...builds high‑performance GPU kernels and custom libraries...  ...heart of our on‑vehicle ML inference for ADAS and autonomous driving...  ..., benchmark, and iterate on CUDA-based kernels and custom operators... 
    Senior
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    5 days ago
  •  ...shipping excellence. We seek engineers with strong intrinsic drive,...  ...experience to help scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding...  ...or LA offices Tech Stack CUDA/C++, GPGPU, Python, Linux Key... 
    Performance
    Full time
    Work at office

    Vast.ai Inc.

    San Francisco, CA
    14 hours ago
  •  ...boundaries of what's possible in video generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and...  ...solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our... 
    Performance

    Genmo

    San Francisco, CA
    1 day ago
  • Unity3D in San Francisco is seeking a Senior Graphics Engineer to enhance Unity's Metal graphics pipeline. You will develop and maintain a high-performance API that optimizes the graphics performance on Apple devices. This role requires over 10 years of C++ programming... 
    Senior
    Performance

    unity3d

    San Francisco, CA
    5 days ago
  • $160k - $200k

     ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems...  .... Improve model performance. Research and implement...  ...validation, release gates, inference wrappers, ONNX/TensorRT...  ...with ONNX, TensorRT, CUDA, quantization, model profiling... 
    Senior
    Performance
    Full time

    Simbe Robotics Inc

    San Francisco, CA
    1 day ago
  •  ...infrastructure company is seeking a Senior Engineer 2 to join their AI Inference Optimization team. The role...  ...leading the technical strategy for performance architecture and addressing complex...  ...computing and a strong understanding of GPU architectures. The position offers... 
    Senior
    Performance
    Remote work

    DigitalOcean

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Performance Engineer - GPU & CUDA. Be the first to apply!