Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Performance Engineer - GPU & CUDA

$220k - $320k

inference.net

inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions. #J-18808-Ljbffr inference.net

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Inference Performance Engineer - GPU & CUDA in San Francisco, CA vacancy
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation... 
    Senior
    Performance
    Local area

    Inference

    San Francisco, CA
    3 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...work helps care teams perform with greater...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics...  ...to real-time onboard inference—while serving as a core...  ...in GPU Compute API - CUDA, OpenCL• Proficiency... 
    Senior
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    3 days ago
  • $167.2k - $209k

     ...DigitalOcean is seeking a Senior Engineer 2 to play a key technical role in our AI Inference Optimization team....  ...the industry-leading performance for our inference services...  ...inference engine and GPU kernel layers,...  ...their software stacks (CUDA, ROCm, TensorRT, OpenAI... 
    Senior
    Performance
    Local area
    Remote work
    Worldwide
    Flexible hours

    DigitalOcean

    San Francisco, CA
    20 hours ago
  •  ...Auto is hiring an experienced GPU-focused engineer to advance autonomous...  ...will optimize end-to-end GPU performance, including sensor processing...  ...LiDAR) and neural network inference, and collaborate with software...  .... The role emphasizes CUDA-based development, profiling... 
    Performance
    Relocation

    Bot-Auto

    San Francisco, CA
    3 days ago
  • MakerMaker.AI is looking for a Senior Machine Learning Systems Engineer in San Francisco. In this...  ...build and operate production inference systems, optimizing for performance and reliability. The ideal candidate...  ...and have strong knowledge in GPU-accelerated inference.... 
    Senior
    Performance

    MakerMaker.AI

    San Francisco, CA
    20 hours ago
  • LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse...  ...emphasizes deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with exposure... 
    Senior
    Performance

    Leoforce

    San Francisco, CA
    4 days ago
  •  ...San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern... 
    Senior
    Performance

    Reflection AI

    San Francisco, CA
    4 days ago
  •  ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of... 
    Performance

    Baseten

    San Francisco, CA
    3 days ago
  • $250k

     ...building a next-generation GPU platform designed for AI...  ..., experimentation, and inference at scale. The company is...  ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Senior
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack...  ...You will primarily work on uzu, our inference engine, and focus on supporting new...  ...work, and experience in writing high-performance GPU kernels or Rust systems programming.... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    3 days ago
  • $190k - $235k

     ...Francisco, CaliforniaSoftware Engineering /Full time Exempt /...  ...you. ABOUT THE ROLEAs a senior Robot Perception...  ...learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT, and related acceleration...  ...)C/C++ experience for performance-critical... 
    Senior
    Performance
    Full time

    Bright Machines

    San Francisco, CA
    4 days ago
  • Inferact is seeking a hands-on cluster administration engineer to own and operate its high-performance GPU compute infrastructure. You will ensure health, availability, observability, and usability around the clock for HP compute clusters across providers. You will collaborate... 
    Senior
    Performance

    Inferact

    San Francisco, CA
    3 days ago
  • SF Tensor in San Francisco is hiring a Member of Technical Staff for GPU Kernel Engineering to push the limits of what the hardware can do before any search begins. You will hand-write and optimize kernels to achieve unprecedented throughput, profile at the microarchitectural... 
    Senior
    Performance
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    4 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Senior
    Performance

    Causal Labs

    San Francisco, CA
    2 days ago
  •  ...growing infrastructure company in San Francisco hire a Senior Software Engineer to build and operate GPU-backed sandboxes for AI agents. You will own the...  ...lifecycle from architecture to deployment, with focus on performance and security. The role demands deep Linux... 
    Senior
    Performance

    Recruiting from Scratch

    San Francisco, CA
    20 hours ago
  • Inferact Inc. is building a world-class GPU compute platform powering vLLM. We seek a hands-on cluster administration engineer to own the high-performance infrastructure that keeps our engineers productive. You will monitor GPU servers, manage scheduling with SLURM/Kubernetes... 
    Senior
    Performance
    Remote work

    Inferact Inc.

    San Francisco, CA
    1 day ago
  • $220k

    Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience... 
    Senior

    Perplexity

    San Francisco, CA
    3 days ago
  •  ...of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems... 
    Senior
    Performance

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • $160k - $230k

     ...efficient and scalable inference for large language models...  ...the boundaries of performance, scalability, and cost-...  ...Frameworks and Optimization Engineer to design, develop, and...  ...-throughput inference, GPU/accelerator...  ...performance serving.Apply CUDA graph optimizations, TensorRT... 
    Performance
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  • San Francisco Tensor Company is seeking a Member of Technical Staff for GPU Compiler Engineering to build the machine that searches and to optimize the entire pipeline from StableHLO ingestion to backend code generation. You will work on MLIR dialects, LLVM backend, and... 
    Senior

    San Francisco Tensor Company

    San Francisco, CA
    4 days ago
  •  ...seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime...  ...software, focusing on reliable long-running workflows, reproducible performance results, and secure handling of model and hardware data. This... 
    Senior
    Performance

    Slope

    San Francisco, CA
    20 hours ago
  •  ...intelligence. We are seeking a Software Engineer to join Crusoe’s Data Center Infrastructure...  ...on software for managing a fleet of GPU servers and the data centers that house...  ...automation and repair tooling for high‑performance GPU compute clusters. The successful candidate... 
    Senior
    Performance

    Crusoe

    San Francisco, CA
    4 days ago
  • An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure...  ...for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate... 
    Performance

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 
    Performance

    Anthropic

    San Francisco, CA
    3 days ago
  •  ...company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure)...  ...role involves optimizing APIs, managing GPU workloads, and collaborating with...  ...with TypeScript/Node, strong skills in performance tuning and distributed systems. This position... 
    Senior
    Performance

    Vizcom

    San Francisco, CA
    2 days ago
  •  ...experienced AI Infrastructure Engineer to design, deploy, and manage large-scale GPU clusters for AI training and inference workloads. You will...  ...GPU utilization, tune NCCL, CUDA, UCX, and Slurm, and work...  ...and software layers to push performance and scalability. This role... 
    Senior
    Performance

    STN Inc

    San Francisco, CA
    20 hours ago
  • Sciforium is seeking a GPU Kernel Engineer to push performance on modern accelerators. You will design and optimize...  ...used for large-scale training and inference. Ideal candidates have 5+ years in...  ...strong C++/Python skills, and deep CUDA/ROCm expertise. Collaboration with... 
    Performance

    Sciforium

    San Francisco, CA
    3 days ago
  •  ...build and operate the inference systems that serve...  .... This is an engineering role, not a research...  ...Own the performance characteristics of...  ...WE'RE LOOKING FOR Senior ML systems engineer...  ...Experience with GPU‑accelerated inference...  ...following languages: C++, CUDA, ROCm or Triton... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  • $100k - $120k

     ...foundation models. As training and inference workloads grow, we need kernel‑level...  ...Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement, and...  ...kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators... 
    Performance

    Coda Robotics

    San Francisco, CA
    1 day ago
  • $175k - $250k

    Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $250,000 per year Job Details...  ...build and maintain a high-performance inference library designed to...  ...Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies... 
    Performance

    LeoForce

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Performance Engineer - GPU & CUDA. Be the first to apply!