Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Inference Performance Engineer - GPU & CUDA

$220k - $320k

inference.net

inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions. #J-18808-Ljbffr inference.net

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Inference Performance Engineer - GPU & CUDA in San Francisco, CA vacancy
  • $250k

     ...building a next-generation GPU platform designed for AI...  ..., experimentation, and inference at scale. The company is...  ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large...  ...reliability, scalability, and performance of HPC and cloud... 
    Senior
    Performance
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for...  ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of... 
    Performance

    Baseten

    San Francisco, CA
    1 day ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack...  ...You will primarily work on uzu, our inference engine, and focus on supporting new...  ...work, and experience in writing high-performance GPU kernels or Rust systems programming.... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    8 hours ago
  • SF Tensor in San Francisco is hiring a Member of Technical Staff for GPU Kernel Engineering to push the limits of what the hardware can do before any search begins. You will hand-write and optimize kernels to achieve unprecedented throughput, profile at the microarchitectural... 
    Senior
    Performance
    Relocation package

    San Francisco Tensor Company

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...artificial intelligence, high-performance computing, and specialized...  ...We're looking for an engineer to help build and maintain a high-performance inference library designed to support...  ...Hands-on experience with CUDA, ROCm, Triton, or similar GPU/accelerator programming technologies... 
    Performance
    Local area

    Jobot

    San Francisco, CA
    3 days ago
  • $315k

    A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams... 
    Performance

    Anthropic

    San Francisco, CA
    8 hours ago
  • $190k - $250k

     ...hands-on support from AMD engineers the team is scaling...  ...a highly skilled GPU Kernel Engineer who is...  ...pushing the limits of performance on modern accelerators...  ...large-scale training and inference. This role is ideal...  ...using C++, PTX, CUDA, ROCm, Triton, and/or... 
    Performance
    Full time
    Flexible hours

    Sciforium

    San Francisco, CA
    2 days ago
  • Jobot seeks a highly skilled AI/ML infrastructure engineer to design and optimize a high-performance inference library supporting modern AI models across...  ...Python and systems languages (C++, Rust) and GPU programming (CUDA, ROCm, Triton) to push peak performance. Join... 
    Senior
    Performance

    Jobot

    San Francisco, CA
    1 day ago
  • Member of Technical Staff, Inference Performance 200k base + equity I’m working...  ...path from model and serving engine through kernels, accelerators...  ...decoding Developing and tuning CUDA, HIP, or Triton kernels Operating...  ...or model-serving systems GPU kernel development using CUDA... 
    Performance

    Acceler8 Talent

    San Francisco, CA
    3 days ago
  • $100k - $120k

     ...foundation models. As training and inference workloads grow, we need kernel‑level...  ...Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement, and...  ...kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators... 
    Performance

    Coda Robotics

    San Francisco, CA
    3 days ago
  •  ...build and operate the inference systems that serve...  .... This is an engineering role, not a research...  ...Own the performance characteristics of...  ...WE'RE LOOKING FOR Senior ML systems engineer...  ...Experience with GPU‑accelerated inference...  ...following languages: C++, CUDA, ROCm or Triton... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    3 days ago
  • $250k

     ...building a serverless inference platform, beginning...  ...chance to join as a Senior Inference Platform Engineer at an early stage...  ...systems to maximise GPU utilisation and minimise...  ...best practices in performance and efficiency. Implement...  ...software stacks (CUDA, Triton, NCCL) and... 
    Senior
    Performance
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...Senior Systems Engineer San Francisco, California Onsite or Remote At Evidently, we are raising...  ...and compliance posture, database performance, AI inference infrastructure: you can cover these...  ...serving platforms (vLLM, Cloud Run, GPU provisioning), latency SLOs, cost... 
    Senior
    Performance
    Remote work
    Work from home

    Evidently

    San Francisco, CA
    3 days ago
  • Unity3D in San Francisco is seeking a Senior Graphics Engineer to enhance Unity's Metal graphics pipeline. You will develop and maintain a high-performance API that optimizes the graphics performance on Apple devices. This role requires over 10 years of C++ programming... 
    Senior
    Performance

    unity3d

    San Francisco, CA
    3 days ago
  • $190k - $235k

     ...Francisco, CaliforniaSoftware Engineering /Full time Exempt /...  ...you. ABOUT THE ROLEAs a senior Robot Perception...  ...learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT, and related acceleration...  ...)C/C++ experience for performance-critical... 
    Senior
    Performance
    Full time

    Bright Machines

    San Francisco, CA
    1 day ago
  • $182k - $242k

     ...combines superior infrastructure performance with deep technical...  ...Do: We're looking for a Senior Storage Engineer, File & Block to help build...  ...our largest AI training and inference workloads and internal stateful...  ...storage from a follower of GPU growth into a multiplier of... 
    Senior
    Performance
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    6 days ago
  • $205.9k - $407.5k

     ...looking to bring on a Senior Principal ML GPU Architect to lead...  ...the Director of ML Engineering who is responsible for...  ...changes in training and inference speed/scale for all...  ...backward passes in CUDA/CuTe. Write...  ...with FP8. Quality and performance analysis between data... 
    Senior
    Performance

    Adobe

    San Francisco, CA
    2 days ago
  •  ...neocloud for alternative chips. Inference is fragmenting: purpose-built...  ...5-7× faster than existing GPU-based competitors. Our...  ...just the backlog. As a founding engineer, you'll help decide what we build...  ...deep enough to reason about performance at the hardware level, not just... 
    Performance

    General Compute

    San Francisco, CA
    3 days ago
  • Distributed LLM Inference Engineer At Anyscale, we're on a mission to democratize distributed computing...  ...that push the boundaries of performance for inference at large scale. This is an...  ...MLIR) Prior experience working on GPUs / CUDA Compensation At Anyscale, we take a market... 
    Performance
    Work at office

    Anyscale

    San Francisco, CA
    18 hours ago
  • $106.9k - $200.6k

     ...We are seeking an AI Systems Engineer to own the delivery, model-...  ...gapped environments. Works with senior engineers to test and...  ...delivery pipelines, operating high-performance inference (GPUs, model servers,...  ...Kafka sink), DCGM exporter for GPU telemetry, processor batching... 
    Senior
    Performance
    Full time
    Summer holiday
    Flexible hours

    EY

    San Francisco, CA
    2 days ago
  • $188k - $275k

     ...combines superior infrastructure performance with deep technical expertise...  ...You'll Do: The Field Engineering organization at CoreWeave is...  ...that customers can train and inference on at scale, spanning infrastructure...  ...lifecycle: leading new GPU cluster bring-up and... 
    Senior
    Performance
    Permanent employment
    Full time
    Contract work
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    San Francisco, CA
    18 days ago
  • $119.8k - $234.7k

     ...Senior Software Engineer, Forward DeployedWindows is where most of the world's developers build, and...  ...them. Coding agents, local inference, containerized toolchains and AI-native...  ...configuration at scale, toolchain and build performance, and the security, identity and policy... 
    Senior
    Performance
    Ongoing contract
    Work at office
    Local area
    Shift work
    3 days per week

    Microsoft Corporation

    San Francisco, CA
    1 day ago
  • $170k - $240k

     ...what makes them reliable. As a Perception Engineer, you will own the full stack of how our...  ...models, to ensuring those models perform accurately and efficiently in real-time...  ...segmentation of deformable food items, real-time inference under tight latency constraints, sensor... 
    Senior
    Performance
    Flexible hours

    Alumni Ventures

    San Francisco, CA
    18 hours ago
  • $180k - $240k

     ...Senior Systems Engineer We are looking for a Senior Systems Developer to lead the design, development...  ...distributed systems powering high-performance infrastructure. This role goes beyond...  ...environments Familiarity with server and GPU hardware architecture and system-level... 
    Senior
    Performance
    Flexible hours

    Nscale

    San Francisco, CA
    2 days ago
  • C++ Graphics Engineer Metal API GPU Rendering Pipeline We are looking for a Senior Graphics Engineer to join our Metal team, where you will be at the heart of Unity...  ...engine. You will develop and maintain a high-performance API layer that ensures efficient communication... 
    Senior
    Performance

    unity3d

    San Francisco, CA
    3 days ago
  • $175k - $250k

     ...0.00/yr - $250,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote...  ...vision while ensuring scalability, performance, and reliability across environments...  ...workloads at scale Manage and automate GPU compute clusters using tools such as... 
    Senior
    Performance
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    18 hours ago
  • $140k - $200k

     ...efficient world.   The Sr. Data Infrastructure / Quality Engineering role will play a crucial part in architecting, building,...  ...systems Experience deploying and optimizing high-performance GPU cloud inference services, with specific expertise utilizing the NVIDIA architecture... 
    Senior
    Performance
    Full time
    Work experience placement
    Local area

    Ouster

    San Francisco, CA
    a month ago
  • Strava is seeking a senior Data Scientist to design and deploy measurement strategies for its most impactful...  ...between product development, marketing, and performance metrics. You will advance industry-standard causal inference frameworks, collaborate across DS, Analytics,... 
    Senior
    Performance

    Apply

    San Francisco, CA
    4 days ago
  • $170.1k - $258.3k

     ...export, kernel development, and performance engineering so that every cycle on our...  ...builds high‑performance GPU kernels and custom libraries...  ...heart of our on‑vehicle ML inference for ADAS and autonomous driving...  ..., benchmark, and iterate on CUDA-based kernels and custom operators... 
    Senior
    Performance
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    San Francisco, CA
    18 hours ago
  • $229.5k - $255k

     ...'re helping autonomy teams and engineers overcome is the sim-to-real gap...  ...embodied AI. We're looking for a Senior Computer Vision Engineer to...  ...leave implicit. Improve Performance, Throughput, and Cost - Profile and optimize GPU and distributed workloads, reduce... 
    Senior
    Performance
    Full time
    Work at office
    3 days per week

    Niantic Spatial

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Inference Performance Engineer - GPU & CUDA. Be the first to apply!