Senior Inference Performance Engineer - GPU & CUDA
$220k - $320kinference.net
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques and debugging inference frameworks. The role offers a competitive salary of $220,000 - $320,000 plus equity, and values curiosity and fast learning. You will join a team devoted to innovative AI solutions. #J-18808-Ljbffr inference.net
- ...worldwide.We’re a team of engineers, clinicians, and... ...work helps care teams perform with greater... ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics... ...to real-time onboard inference—while serving as a core... ...in GPU Compute API - CUDA, OpenCL• Proficiency...SeniorPerformanceLocal areaWorldwideFlexible hours
- ...San Francisco is seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the... ...requires deep hands-on experience with inference engines such as SGLang, vLLM or TensorRT, GPU architectures, and software stacks using #J...SeniorPerformance
$250k
...building a next-generation GPU platform designed for AI... ..., experimentation, and inference at scale. The company is... ...company is looking for a Senior / Staff Site Reliability Engineer to support and scale large... ...reliability, scalability, and performance of HPC and cloud...SeniorPerformanceFull timeRemote work- ...San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate will have hands-on experience with modern...SeniorPerformance
- ...AI acceleration company in San Francisco is seeking a GPU Kernel Engineer to optimize performance for machine learning models. You will be responsible for... ...computation efficiency. Ideal candidates have 1-5 years of CUDA development experience and a strong understanding of...Performance
$190k - $235k
...Francisco, CaliforniaSoftware Engineering /Full time Exempt /... ...you. ABOUT THE ROLEAs a senior Robot Perception... ...learningOptimize model inference for GPU deployment, leveraging CUDA, TensorRT, and related acceleration... ...)C/C++ experience for performance-critical...SeniorPerformanceFull time- ...building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You... ...and drive reliability, observability, and performance across workloads. You’ll manage the...SeniorPerformance
- Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack... ...You will primarily work on uzu, our inference engine, and focus on supporting new... ...work, and experience in writing high-performance GPU kernels or Rust systems programming....PerformanceLocal area
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...SeniorPerformance
$220k
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience...Senior- ...of Technical Staff to design and optimize inference systems. The role involves managing KV cache allocation and improving execution performance across various components. Ideal candidates should have strong software engineering skills and experience with ML inference systems...SeniorPerformance
$160k - $230k
...efficient and scalable inference for large language models... ...the boundaries of performance, scalability, and cost-... ...Frameworks and Optimization Engineer to design, develop, and... ...-throughput inference, GPU/accelerator... ...performance serving.Apply CUDA graph optimizations, TensorRT...PerformanceFull time- ...seeking an experienced systems generalist to build an automated inference optimization platform across hardware, compiler, and runtime... ...software, focusing on reliable long-running workflows, reproducible performance results, and secure handling of model and hardware data. This...SeniorPerformance
- ...seeking an expert in high‑performance LLM serving systems and inference optimization. In this... ...optimizing major inference engines such as SGLang, vLLM, or... ...knowledge of state‑of‑the‑art GPU architectures, and... ...using PyTorch, Triton, CuTe, CUDA, etc. Proven track record...Performance
$300k
...GPU Optimisation Engineer — Real-Time Inference Want to push GPU performance to its limits — not in theory, but in production systems handling real-time speech and multimodal... ..., and scheduling Writing and tuning custom CUDA / Triton kernels for performance-critical paths...PerformanceRelocationVisa sponsorshipFree visa- An innovative company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing and implementing infrastructure... ...for large-scale multimodal models, focusing on high-performance delivery of audio and image inputs. You'll collaborate...Performance
$100k - $150k
...vertically integrated AI cloud engineered for AI. We own and... ...energy, data centres, GPU superclusters,... ...services — delivering high-performance infrastructure to AI-... ...(Job Purpose) Senior Infrastructure Support... ...stacks on AI training and inference clusters. Confident...SeniorPerformanceFull timeRemote workFlexible hours$315k
A leading AI research company in San Francisco is seeking a mid-senior GPU Performance Engineer. In this role, you'll architect systems that enhance GPU performance for groundbreaking AI models. Responsibilities include developing optimizations, collaborating with teams...Performance$250k
...building a serverless inference platform, beginning... ...chance to join as a Senior Inference Platform Engineer at an early stage... ...systems to maximise GPU utilisation and minimise... ...best practices in performance and efficiency. Implement... ...software stacks (CUDA, Triton, NCCL) and...SeniorPerformanceFull time- ...build and operate the inference systems that serve... .... This is an engineering role, not a research... ...Own the performance characteristics of... ...WE'RE LOOKING FOR Senior ML systems engineer... ...Experience with GPU‑accelerated inference... ...following languages: C++, CUDA, ROCm or Triton...Performance
$100k - $120k
...foundation models. As training and inference workloads grow, we need kernel‑level... ...Responsibilities Lead a team of kernel and system engineers focused on performance-critical code Design, implement, and... ...kernels for CPU (AVX/ARM NEON), GPU (CUDA/ROCm), and hardware accelerators...Performance$250k
A Series A Funded start-up in California is seeking a Systems Engineer to design and optimize systems handling complex ML pipelines.... ...infrastructure, developing CI/CD pipelines, and ensuring system performance. Key qualifications include strong skills in Go and Python,...SeniorPerformance- ...and unicorn founders and senior engineers with deep expertise in 3... ...a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This is... ...using torch.compile, custom CUDA kernels, and specialized... ...Working knowledge of GPU hardware (NVIDIA) and...PerformanceRelocationVisa sponsorshipRelocation package
$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments... ...role involves collaboration with researchers and focus on performance optimization. Compensation ranges from $325K to $490K. #J-1880...SeniorPerformance$170.1k - $258.3k
...export, kernel development, and performance engineering so that every cycle on our... ...builds high‑performance GPU kernels and custom libraries... ...heart of our on‑vehicle ML inference for ADAS and autonomous driving... ..., benchmark, and iterate on CUDA-based kernels and custom operators...SeniorPerformanceFull timeLocal areaRemote workWork from homeRelocation packageFlexible hours- ...shipping excellence. We seek engineers with strong intrinsic drive,... ...experience to help scale AI inference. You’ll leverage your knowledge of high-performance systems to optimize GPU performance at the bleeding... ...or LA offices Tech Stack CUDA/C++, GPGPU, Python, Linux Key...PerformanceFull timeWork at office
- ...boundaries of what's possible in video generation.We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and... ...solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our...Performance
- Unity3D in San Francisco is seeking a Senior Graphics Engineer to enhance Unity's Metal graphics pipeline. You will develop and maintain a high-performance API that optimizes the graphics performance on Apple devices. This role requires over 10 years of C++ programming...SeniorPerformance
$160k - $200k
...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems... .... Improve model performance. Research and implement... ...validation, release gates, inference wrappers, ONNX/TensorRT... ...with ONNX, TensorRT, CUDA, quantization, model profiling...SeniorPerformanceFull time- ...infrastructure company is seeking a Senior Engineer 2 to join their AI Inference Optimization team. The role... ...leading the technical strategy for performance architecture and addressing complex... ...computing and a strong understanding of GPU architectures. The position offers...SeniorPerformanceRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Inference Performance Engineer - GPU & CUDA. Be the first to apply!
- senior operations associate San Francisco, CA
- senior safety specialist San Francisco, CA
- senior technology project manager San Francisco, CA
- remote senior business analyst San Francisco, CA
- senior director fp&a San Francisco, CA
- senior manager clinical operations San Francisco, CA
- senior supervisor San Francisco, CA
- senior researcher San Francisco, CA
- senior leadership San Francisco, CA
- senior title examiner San Francisco, CA



