Senior AI Inference Engineer - GPU, Rust & CUDA
$220kPerplexity
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr
- LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures... ...deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with...Senior
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...Senior- ...About the Team Our Inference team brings OpenAI’s most... ...our state-of-the-art AI models, allowing them to... ...the Role We’re hiring engineers to scale and optimize... ...infrastructure across emerging GPU platforms. You’ll work... ...GPU kernels using HIP, CUDA, or Triton, and care...SuggestedFull time
$220k - $320k
A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation...SeniorLocal area- An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team...SuggestedWorldwide
- OpenAI is seeking an experienced software engineer to join the GPT Infrastructure team and help build an automated inference optimization platform that scales research prototypes into production‑ready solutions. You will design durable APIs and control‑plane services,...Senior
- A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate...Senior
$240k - $280k
...looking for a Software Engineer to build the systems... ...of GPUs into running inference clusters without a human... ...a fully functioning AI cluster for training or... ...inference bring-up to GPU driver/CUDA stack, health... ...background in Go, Python, Rust, or similar — you write...Full time$200k - $350k
...systems and platforms. We are seeking a Senior AI Engineer to develop, deploy, and optimize... ...AI, and modern ML frameworks. Build inference and evaluation pipelines. Optimize model... ...and cloud infrastructure. GPU optimization. RAG, fine-tuning, or evaluation...SeniorRemote jobFull timeImmediate start- ...Accellor is an AI-native services... ...advanced AI, data, and engineering capabilities.... ...— AI Systems, Inference & Platform Internals... ...model serving, GPU infrastructure,... ...candidate is a senior hands-on... ...-style serving, CUDA/Triton kernels,... ...such as C++, Go, Rust, Java, or TypeScript...
$215k - $260k
...the only vertically integrated AI infrastructure company built... ...production. That means owning the inference stack end to end: profiling... ...also work directly with customer engineering teams to tailor deployments to... ...like vLLM and SGLang to the CUDA kernels underneath, profiling...Temporary work- ...Our client is a well-funded AI startup building production-... ...customers. They are looking for a Senior AI/ML Engineer to own model training... ...pipelines, evaluation systems, and inference serving at scale. Full-time,... ...with distributed training, GPU optimization, or inference...SeniorFull time
$211k - $235k
...an applied science company building GPU-resident distributed data systems... ...Graph Neural Networks, and causal inference to deliver real-time analytics that... ...impossible.The RoleWe're hiring a Senior Software Engineer onto our Applied AI team to build and extend the backend...SeniorWork at office$250k
...Ready to architect AI infrastructure... ...building a serverless inference platform,... ...chance to join as a Senior Inference Platform Engineer at an early stage... ...systems to maximise GPU utilisation and... ...Proficiency in Python, Go, Rust, or a comparable... ...software stacks (CUDA, Triton, NCCL)...SeniorFull time- ...leading design technology company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure). You will architect... ...demand grows. This role involves optimizing APIs, managing GPU workloads, and collaborating with cross-functional teams....Senior
$160k - $200k
...Simbe is building the AI powered operating system for... ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems... ..., release gates, inference wrappers, ONNX/TensorRT exports... ...with ONNX, TensorRT, CUDA, quantization, model profiling...SeniorFull time- Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest... ...design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every...
$229.9k - $262.4k
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real...SeniorFull timePart timeLocal area- iframe.ai is hiring a Customer Cluster Engineer to own three to five reserved-capacity accounts, managing their training and inference performance end-to-end. You will be embedded with the accounts... .../FSDP/Megatron-LM debugging, CUDA and Triton/CUTLASS familiarity, and...SeniorRemote job
- A leading data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building large-scale distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates...Senior
$195k - $255k
...the latest advancements in AI and IoTWho we areThe challenge... ...detection and beyond.As a Senior Computer Vision Engineer, you will lead the design,... ...computer vision models and inference pipelines running on both... ...learning models for ARM64, CUDA, TensorRT, ONNX, and NVIDIA...SeniorFull timeLocal areaRemote workWorldwide- ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,... ...help build the platform engineers turn to to ship AI... ...ROLE We’re seeking a GPU Kernel Engineer to join our... ...Write and optimize code using CUDA, PTX assembly, and architecture...Full timeFlexible hours
- ...Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,... ...production environment. You’ll contribute to inference platforms, model serving, and high-...
$210k - $300k
...Nimble Nimble is an AI robotics company... ...the world's best engineers and operators. If you... ...the training and inference systems that power... ...systems that turn our GPU clusters into a... ...optimize low-level CUDA kernels. Design... ...languages such as Rust, Go, Python, or C++...SeniorLocal areaFlexible hours$207k - $290k
...Description About JazzX AI: Vision:... ...seeking an experienced AI Engineer with deep expertise in... ...to join our team as a Senior Staff Architect. In this... ...techniques , including inference-time search, chain-of-thought... ...infrastructure (Kubernetes, GPU/TPU clusters, and cloud...SeniorWorldwideFlexible hours- Give every AI agent its own GPU Senior Software Engineer – GPU Sandboxes San Francisco onsite | Relocation considered... ...PCIe isolation NVIDIA drivers and CUDA lifecycle VM boot paths, snapshots... ...and physical hardware. Go, Rust or C/C++ can all work. This is rare...SeniorFull timeRelocation
- ...and scale revolutionary AI‑powered enterprise... ...seeking an experienced AI Engineer with deep expertise in... ...to join our team as a Senior Staff Architect. In this... ...techniques , including inference‑time search, chain‑of‑thought... ...(Kubernetes, GPU/TPU clusters, and cloud...SeniorFlexible hours
- ...About the Team Our Inference team brings OpenAI’s most... ...access our start-of-the-art AI models, allowing them... ...We are looking for an engineer who wants to take the... ...every FLOP and every GB of GPU RAM of our hardware.... ...optimize them (e.g. NCCL, CUDA), as well as HPC...Full time
$200k - $350k
...platforms. We are seeking a Staff AI Engineer to lead the design and deployment... ...systems. Design model serving, inference, evaluation, and optimization... ...Nice to Have vLLM, TGI, Triton, CUDA, PyTorch. Kubernetes and GPU infrastructure. RAG and fine-tuning...Remote jobFull timeImmediate start- ...Francisco Tensor Company is seeking a Member of Technical Staff for GPU Compiler Engineering to build the machine that searches and to optimize the... ..., and cross-architecture targets to push high-throughput AI workloads. Join a small, hands-on team in San Francisco focused...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!
- ai research engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior technical analyst San Francisco, CA
- senior associate attorney San Francisco, CA




