Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Engineer - GPU, Rust & CUDA

$220k

Perplexity

Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Engineer - GPU, Rust & CUDA in San Francisco, CA vacancy
  • LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You will work across diverse compute architectures...  ...deep knowledge of AI infrastructure, GPU programming, and low-level kernels, with... 
    Senior

    Leoforce

    San Francisco, CA
    5 days ago
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Senior

    inference.net

    San Francisco, CA
    5 days ago
  •  ...About the Team Our Inference team brings OpenAI’s most...  ...our state-of-the-art AI models, allowing them to...  ...the Role We’re hiring engineers to scale and optimize...  ...infrastructure across emerging GPU platforms. You’ll work...  ...GPU kernels using HIP, CUDA, or Triton, and care... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation... 
    Senior
    Local area

    Inference

    San Francisco, CA
    4 days ago
  • An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team... 
    Suggested
    Worldwide

    Spellbrush

    San Francisco, CA
    3 days ago
  • OpenAI is seeking an experienced software engineer to join the GPT Infrastructure team and help build an automated inference optimization platform that scales research prototypes into production‑ready solutions. You will design durable APIs and control‑plane services,... 
    Senior

    OpenAI

    San Francisco, CA
    5 days ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate... 
    Senior

    Reflection AI

    San Francisco, CA
    5 days ago
  • $240k - $280k

     ...looking for a Software Engineer to build the systems...  ...of GPUs into running inference clusters without a human...  ...a fully functioning AI cluster for training or...  ...inference bring-up to GPU driver/CUDA stack, health...  ...background in Go, Python, Rust, or similar — you write... 
    Full time

    Together Ai

    San Francisco, CA
    1 day ago
  • $200k - $350k

     ...systems and platforms. We are seeking a Senior AI Engineer to develop, deploy, and optimize...  ...AI, and modern ML frameworks. Build inference and evaluation pipelines. Optimize model...  ...and cloud infrastructure. GPU optimization. RAG, fine-tuning, or evaluation... 
    Senior
    Remote job
    Full time
    Immediate start

    Pragmatike

    San Francisco, CA
    1 day ago
  •  ...Accellor is an AI-native services...  ...advanced AI, data, and engineering capabilities....  ...— AI Systems, Inference & Platform Internals...  ...model serving, GPU infrastructure,...  ...candidate is a senior hands-on...  ...-style serving, CUDA/Triton kernels,...  ...such as C++, Go, Rust, Java, or TypeScript... 

    Accellor

    San Francisco, CA
    21 days ago
  • $215k - $260k

     ...the only vertically integrated AI infrastructure company built...  ...production. That means owning the inference stack end to end: profiling...  ...also work directly with customer engineering teams to tailor deployments to...  ...like vLLM and SGLang to the CUDA kernels underneath, profiling... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  •  ...Our client is a well-funded AI startup building production-...  ...customers. They are looking for a Senior AI/ML Engineer to own model training...  ...pipelines, evaluation systems, and inference serving at scale. Full-time,...  ...with distributed training, GPU optimization, or inference... 
    Senior
    Full time

    Clera

    San Francisco, CA
    1 day ago
  • $211k - $235k

     ...an applied science company building GPU-resident distributed data systems...  ...Graph Neural Networks, and causal inference to deliver real-time analytics that...  ...impossible.The RoleWe're hiring a Senior Software Engineer onto our Applied AI team to build and extend the backend... 
    Senior
    Work at office

    Alembic

    San Francisco, CA
    5 days ago
  • $250k

     ...Ready to architect AI infrastructure...  ...building a serverless inference platform,...  ...chance to join as a Senior Inference Platform Engineer at an early stage...  ...systems to maximise GPU utilisation and...  ...Proficiency in Python, Go, Rust, or a comparable...  ...software stacks (CUDA, Triton, NCCL)... 
    Senior
    Full time
    San Francisco, CA
    more than 2 months ago
  •  ...leading design technology company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure). You will architect...  ...demand grows. This role involves optimizing APIs, managing GPU workloads, and collaborating with cross-functional teams.... 
    Senior

    Vizcom

    San Francisco, CA
    3 days ago
  • $160k - $200k

     ...Simbe is building the AI powered operating system for...  ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems...  ..., release gates, inference wrappers, ONNX/TensorRT exports...  ...with ONNX, TensorRT, CUDA, quantization, model profiling... 
    Senior
    Full time

    Simbe Robotics Inc

    San Francisco, CA
    1 day ago
  • Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest...  ...design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every... 

    SAIL

    San Francisco, CA
    4 days ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real... 
    Senior
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    2 days ago
  • iframe.ai is hiring a Customer Cluster Engineer to own three to five reserved-capacity accounts, managing their training and inference performance end-to-end. You will be embedded with the accounts...  .../FSDP/Megatron-LM debugging, CUDA and Triton/CUTLASS familiarity, and... 
    Senior
    Remote job

    iFrame

    San Francisco, CA
    3 days ago
  • A leading data and AI company in San Francisco is seeking a Senior Engineer to enhance their Model Serving platform. This role requires expertise in building large-scale distributed systems and collaboration across teams to optimize performance and reliability. Ideal candidates... 
    Senior

    Jobleads-US

    San Francisco, CA
    3 days ago
  • $195k - $255k

     ...the latest advancements in AI and IoTWho we areThe challenge...  ...detection and beyond.As a Senior Computer Vision Engineer, you will lead the design,...  ...computer vision models and inference pipelines running on both...  ...learning models for ARM64, CUDA, TensorRT, ONNX, and NVIDIA... 
    Senior
    Full time
    Local area
    Remote work
    Worldwide

    Pano AI

    San Francisco, CA
    3 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,...  ...help build the platform engineers turn to to ship AI...  ...ROLE We’re seeking a GPU Kernel Engineer to join our...  ...Write and optimize code using CUDA, PTX assembly, and architecture... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,...  ...production environment. You’ll contribute to inference platforms, model serving, and high-... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    3 days ago
  • $210k - $300k

     ...Nimble Nimble is an AI robotics company...  ...the world's best engineers and operators. If you...  ...the training and inference systems that power...  ...systems that turn our GPU clusters into a...  ...optimize low-level CUDA kernels. Design...  ...languages such as Rust, Go, Python, or C++... 
    Senior
    Local area
    Flexible hours

    Nimble Robotics

    San Francisco, CA
    24 days ago
  • $207k - $290k

     ...Description About JazzX AI:   Vision:...  ...seeking an experienced AI Engineer with deep expertise in...  ...to join our team as a Senior Staff Architect. In this...  ...techniques , including inference-time search, chain-of-thought...  ...infrastructure (Kubernetes, GPU/TPU clusters, and cloud... 
    Senior
    Worldwide
    Flexible hours

    JazzX AI

    San Francisco, CA
    8 days ago
  • Give every AI agent its own GPU Senior Software Engineer – GPU Sandboxes San Francisco onsite | Relocation considered...  ...PCIe isolation NVIDIA drivers and CUDA lifecycle VM boot paths, snapshots...  ...and physical hardware. Go, Rust or C/C++ can all work. This is rare... 
    Senior
    Full time
    Relocation

    nineDots Technology Recruitment

    San Francisco, CA
    15 days ago
  •  ...and scale revolutionary AI‑powered enterprise...  ...seeking an experienced AI Engineer with deep expertise in...  ...to join our team as a Senior Staff Architect. In this...  ...techniques , including inference‑time search, chain‑of‑thought...  ...(Kubernetes, GPU/TPU clusters, and cloud... 
    Senior
    Flexible hours

    JazzX AI

    San Francisco, CA
    3 days ago
  •  ...About the Team Our Inference team brings OpenAI’s most...  ...access our start-of-the-art AI models, allowing them...  ...We are looking for an engineer who wants to take the...  ...every FLOP and every GB of GPU RAM of our hardware....  ...optimize them (e.g. NCCL, CUDA), as well as HPC... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $200k - $350k

     ...platforms. We are seeking a Staff AI Engineer to lead the design and deployment...  ...systems. Design model serving, inference, evaluation, and optimization...  ...Nice to Have vLLM, TGI, Triton, CUDA, PyTorch. Kubernetes and GPU infrastructure. RAG and fine-tuning... 
    Remote job
    Full time
    Immediate start

    Pragmatike

    San Francisco, CA
    1 day ago
  •  ...Francisco Tensor Company is seeking a Member of Technical Staff for GPU Compiler Engineering to build the machine that searches and to optimize the...  ..., and cross-architecture targets to push high-throughput AI workloads. Join a small, hands-on team in San Francisco focused... 
    Senior

    San Francisco Tensor Company

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!