Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Engineer - GPU, Rust & CUDA

$220k

Perplexity

Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr Perplexity

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Engineer - GPU, Rust & CUDA in San Francisco, CA vacancy
  • $220k - $320k

    inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques... 
    Senior

    inference.net

    San Francisco, CA
    3 days ago
  • $250k - $300k

     ...the only vertically integrated AI infrastructure company built...  ...production. That means owning the inference stack end to end: profiling...  ...also work directly with customer engineering teams to tailor deployments to...  ...like vLLM and SGLang to the CUDA kernels underneath, profiling... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  •  ...in San Francisco, CA, is seeking a Global Inference Library Engineer to build and maintain a high-performance inference library for modern AI models across diverse compute...  ...programming with Python and C++, and hands-on CUDA/ROCm/Triton experience, plus familiarity... 
    Senior

    Australia-Employment

    San Francisco, CA
    2 days ago
  •  ...scale. As part of the inference team, you’ll be responsible...  ...for a kernel-focused engineer to lead efforts in...  ...porting, and optimizing GPU kernels used in inference...  ...deep familiarity with CUDA or equivalent kernel programming...  ...OpenAI OpenAI is an AI research and deployment... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    13 hours ago
  •  ...About the Team Our Inference team brings OpenAI’s most...  ...our state-of-the-art AI models, allowing them to...  ...the Role We’re hiring engineers to scale and optimize...  ...infrastructure across emerging GPU platforms. You’ll work...  ...GPU kernels using HIP, CUDA, or Triton, and care... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    13 hours ago
  • $220k

    We build and run the inference engine behind every Perplexity query and deploy dozens of model...  ...and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer...  ...management to support in API Gateway. GPU kernels migration to CuTe DSL. Port... 

    Perplexity

    San Francisco, CA
    3 days ago
  • Perplexity is building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You will own the platform surface and drive reliability, observability, and performance... 
    Senior

    Neura Market

    San Francisco, CA
    4 days ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation... 
    Senior
    Local area

    Inference

    San Francisco, CA
    2 days ago
  • An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team... 
    Worldwide

    Spellbrush

    San Francisco, CA
    1 day ago
  • $160k - $225k

    Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of... 
    Senior

    Cacheflow

    San Francisco, CA
    5 days ago
  • OpenAI is seeking an experienced software engineer to join the GPT Infrastructure team and help build an automated inference optimization platform that scales research prototypes into production‑ready solutions. You will design durable APIs and control‑plane services,... 
    Senior

    OpenAI

    San Francisco, CA
    3 days ago
  • $200k - $350k

     ...systems and platforms. We are seeking a Senior AI Engineer to develop, deploy, and optimize...  ...AI, and modern ML frameworks. Build inference and evaluation pipelines. Optimize model...  ...and cloud infrastructure. GPU optimization. RAG, fine-tuning, or evaluation... 
    Senior
    Remote job
    Full time
    Immediate start

    Pragmatike

    San Francisco, CA
    13 hours ago
  • A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate... 
    Senior

    Reflection AI

    San Francisco, CA
    3 days ago
  •  ...Accellor is an AI-native services...  ...advanced AI, data, and engineering capabilities....  ...— AI Systems, Inference & Platform Internals...  ...model serving, GPU infrastructure,...  ...candidate is a senior hands-on...  ...-style serving, CUDA/Triton kernels,...  ...such as C++, Go, Rust, Java, or TypeScript... 

    Accellor

    San Francisco, CA
    5 days ago
  •  ...Our client is a well-funded AI startup building production-...  ...customers. They are looking for a Senior AI/ML Engineer to own model training...  ...pipelines, evaluation systems, and inference serving at scale. Full-time,...  ...with distributed training, GPU optimization, or inference... 
    Senior
    Full time

    Clera

    San Francisco, CA
    13 hours ago
  •  ...The role: SoFi’s Staff AI Engineer is a hands-on AI engineering...  ...organization. This is a critical, senior role responsible for setting...  ...high-throughput, low-latency inference across diverse hardware...  ...managing the underlying Kubernetes/GPU orchestration for custom... 
    Senior
    Full time

    Sofi

    San Francisco, CA
    13 hours ago
  • $128k - $252.5k

     ...executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte Business, is looking to add a Senior AI Engineer to their vibrant...  ...)2+ years of experience with GPU computing (CUDA, OpenCL) and HPC system... 
    Senior
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    4 days ago
  • $211k - $235k

     ...an applied science company building GPU-resident distributed data systems...  ...Graph Neural Networks, and causal inference to deliver real-time analytics that...  ...impossible.The RoleWe're hiring a Senior Software Engineer onto our Applied AI team to build and extend the backend... 
    Senior
    Work at office

    Alembic

    San Francisco, CA
    3 days ago
  • $250k

     ...Ready to architect AI infrastructure...  ...building a serverless inference platform,...  ...chance to join as a Senior Inference Platform Engineer at an early stage...  ...systems to maximise GPU utilisation and...  ...Proficiency in Python, Go, Rust, or a comparable...  ...software stacks (CUDA, Triton, NCCL)... 
    Senior
    Full time
    San Francisco, CA
    more than 2 months ago
  • $160k - $200k

     ...Simbe is building the AI powered operating system for...  ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems...  ..., release gates, inference wrappers, ONNX/TensorRT exports...  ...with ONNX, TensorRT, CUDA, quantization, model profiling... 
    Senior
    Full time

    Simbe Robotics Inc

    San Francisco, CA
    13 hours ago
  •  ...leading design technology company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure). You will architect...  ...demand grows. This role involves optimizing APIs, managing GPU workloads, and collaborating with cross-functional teams.... 
    Senior

    Vizcom

    San Francisco, CA
    1 day ago
  • Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest...  ...design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every... 

    Sail

    San Francisco, CA
    2 days ago
  •  ...are a small, fast-growing team of engineers in San Francisco powering Fortune 1...  ...Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own...  ...~ Strong Python, plus C++ or CUDA exposure ~ Experience with GPU profiling and model serving Nice... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Pulse

    San Francisco, CA
    13 hours ago
  • iframe.ai is hiring a Customer Cluster Engineer to own three to five reserved-capacity accounts, managing their training and inference performance end-to-end. You will be embedded with the accounts...  .../FSDP/Megatron-LM debugging, CUDA and Triton/CUTLASS familiarity, and... 
    Senior
    Remote job

    iFrame

    San Francisco, CA
    1 day ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,...  ...help build the platform engineers turn to to ship AI...  ...ROLE We’re seeking a GPU Kernel Engineer to join our...  ...Write and optimize code using CUDA, PTX assembly, and architecture... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    13 hours ago
  •  ...access state-of-the-art AI models - unlocking new...  ...high-performance model inference and accelerating research...  ...this role, you’ll lead engineering efforts to ensure our...  ...responsible for shaping our CUDA strategy, driving...  ...Mentor engineers on GPU performance, CUDA development... 
    Full time

    OpenAI

    San Francisco, CA
    13 hours ago
  • $320k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...Our mandate is to make inference deployment boring and...  ...into production across GPU, TPU, and Trainium fleets...  ...with Python and/or Rust in production systems... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    13 hours ago
  • $300k

     ...interpretable, and steerable AI systems. We want AI to be safe...  ...group of committed researchers, engineers, policy experts, and business...  ...About the Role The Cloud Inference team scales and optimizes Claude...  ...Proficiency in Python or Rust   The annual compensation... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    13 hours ago
  •  ...Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,...  ...production environment. You’ll contribute to inference platforms, model serving, and high-... 

    Hamilton Barnes Associates Limited

    San Francisco, CA
    1 day ago
  •  ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within...  ...platforms optimized for AI/ML workloads.Partner with AI...  ...architecture, ML training, and inference.Experience with...  ...skillsFoundational understanding of NVIDIA GPU infrastructure software (e.... 
    Senior
    For contractors

    JP Morgan Chase

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!