Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior AI Inference Engineer - GPU, Rust & CUDA

$220k

Perplexity

Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior AI Inference Engineer - GPU, Rust & CUDA in San Francisco, CA vacancy
  • $250k - $300k

     ...the only vertically integrated AI infrastructure company built...  ...production. That means owning the inference stack end to end: profiling...  ...also work directly with customer engineering teams to tailor deployments to...  ...like vLLM and SGLang to the CUDA kernels underneath, profiling... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  •  ...About the Team Our Inference team brings OpenAI’s most...  ...our state-of-the-art AI models, allowing them to...  ...the Role We’re hiring engineers to scale and optimize...  ...infrastructure across emerging GPU platforms. You’ll work...  ...GPU kernels using HIP, CUDA, or Triton, and care... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...scale. As part of the inference team, you’ll be responsible...  ...for a kernel-focused engineer to lead efforts in...  ...porting, and optimizing GPU kernels used in inference...  ...deep familiarity with CUDA or equivalent kernel programming...  ...OpenAI OpenAI is an AI research and deployment... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  • $300k

     ...startup building an AI and cloud platform,...  ...model training, or inference.  Our client...  ...operates high-performance GPU clusters powering...  ...operate inference engines such as vLLM, SGLang...  ...in Python, Go, Rust, or a comparable language...  ...software stacks (CUDA, Triton, NCCL) and... 
    Senior
    Permanent employment
    Worldwide
    San Francisco, CA
    more than 2 months ago
  • $128k - $252.5k

     ...executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte Business, is looking to add a Senior AI Engineer to their vibrant...  ...)2+ years of experience with GPU computing (CUDA, OpenCL) and HPC system... 
    Senior
    Local area
    Visa sponsorship

    Deloitte

    San Francisco, CA
    4 days ago
  • $211k - $235k

     ...an applied science company building GPU-resident distributed data systems...  ...Graph Neural Networks, and causal inference to deliver real-time analytics that...  ...impossible.The RoleWe're hiring a Senior Software Engineer onto our Applied AI team to build and extend the backend... 
    Senior
    Work at office

    Alembic

    San Francisco, CA
    3 days ago
  •  ....The role:SoFi’s SeniorStaff AI Engineer is a hands-on AI engineering...  ...organization. This is a critical, senior role responsible for setting...  ...high-throughput, low-latency inference across diverse hardware...  ...managing the underlying Kubernetes/GPU orchestration for custom... 
    Senior
    Remote work

    SoFi

    San Francisco, CA
    1 day ago
  •  ...Our client is a well-funded AI startup building production-...  ...customers. They are looking for a Senior AI/ML Engineer to own model training...  ...pipelines, evaluation systems, and inference serving at scale. Full-time,...  ...with distributed training, GPU optimization, or inference... 
    Senior
    Full time

    Clera

    San Francisco, CA
    23 hours ago
  • $160k - $200k

     ...Simbe is building the AI powered operating system for...  ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems...  ..., release gates, inference wrappers, ONNX/TensorRT exports...  ...with ONNX, TensorRT, CUDA, quantization, model profiling... 
    Senior
    Full time

    Simbe Robotics Inc

    San Francisco, CA
    23 hours ago
  •  ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within...  ...platforms optimized for AI/ML workloads.Partner with AI...  ...architecture, ML training, and inference.Experience with...  ...skillsFoundational understanding of NVIDIA GPU infrastructure software (e.... 
    Senior
    For contractors

    JP Morgan Chase

    San Francisco, CA
    2 days ago
  • $142.2k - $204.6k

    P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize...  ..., etc.Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS...  ...DatabricksDatabricks is the data and AI company. More than 10,000 organizations... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $320k

     ...interpretable, and steerable AI systems. We want AI to...  ...committed researchers, engineers, policy experts, and...  ...Our mandate is to make inference deployment boring and...  ...into production across GPU, TPU, and Trainium fleets...  ...with Python and/or Rust in production systems... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    23 hours ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,...  ...help build the platform engineers turn to to ship AI...  ...ROLE We’re seeking a GPU Kernel Engineer to join our...  ...Write and optimize code using CUDA, PTX assembly, and architecture... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    23 hours ago
  •  ...access state-of-the-art AI models - unlocking new...  ...high-performance model inference and accelerating research...  ...this role, you’ll lead engineering efforts to ensure our...  ...responsible for shaping our CUDA strategy, driving...  ...Mentor engineers on GPU performance, CUDA development... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  •  ...are a small, fast-growing team of engineers in San Francisco powering Fortune 1...  ...Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own...  ...~ Strong Python, plus C++ or CUDA exposure ~ Experience with GPU profiling and model serving Nice... 
    Full time
    Work at office
    Visa sponsorship
    Relocation package

    Pulse

    San Francisco, CA
    23 hours ago
  • $300k

     ...interpretable, and steerable AI systems. We want AI to be safe...  ...group of committed researchers, engineers, policy experts, and business...  ...About the Role The Cloud Inference team scales and optimizes Claude...  ...Proficiency in Python or Rust   The annual compensation... 
    Senior
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    23 hours ago
  •  ...with the web by building AI agents that can...  ...Scale infra for agentic inference (throughput and latency...  ...Work closely with product engineers to translate cutting‑edge...  ...with ML infrastructure (GPU clusters) and...  ...systems experience (Triton, CUDA) High IQ, high EQ, high... 
    Work at office
    Relocation
    Visa sponsorship

    Yutori

    San Francisco, CA
    1 day ago
  •  ...access our start-of-the-art AI models, allowing them...  ...and efficient model inference, as well as accelerating...  ...We are looking for an engineer who wants to take the world...  ...FLOP and every GB of GPU RAM of our hardware....  ...optimize them (e.g. NCCL, CUDA), as well as HPC... 
    Full time

    OpenAI

    San Francisco, CA
    23 hours ago
  • $149k - $240k

    Who We AreHP IQ is HP’s new AI innovation lab. Combining startup...  ...a diverse, world-class team—engineers, designers, researchers, and product...  .... We are looking for a Senior Software Engineer to design and...  ...storage solutions for real-time AI inference and processing.Implement... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    HP IQ

    San Francisco, CA
    6 days ago
  • $250k - $300k

     ...About Us At You.com, we are building the AI Search Infrastructure that powers modern...  ..., and useful. Our team includes engineers, researchers, product builders, and operators...  ...APIs — improving agentic results, cutting inference cost and token usage, and getting strong... 
    Senior
    Full time
    Immediate start
    Remote work
    Work from home
    Flexible hours

    You.com

    San Francisco, CA
    23 hours ago
  • $163.2k - $220.8k

     ...achievement and career growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team. The Senior AI...  ...secrets management for model API keys, network isolation for AI inference endpoints, and identity-aware proxy patterns for LLM... 
    Senior
    Full time
    Work experience placement
    Remote work
    Worldwide
    Shift work

    Wilson Sonsini Goodrich & Rosati

    San Francisco, CA
    4 days ago
  • $314.8k - $359.3k

    {"description": "Senior Distinguished AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking...  ...including foundation model training, large language model inference, similarity search, guardrails, model evaluation,... 
    Senior
    Full time
    Part time
    Local area

    Capital One Financial Corporation

    San Francisco, CA
    23 hours ago
  •  ...worldwide.We’re a team of engineers, clinicians, and innovators...  ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be responsible...  ...to real-time onboard inference—while serving as a core...  ...in GPU Compute API - CUDA, OpenCL• Proficiency in multiple... 
    Senior
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    1 day ago
  • ~ Senior Software Engineer (Rust) at Symbolica – San Francisco, US Senior Software Engineer (Rust) at Symbolica – San Francisco, US About Us Symbolica is an AI research lab pioneering the application of category theory to enable logical reasoning in machines... 
    Senior
    Work at office
    Shift work

    Victrays

    San Francisco, CA
    3 days ago
  •  ...deploy, monitor, and scale autonomous AI agents with full visibility and governance...  ...Overview We are looking for a Senior Applied AI Engineer to join our on-site Research & Intelligence...  ...to fine-tuning, PEFT, or cost-aware inference strategies. Experience working in... 
    Senior
    Full time

    Alterion, Inc.

    San Francisco, CA
    16 hours ago
  •  ...* OF OUR JOB OPENINGS!Senior Software Engineer — Backend PerformanceAs...  ...C/C++, Cython, or Rust when Python runs out of...  ...room, and reach for the GPU when it earns its keep...  ...Apply GPU acceleration (CUDA) to pipeline execution...  ...code yourself. AI assistants speed you up... 
    Senior

    Three Pillars Recruiting

    San Francisco, CA
    5 days ago
  •  ...that sit at the intersection of AI, biology, chemistry, and large-scale engineering. Our goal is to translate complex...  ...those systems. The Role As a Senior AI/ML Engineer, you will lead the...  ...decisions around model serving, inference efficiency, and lifecycle... 
    Senior
    Full time
    Remote work
    Flexible hours

    Absentia Labs

    San Francisco, CA
    23 hours ago
  •  ...AI Research Engineer Opportunity Poly is building a better file storage platform for...  ...Extensive experience with modeling / inference tools such as pytorch and CUDA. Pragmatic and product focused...  ...is a must, and enthusiasm for Rust is highly encouraged. An... 
    Work at office

    Poly

    San Francisco, CA
    23 hours ago
  • $100k - $250k

     ...Solutions Engineer Specializing In Ai Infrastructure And Platforms We are seeking...  ...requires deep expertise in GPU-accelerated computing, data...  ...solutions for AI training, inference, and high-performance...  ...systems, GPU technologies (e.g., CUDA, NCCL), and AI data center... 
    Work experience placement
    Work at office
    Flexible hours

    SHI GmbH

    San Francisco, CA
    2 days ago
  • $149k - $240k

     ...HP IQ is HP’s new AI innovation lab. Combining startup agility...  ...assembling a diverse, world‑class team—engineers, designers, researchers, and...  .... We are looking for a Senior Software Engineer to design and...  ...storage solutions for real‑time AI inference and processing. Implement... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    SupportFinity

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!