Senior AI Inference Engineer - GPU, Rust & CUDA
$220kPerplexity
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr
$250k - $300k
...the only vertically integrated AI infrastructure company built... ...production. That means owning the inference stack end to end: profiling... ...also work directly with customer engineering teams to tailor deployments to... ...like vLLM and SGLang to the CUDA kernels underneath, profiling...SeniorTemporary work- ...About the Team Our Inference team brings OpenAI’s most... ...our state-of-the-art AI models, allowing them to... ...the Role We’re hiring engineers to scale and optimize... ...infrastructure across emerging GPU platforms. You’ll work... ...GPU kernels using HIP, CUDA, or Triton, and care...SuggestedFull time
- ...scale. As part of the inference team, you’ll be responsible... ...for a kernel-focused engineer to lead efforts in... ...porting, and optimizing GPU kernels used in inference... ...deep familiarity with CUDA or equivalent kernel programming... ...OpenAI OpenAI is an AI research and deployment...SuggestedFull time
$300k
...startup building an AI and cloud platform,... ...model training, or inference. Our client... ...operates high-performance GPU clusters powering... ...operate inference engines such as vLLM, SGLang... ...in Python, Go, Rust, or a comparable language... ...software stacks (CUDA, Triton, NCCL) and...SeniorPermanent employmentWorldwide$128k - $252.5k
...executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte Business, is looking to add a Senior AI Engineer to their vibrant... ...)2+ years of experience with GPU computing (CUDA, OpenCL) and HPC system...SeniorLocal areaVisa sponsorship$211k - $235k
...an applied science company building GPU-resident distributed data systems... ...Graph Neural Networks, and causal inference to deliver real-time analytics that... ...impossible.The RoleWe're hiring a Senior Software Engineer onto our Applied AI team to build and extend the backend...SeniorWork at office- ....The role:SoFi’s SeniorStaff AI Engineer is a hands-on AI engineering... ...organization. This is a critical, senior role responsible for setting... ...high-throughput, low-latency inference across diverse hardware... ...managing the underlying Kubernetes/GPU orchestration for custom...SeniorRemote work
- ...Our client is a well-funded AI startup building production-... ...customers. They are looking for a Senior AI/ML Engineer to own model training... ...pipelines, evaluation systems, and inference serving at scale. Full-time,... ...with distributed training, GPU optimization, or inference...SeniorFull time
$160k - $200k
...Simbe is building the AI powered operating system for... ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems... ..., release gates, inference wrappers, ONNX/TensorRT exports... ...with ONNX, TensorRT, CUDA, quantization, model profiling...SeniorFull time- ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within... ...platforms optimized for AI/ML workloads.Partner with AI... ...architecture, ML training, and inference.Experience with... ...skillsFoundational understanding of NVIDIA GPU infrastructure software (e....SeniorFor contractors
$142.2k - $204.6k
P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize... ..., etc.Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS... ...DatabricksDatabricks is the data and AI company. More than 10,000 organizations...Local areaWorldwide$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...Our mandate is to make inference deployment boring and... ...into production across GPU, TPU, and Trainium fleets... ...with Python and/or Rust in production systems...Full timeWork at officeVisa sponsorshipFlexible hoursShift work- ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,... ...help build the platform engineers turn to to ship AI... ...ROLE We’re seeking a GPU Kernel Engineer to join our... ...Write and optimize code using CUDA, PTX assembly, and architecture...Full timeFlexible hours
- ...access state-of-the-art AI models - unlocking new... ...high-performance model inference and accelerating research... ...this role, you’ll lead engineering efforts to ensure our... ...responsible for shaping our CUDA strategy, driving... ...Mentor engineers on GPU performance, CUDA development...Full time
- ...are a small, fast-growing team of engineers in San Francisco powering Fortune 1... ...Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own... ...~ Strong Python, plus C++ or CUDA exposure ~ Experience with GPU profiling and model serving Nice...Full timeWork at officeVisa sponsorshipRelocation package
$300k
...interpretable, and steerable AI systems. We want AI to be safe... ...group of committed researchers, engineers, policy experts, and business... ...About the Role The Cloud Inference team scales and optimizes Claude... ...Proficiency in Python or Rust The annual compensation...SeniorFull timeWork at officeVisa sponsorshipFlexible hours- ...with the web by building AI agents that can... ...Scale infra for agentic inference (throughput and latency... ...Work closely with product engineers to translate cutting‑edge... ...with ML infrastructure (GPU clusters) and... ...systems experience (Triton, CUDA) High IQ, high EQ, high...Work at officeRelocationVisa sponsorship
- ...access our start-of-the-art AI models, allowing them... ...and efficient model inference, as well as accelerating... ...We are looking for an engineer who wants to take the world... ...FLOP and every GB of GPU RAM of our hardware.... ...optimize them (e.g. NCCL, CUDA), as well as HPC...Full time
$149k - $240k
Who We AreHP IQ is HP’s new AI innovation lab. Combining startup... ...a diverse, world-class team—engineers, designers, researchers, and product... .... We are looking for a Senior Software Engineer to design and... ...storage solutions for real-time AI inference and processing.Implement...SeniorFull timeTemporary workLocal areaFlexible hours$250k - $300k
...About Us At You.com, we are building the AI Search Infrastructure that powers modern... ..., and useful. Our team includes engineers, researchers, product builders, and operators... ...APIs — improving agentic results, cutting inference cost and token usage, and getting strong...SeniorFull timeImmediate startRemote workWork from homeFlexible hours$163.2k - $220.8k
...achievement and career growth. Wilson Sonsini is looking for a Senior AI Security Engineer to join the Security Operations team. The Senior AI... ...secrets management for model API keys, network isolation for AI inference endpoints, and identity-aware proxy patterns for LLM...SeniorFull timeWork experience placementRemote workWorldwideShift work$314.8k - $359.3k
{"description": "Senior Distinguished AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing banking... ...including foundation model training, large language model inference, similarity search, guardrails, model evaluation,...SeniorFull timePart timeLocal area- ...worldwide.We’re a team of engineers, clinicians, and innovators... ...Function of PositionAs a Senior Systems GPU Engineer - AI & Robotics, you will be responsible... ...to real-time onboard inference—while serving as a core... ...in GPU Compute API - CUDA, OpenCL• Proficiency in multiple...SeniorLocal areaWorldwideFlexible hours
- ~ Senior Software Engineer (Rust) at Symbolica – San Francisco, US Senior Software Engineer (Rust) at Symbolica – San Francisco, US About Us Symbolica is an AI research lab pioneering the application of category theory to enable logical reasoning in machines...SeniorWork at officeShift work
- ...deploy, monitor, and scale autonomous AI agents with full visibility and governance... ...Overview We are looking for a Senior Applied AI Engineer to join our on-site Research & Intelligence... ...to fine-tuning, PEFT, or cost-aware inference strategies. Experience working in...SeniorFull time
- ...* OF OUR JOB OPENINGS!Senior Software Engineer — Backend PerformanceAs... ...C/C++, Cython, or Rust when Python runs out of... ...room, and reach for the GPU when it earns its keep... ...Apply GPU acceleration (CUDA) to pipeline execution... ...code yourself. AI assistants speed you up...Senior
- ...that sit at the intersection of AI, biology, chemistry, and large-scale engineering. Our goal is to translate complex... ...those systems. The Role As a Senior AI/ML Engineer, you will lead the... ...decisions around model serving, inference efficiency, and lifecycle...SeniorFull timeRemote workFlexible hours
- ...AI Research Engineer Opportunity Poly is building a better file storage platform for... ...Extensive experience with modeling / inference tools such as pytorch and CUDA. Pragmatic and product focused... ...is a must, and enthusiasm for Rust is highly encouraged. An...Work at office
$100k - $250k
...Solutions Engineer Specializing In Ai Infrastructure And Platforms We are seeking... ...requires deep expertise in GPU-accelerated computing, data... ...solutions for AI training, inference, and high-performance... ...systems, GPU technologies (e.g., CUDA, NCCL), and AI data center...Work experience placementWork at officeFlexible hours$149k - $240k
...HP IQ is HP’s new AI innovation lab. Combining startup agility... ...assembling a diverse, world‑class team—engineers, designers, researchers, and... .... We are looking for a Senior Software Engineer to design and... ...storage solutions for real‑time AI inference and processing. Implement...SeniorFull timeTemporary workLocal areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!
- ai engineer remote San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- senior manufacturing manager San Francisco, CA
- senior business analyst San Francisco, CA



