Senior AI Inference Engineer - GPU, Rust & CUDA
$220kPerplexity
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr
- Acceler8 Talent in San Francisco is partnering with a rapidly growing AI infrastructure company to own the core cluster infrastructure powering a heterogeneous AI cloud. You will manage large CPU/GPU/accelerator clusters, bare‑metal provisioning, and production readiness...Senior
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...Senior- OpenInfer is looking for a Full-Time AI Software Engineer specialized in Rust to enhance our local AI inference systems. You will work on developing and maintaining the LLM inference engine and collaborating with product teams to integrate new features across major client...SuggestedFull timeLocal areaFlexible hours
- A leading cloud infrastructure company is seeking a Senior Engineer 2 to join their AI Inference Optimization team. The role involves leading the technical... ...high-performance computing and a strong understanding of GPU architectures. The position offers a competitive salary...SeniorRemote job
$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model... ...and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer... ...management to support in API Gateway. GPU kernels migration to CuTe DSL. Port...Suggested$220k - $320k
A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation...SeniorLocal area$160k - $250k
Together AI is building the Inference Platform that brings the most advanced generative... ...data centers and model engine pods. Develop auto‑... ...in one or more of: Rust, Go, Python, or TypeScript... ...plus. Familiarity with GPU software stacks (CUDA, Triton, NCCL) and HPC technologies...SeniorFull timeLocal area- An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team...Worldwide
- Gravity Engineering Services Pvt Ltd. is seeking a Senior Software Engineer for AI Runtime in San Francisco, California. The role involves driving the architecture of a managed GPU training platform, solving complex training challenges, and enhancing performance and reliability...Senior
$160k - $225k
Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of...Senior- Accenture is seeking a seasoned AI Infrastructure Architect in San Francisco to design and implement scalable AI infrastructure, including... ...for enterprise clients. The role requires deep experience with GPU/DPUs, networking, storage, and orchestration tools, plus strong...Senior
$167.2k - $209k
A leading cloud service provider is seeking a Senior Engineer 2 for their AI Inference Data Plane team. This remote role focuses on designing and developing high-scale, resilient data plane services that enhance AI-driven applications. The ideal candidate will have strong...SeniorRemote job- Acceler8 Talent is seeking a Senior Software Engineer to design and optimize AI inference systems for production workloads. You’ll work across runtime behavior, scheduling, memory management and system performance to deliver faster, more scalable AI inference. You’ll collaborate...Senior
- A technology company in San Francisco is seeking an experienced Infrastructure Engineer to ensure 99.99% uptime while working with custom inference stacks and managing GPU loads. The ideal candidate will have deep infrastructure expertise and a passion for tackling complex...Senior
- Lightning AI in San Francisco or Seattle is seeking a Senior Application Security Engineer to secure our AI/ML platforms and inference services. You will work with platform, ML, and infrastructure teams to identify risks and implement secure architectures. The role emphasizes...Senior
$300k
...startup building an AI and cloud platform,... ...model training, or inference. Our client... ...operates high-performance GPU clusters powering... ...operate inference engines such as vLLM, SGLang... ...in Python, Go, Rust, or a comparable language... ...software stacks (CUDA, Triton, NCCL) and...SeniorPermanent employmentWorldwide- A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate...Senior
$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...’s mandate is to make inference deployment boring and... ...into production across GPU, TPU, and Trainium fleets... ...with Python and/or Rust in production systems...SeniorWork at officeVisa sponsorshipFlexible hoursShift work$216k - $270k
As a Software Engineer on the Machine Learning Infrastructure... ...” for our large-scale GPU clusters. You will... ...compute into breakthrough AI. You will: Architect and... ...languages (e.g., Python, Go, Rust, C++). Experience with... ...and hardware stack (CUDA, NCCL). Experience with...SeniorFull timeFor contractors- ...leading design technology company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure). You will architect... ...demand grows. This role involves optimizing APIs, managing GPU workloads, and collaborating with cross-functional teams....Senior
- ...role Anthropic's Inference organization... ...efficiency that frontier AI demands. We... ...for a Staff Engineer to be a technical... ...builds on. This is a senior IC role with... ...performance‑sensitive Rust and Python... ...management - across GPU , TPU , and... ...accelerator ecosystem ( CUDA/GPU , TPU , or...
- Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest... ...design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every...
- Together AI is looking for a Customer Support Engineer in San Francisco, CA. In this role, you will assist customers with complex technical challenges involving cutting-edge AI solutions. The position requires a strong technical background and excellent communication skills...Remote work
$320k
Anthropic in San Francisco is seeking a Software Engineer for their Launch Engineering team. You'll design infrastructure to deploy AI models efficiently and manage resource constraints while ensuring high reliability. Ideal candidates should have robust software engineering...Senior- Luma AI in San Francisco is seeking a leader to define reliability for a frontier AI infrastructure. You will architect and operate large GPU environments, pushing the limits of training and inference while partnering with research and product to scale systems and improve...
$96.8k - $306.4k
...Job Description The Senior Principal AI Agent / ML Software Engineer is a Senior Staff-level, hands‑on technical leadership... ..., autonomous workflows, scalable inference infrastructure, and enterprise AI... ...for low latency, high throughput, GPU efficiency, reliability, cost,...SeniorTemporary workFlexible hours- Gravity Engineering Services Pvt Ltd. is seeking a Staff Software Engineer for AI Research Infrastructure. You will develop and run a... ...orchestrate large-scale training and inference workloads across thousands of... ...languages like C++, Rust, and Go. Join us to influence...
- ...and scaling predictive analytics and AI systems on top of Amperesand’s IoT platform... ...including data ingestion, feature engineering, model training, inference, deployment, and monitoring. Develop... ...skills in Python ; Go, Java, or Rust is a plus. Experience with ML frameworks...SeniorWork experience placement
- ...and scale revolutionary AI‑powered enterprise... ...seeking an experienced AI Engineer with deep expertise in... ...to join our team as a Senior Staff Architect. In this... ...techniques , including inference‑time search, chain‑of‑thought... ...(Kubernetes, GPU/TPU clusters, and cloud...SeniorFlexible hours
- Perplexity seeks an experienced platform engineer to own and evolve a self-serve GPU compute platform. You will design and operate GPU provisioning, cluster... ...orchestration to support both training and real-time inference workloads. You’ll implement fault-tolerant scheduling,...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!
- ai engineer San Francisco, CA
- ai ml engineer San Francisco, CA
- ai research engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai developer San Francisco, CA
- ai prompt engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- senior trade analyst San Francisco, CA
- senior app developer San Francisco, CA

