Senior AI Inference Engineer - GPU, Rust & CUDA
$220kPerplexity
Perplexity is looking for an engineer to join their team in San Francisco. You will work on building and operating the inference engine, supporting new models, migrating GPU kernels, and developing a Rust-based serving runtime. The ideal candidate has 3+ years of experience in software engineering with a focus on ML inference, familiarity with deep learning frameworks, and a strong understanding of GPU architectures. Compensation ranges from $220K to $485K. #J-18808-Ljbffr Perplexity
$220k - $320k
inference.net, a growing company in San Francisco, seeks an experienced engineer to optimize AI inference performance. The ideal candidate will have over 2 years of experience in ML systems and GPU programming. Key responsibilities include implementing optimization techniques...Senior$250k - $300k
...the only vertically integrated AI infrastructure company built... ...production. That means owning the inference stack end to end: profiling... ...also work directly with customer engineering teams to tailor deployments to... ...like vLLM and SGLang to the CUDA kernels underneath, profiling...SeniorTemporary work- ...in San Francisco, CA, is seeking a Global Inference Library Engineer to build and maintain a high-performance inference library for modern AI models across diverse compute... ...programming with Python and C++, and hands-on CUDA/ROCm/Triton experience, plus familiarity...Senior
- ...scale. As part of the inference team, you’ll be responsible... ...for a kernel-focused engineer to lead efforts in... ...porting, and optimizing GPU kernels used in inference... ...deep familiarity with CUDA or equivalent kernel programming... ...OpenAI OpenAI is an AI research and deployment...SuggestedFull time
- ...About the Team Our Inference team brings OpenAI’s most... ...our state-of-the-art AI models, allowing them to... ...the Role We’re hiring engineers to scale and optimize... ...infrastructure across emerging GPU platforms. You’ll work... ...GPU kernels using HIP, CUDA, or Triton, and care...SuggestedFull time
$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model... ...and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer... ...management to support in API Gateway. GPU kernels migration to CuTe DSL. Port...- Perplexity is building a self-serve compute platform that lets inference engineers run training jobs and inference services without worrying about GPU provisioning or cluster configuration. You will own the platform surface and drive reliability, observability, and performance...Senior
$220k - $320k
A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience in ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation...SeniorLocal area- An innovative studio is seeking an AI Infrastructure Engineer to enhance their ML infrastructure for groundbreaking anime games. This role involves designing and implementing cutting-edge inference architectures to support various platforms. As part of a small, agile team...Worldwide
$160k - $225k
Cacheflow is seeking a Senior Software Engineer for AI Runtime at Databricks, located in San Francisco. You will be instrumental in building and scaling systems for large-scale GPU training, ensuring high throughput and resilience in training across expansive fleets of...Senior- OpenAI is seeking an experienced software engineer to join the GPT Infrastructure team and help build an automated inference optimization platform that scales research prototypes into production‑ready solutions. You will design durable APIs and control‑plane services,...Senior
$200k - $350k
...systems and platforms. We are seeking a Senior AI Engineer to develop, deploy, and optimize... ...AI, and modern ML frameworks. Build inference and evaluation pipelines. Optimize model... ...and cloud infrastructure. GPU optimization. RAG, fine-tuning, or evaluation...SeniorRemote jobFull timeImmediate start- A cutting-edge AI technology company based in San Francisco is seeking a specialist to design and operate large-scale GPU infrastructure. This role requires expertise in deploying GPU systems for high-throughput inference and model performance optimization. The ideal candidate...Senior
- ...Accellor is an AI-native services... ...advanced AI, data, and engineering capabilities.... ...— AI Systems, Inference & Platform Internals... ...model serving, GPU infrastructure,... ...candidate is a senior hands-on... ...-style serving, CUDA/Triton kernels,... ...such as C++, Go, Rust, Java, or TypeScript...
- ...Our client is a well-funded AI startup building production-... ...customers. They are looking for a Senior AI/ML Engineer to own model training... ...pipelines, evaluation systems, and inference serving at scale. Full-time,... ...with distributed training, GPU optimization, or inference...SeniorFull time
- ...The role: SoFi’s Staff AI Engineer is a hands-on AI engineering... ...organization. This is a critical, senior role responsible for setting... ...high-throughput, low-latency inference across diverse hardware... ...managing the underlying Kubernetes/GPU orchestration for custom...SeniorFull time
$128k - $252.5k
...executives and data scientists to AI strategists, machine learning specialists, and data engineers. SFL Scientific, a Deloitte Business, is looking to add a Senior AI Engineer to their vibrant... ...)2+ years of experience with GPU computing (CUDA, OpenCL) and HPC system...SeniorLocal areaVisa sponsorship$211k - $235k
...an applied science company building GPU-resident distributed data systems... ...Graph Neural Networks, and causal inference to deliver real-time analytics that... ...impossible.The RoleWe're hiring a Senior Software Engineer onto our Applied AI team to build and extend the backend...SeniorWork at office$250k
...Ready to architect AI infrastructure... ...building a serverless inference platform,... ...chance to join as a Senior Inference Platform Engineer at an early stage... ...systems to maximise GPU utilisation and... ...Proficiency in Python, Go, Rust, or a comparable... ...software stacks (CUDA, Triton, NCCL)...SeniorFull time$160k - $200k
...Simbe is building the AI powered operating system for... ...Simbe is looking for a Senior Computer Vision / Applied AI Engineer to build production AI systems... ..., release gates, inference wrappers, ONNX/TensorRT exports... ...with ONNX, TensorRT, CUDA, quantization, model profiling...SeniorFull time- ...leading design technology company in San Francisco is seeking a Senior Software Engineer for Backend (Systems / Infrastructure). You will architect... ...demand grows. This role involves optimizing APIs, managing GPU workloads, and collaborating with cross-functional teams....Senior
- Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest... ...design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every...
- ...are a small, fast-growing team of engineers in San Francisco powering Fortune 1... ...Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own... ...~ Strong Python, plus C++ or CUDA exposure ~ Experience with GPU profiling and model serving Nice...Full timeWork at officeVisa sponsorshipRelocation package
- iframe.ai is hiring a Customer Cluster Engineer to own three to five reserved-capacity accounts, managing their training and inference performance end-to-end. You will be embedded with the accounts... .../FSDP/Megatron-LM debugging, CUDA and Triton/CUTLASS familiarity, and...SeniorRemote job
- ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,... ...help build the platform engineers turn to to ship AI... ...ROLE We’re seeking a GPU Kernel Engineer to join our... ...Write and optimize code using CUDA, PTX assembly, and architecture...Full timeFlexible hours
- ...access state-of-the-art AI models - unlocking new... ...high-performance model inference and accelerating research... ...this role, you’ll lead engineering efforts to ensure our... ...responsible for shaping our CUDA strategy, driving... ...Mentor engineers on GPU performance, CUDA development...Full time
$320k
...interpretable, and steerable AI systems. We want AI to... ...committed researchers, engineers, policy experts, and... ...Our mandate is to make inference deployment boring and... ...into production across GPU, TPU, and Trainium fleets... ...with Python and/or Rust in production systems...Full timeWork at officeVisa sponsorshipFlexible hoursShift work$300k
...interpretable, and steerable AI systems. We want AI to be safe... ...group of committed researchers, engineers, policy experts, and business... ...About the Role The Cloud Inference team scales and optimizes Claude... ...Proficiency in Python or Rust The annual compensation...SeniorFull timeWork at officeVisa sponsorshipFlexible hours- ...Associates Limited is seeking a Staff-level engineer to architect and evolve the AI infrastructure software stack for large-scale GPU workloads. You’ll drive orchestration,... ...production environment. You’ll contribute to inference platforms, model serving, and high-...
- ...notch technology products.As a Senior Lead Software Engineer at JPMorgan Chase within... ...platforms optimized for AI/ML workloads.Partner with AI... ...architecture, ML training, and inference.Experience with... ...skillsFoundational understanding of NVIDIA GPU infrastructure software (e....SeniorFor contractors
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior AI Inference Engineer - GPU, Rust & CUDA. Be the first to apply!
- ai ml engineer San Francisco, CA
- machine learning ai engineer San Francisco, CA
- ai developer San Francisco, CA
- ai engineer San Francisco, CA
- ai research engineer San Francisco, CA
- ai prompt engineer San Francisco, CA
- senior ai engineer San Francisco, CA
- ai engineer remote San Francisco, CA
- senior lead project manager San Francisco, CA
- senior robotics software engineer San Francisco, CA


