AI Inference Engineer (GPU/Rust/CUDA)
Perplexity
Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure. You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and #J-18808-Ljbffr Perplexity
- ...Senior GPU Systems / AI Infrastructure Engineer (NYC) Location: New York City (Hybrid / On... ...-scale model training and inference. This role sits at the intersection... ...optimise GPU kernels (CUDA / Triton / HIP) for large-... ...high-performance C++ / Rust / Python systems ~...SuggestedFull time
$220k
We build and run the inference engine behind every Perplexity query and deploy dozens of model... ...and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer... ...management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our...Suggested$229.9k - $286.2k
AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing... ...reviews across AI systems, tracking GPU utilization, model throughput, and... ...programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications...SuggestedFull timePart timeLocal area$250k - $300k
Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI... ..., including but not limited to GPU kernel development, novel inference devices... ...engineering skills, especially any of: CUDA/Triton/Pallas/CuTe DSL kernel development...SuggestedWork experience placementWork at officeLocal areaImmediate start$197.3k - $225.1k
...AI Engineer 4 At Capital One, we are creating responsible and reliable... ..., large language model inference, agents and multi-agent workflows... ...engineering to optimize GPU/TPU utilization and accelerate... ...programming with Python, Go, Scala, CUDA, or Java Preferred...SuggestedFull timePart timeLocal area$100k - $150k
...delivering cloud, AI, data, and enterprise... ...Title: AI Platform Engineer Location: 100%... ...operate scalable AI inference platforms for production... ..., LLM serving, GPU optimization, autoscaling... ...in Python and Go, Rust, or C++. Experience... ...workloads using CUDA , NVIDIA GPU technologies...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing banking for good. We use machine learning to craft real-time, personalized customer experiences and build scalable, high...Local area
- ...Job Title: AI DevOps Engineer Job Summary We are seeking an AI DevOps... ...Optimize cloud resources, GPU utilization, and infrastructure... ...such as KServe, NVIDIA Triton Inference Server, Ray Serve, or BentoML... ...enabled infrastructure and NVIDIA CUDA environments....Full timeRemote work
- GPU Kernel Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting... ...routing Write and optimize code using CUDA, PTX assembly, and architecture-specific...Flexible hours
- Inference Runtime Software Engineer LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in... ...a wide range of CPU and GPU targets. You will also contribute... ...reliability across CPU, CUDA, Metal, Vulkan, and ROCm...Work at officeWork from homeFlexible hours
$229.9k - $262.4k
Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized...Full timePart timeLocal area$300k
...Campus AI Research Engineer (Intern) Jump Trading Group is committed to world class research... ...clusters, developing low-latency inference systems, or pushing the boundaries... ...programming languages: C / C++ / Python / CUDA and other low-level GPU languages. Build large-scale AI/...Full timeInternshipVisa sponsorshipWork visaFlexible hours- Job Title: AI Infrastructure Engineer Job Summary We are seeking an AI Infrastructure... ...for model training, inference, and deployment while optimizing... ...compute, storage, networking, and GPU resources. You will work... ...GPU infrastructure and CUDA environments. Experience with...Full timeRemote work
- NVIDIA seeks outstanding Senior High Performance AI Engineers to build the next generation of agentic AI systems for the CUDA ecosystem. You will work across the full stack—from training and improving models, to designing agent architectures and multi-agent systems, to...
- ...Founding AI Engineer (AI + Production) New York City (5 days on-site)... ...debugging eval failures, and scaling inference—then immediately applying... ...-informed models, running GPU-accelerated simulations, or collaborating... ...tools (NeMo, Modulus, CUDA optimization) ~ You're the...Full timeImmediate startWeekend work
- ...Position Title: Principal AI Platform Engineer Location: New York NY (Hybrid... ...Design and operate LLM inference and model-serving infrastructure... ...scalability. Optimize GPU utilization, token consumption... ...with NVIDIA NIM, NeMo, CUDA or GPU scheduling. Experience...
$204k - $259k
...In this hybrid role, you will report to an Engineering Manager. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale... ...: Expertise in C++ programming for GPU (CUDA or similar framework) Bachelor degrees in EECS...Full timeRemote work$176k
...Overview Current PhD, AI Engineering Internship Program - Summer 2027 Key Role Details This is a full-time paid... ...the-art techniques for optimizing AI/ML training or inference pipelines, including distributed GPU computational frameworks. Experience or academic...Full timePart timeSummer workInternshipLocal area$200k - $300k
...Trading is building a centralized AI function and is now hiring for... ...specific tasks, operating the inference infrastructure to run them on-... ...Partner with the agent engineering team to ensure the model layer... ...equivalent) ~ Experience managing GPU infrastructure (provisioning,...Immediate startWorldwideFlexible hours$120k - $160k
...CoreWeave is The Essential Cloud for AI™. Built for pioneers by... ...About the role ASoftware Engineer contributes to the design, implementation... ...hardware teams to evolve our GPU performance testing platform... ...infrastructure and training / inference. Why CoreWeave At...Permanent employmentFull timeTemporary workCasual workWork at officeFlexible hours$225k - $275k
...AI Data EngineerNew York, New York, United StatesSchonfeld Strategic... ...an experienced AI Data Engineer to join our Data Engineering team... ...for AI training and inference pipelines.Optimize data delivery... ...language (e.g. Java, Scala, Go, Rust)Data Engineering: 5+ years of...- Job Description: GPU Programming Software Engineer Role Type: Contractor Location: Remote In this role, you will apply your expertise as a GPU... ...implement, and optimize GPU-based software solutions using CUDA, WebGPU or GLSL. Profile and fine-tune GPU kernels and...For contractorsRemote work
$160k - $240k
Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779... ...or memory-aware serving. Familiarity with PyTorch and GPU software stacks such as CUDA and NCCL. Exposure to high-performance interconnects...Temporary workFor contractorsWork experience placement- ...FMX is seeking a Senior AI Research Engineer to help design, build, and scale Skynapse, FMX's agentic... ...fine-tuning, governed deployment, and inference optimization.Bachelor's degree in... ...Hugging Face TGI, ONNX, batching, caching, GPU utilization, and latency/cost optimization...
- Senior AI Engineer Location: Summit, NJ (Candidate has to go onsite only on need bases like once in a month or twice in a month) Role... ...experience, including feature stores, model deployment, and GPU inference. Experience with AI evaluation, observability, and cost optimization...
$229.9k - $262.4k
...Overview AI Engineer 5 ((AI Foundations, LLM Core and Agentic AI)... ...training, large language model inference, agents and multi-agent workflows... ...across AI systems, tracking GPU utilization, model throughput... ...with Python, Go, Scala, CUDA, or Java Preferred Qualifications...Full timePart timeLocal area- AI/ML Software Engineer At Gallatin, we are rebuilding logistics infrastructure for the national security... ...scale ML pipelines and real-time inference systems—while collaborating with cross... ...frameworks, including: PyTorch, TensorFlow, CUDA, Jupyter Notebooks Large Language...Local area
$152k - $241.5k
We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency... ...investigating, and resolving, training & inference performance end to endDebugging... ...:Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarkingExperience...Full timeRemote work$269.1k - $335.1k
Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing... ...model training, large language model inference, agents and multi-agent workflows,... ...experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications:...Full timePart timeLocal area- Staff AI Engineer This role will be based in Sunnyvale, San Francisco, or New York City. At... ...experiments that prove efficacy, and managing the GPU fleets that run them at scale in... ...efficiency or quality improvement (i.e. inference/training efficiency, engineer velocity,...For contractorsWork at officeImmediate startFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to AI Inference Engineer (GPU/Rust/CUDA). Be the first to apply!


