Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI Inference Engineer (GPU/Rust/CUDA)

Perplexity

Perplexity in New York is seeking an experienced engineer to join our team and own ML inference workloads, building a Rust/Python-based serving runtime and CUDA kernel work that scales to production. You will work with transformer-based models, enabling efficient weight loading, scheduling, and KV-cache management in our in-house infrastructure. You will read research papers, implement kernels, and diagnose production incidents in a fast-moving environment, collaborating across languages and #J-18808-Ljbffr Perplexity

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the AI Inference Engineer (GPU/Rust/CUDA) in New York, NY vacancy
  •  ...Senior GPU Systems / AI Infrastructure Engineer (NYC) Location: New York City (Hybrid / On...  ...-scale model training and inference. This role sits at the intersection...  ...optimise GPU kernels (CUDA / Triton / HIP) for large-...  ...high-performance C++ / Rust / Python systems ~... 
    Suggested
    Full time
    New York, NY
    more than 2 months ago
  • $220k

    We build and run the inference engine behind every Perplexity query and deploy dozens of model...  ...and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL - and we need another engineer...  ...management to support in API Gateway. GPU kernels migration to CuTe DSL. Port our... 
    Suggested

    Perplexity

    New York, NY
    12 hours ago
  • $229.9k - $286.2k

    AI Engineer 5 (FM Hosting, LLM Inference) At Capital One, we are creating responsible and reliable AI systems, changing...  ...reviews across AI systems, tracking GPU utilization, model throughput, and...  ...programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications... 
    Suggested
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    3 days ago
  • $250k - $300k

    Hudson River Trading (HRT) is seeking an AI Research Engineer (Inference) to join the HAIL team. HAIL (HRT AI...  ..., including but not limited to GPU kernel development, novel inference devices...  ...engineering skills, especially any of: CUDA/Triton/Pallas/CuTe DSL kernel development... 
    Suggested
    Work experience placement
    Work at office
    Local area
    Immediate start

    Hudson River Trading

    New York, NY
    3 days ago
  • $197.3k - $225.1k

     ...AI Engineer 4 At Capital One, we are creating responsible and reliable...  ..., large language model inference, agents and multi-agent workflows...  ...engineering to optimize GPU/TPU utilization and accelerate...  ...programming with Python, Go, Scala, CUDA, or Java Preferred... 
    Suggested
    Full time
    Part time
    Local area

    Capital One National Association

    New York, NY
    2 days ago
  • $100k - $150k

     ...delivering cloud, AI, data, and enterprise...  ...Title: AI Platform Engineer Location: 100%...  ...operate scalable AI inference platforms for production...  ..., LLM serving, GPU optimization, autoscaling...  ...in Python and Go, Rust, or C++. Experience...  ...workloads using CUDA , NVIDIA GPU technologies... 
    Full time
    H1b
    Local area
    Immediate start
    Remote work
    Visa sponsorship

    Bright Vision Technologies

    New York, NY
    2 days ago
  •  ...Overview Sr. Lead AI Engineer (FM Hosting, LLM Inference). At Capital One, we are creating responsible and reliable AI systems that are changing banking for good. We use machine learning to craft real-time, personalized customer experiences and build scalable, high... 
    Local area

    Capital One

    New York, NY
    2 days ago
  •  ...Job Title: AI DevOps Engineer Job Summary We are seeking an AI DevOps...  ...Optimize cloud resources, GPU utilization, and infrastructure...  ...such as KServe, NVIDIA Triton Inference Server, Ray Serve, or BentoML...  ...enabled infrastructure and NVIDIA CUDA environments.... 
    Full time
    Remote work

    Ova Technologies

    New York, NY
    3 days ago
  • GPU Kernel Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting...  ...routing Write and optimize code using CUDA, PTX assembly, and architecture-specific... 
    Flexible hours

    Baseten

    New York, NY
    21 hours ago
  • Inference Runtime Software Engineer LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in...  ...a wide range of CPU and GPU targets. You will also contribute...  ...reliability across CPU, CUDA, Metal, Vulkan, and ROCm... 
    Work at office
    Work from home
    Flexible hours

    Lm Studio

    New York, NY
    21 hours ago
  • $229.9k - $262.4k

    Senior Lead AI Engineer (FM Hosting, LLM Inference) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    1 day ago
  • $300k

     ...Campus AI Research Engineer (Intern) Jump Trading Group is committed to world class research...  ...clusters, developing low-latency inference systems, or pushing the boundaries...  ...programming languages: C / C++ / Python / CUDA and other low-level GPU languages. Build large-scale AI/... 
    Full time
    Internship
    Visa sponsorship
    Work visa
    Flexible hours

    Jump Trading

    New York, NY
    4 days ago
  • Job Title: AI Infrastructure Engineer Job Summary We are seeking an AI Infrastructure...  ...for model training, inference, and deployment while optimizing...  ...compute, storage, networking, and GPU resources. You will work...  ...GPU infrastructure and CUDA environments. Experience with... 
    Full time
    Remote work

    Ova Technologies

    New York, NY
    2 days ago
  • NVIDIA seeks outstanding Senior High Performance AI Engineers to build the next generation of agentic AI systems for the CUDA ecosystem. You will work across the full stack—from training and improving models, to designing agent architectures and multi-agent systems, to... 

    NVIDIA

    New York, NY
    2 days ago
  •  ...Founding AI Engineer (AI + Production) New York City (5 days on-site)...  ...debugging eval failures, and scaling inference—then immediately applying...  ...-informed models, running GPU-accelerated simulations, or collaborating...  ...tools (NeMo, Modulus, CUDA optimization) ~ You're the... 
    Full time
    Immediate start
    Weekend work

    Everstar

    New York, NY
    more than 2 months ago
  •  ...Position Title: Principal AI Platform Engineer Location: New York NY (Hybrid...  ...Design and operate LLM inference and model-serving infrastructure...  ...scalability. Optimize GPU utilization, token consumption...  ...with NVIDIA NIM, NeMo, CUDA or GPU scheduling. Experience... 

    Kutir Technologies

    New York, NY
    1 day ago
  • $204k - $259k

     ...In this hybrid role, you will report to an Engineering Manager. You will: Develop high-performance GPU primitives and abstractions to enable Waymo to scale...  ...: Expertise in C++ programming for GPU (CUDA or similar framework) Bachelor degrees in EECS... 
    Full time
    Remote work

    Waymo

    New York, NY
    3 days ago
  • $176k

     ...Overview Current PhD, AI Engineering Internship Program - Summer 2027 Key Role Details This is a full-time paid...  ...the-art techniques for optimizing AI/ML training or inference pipelines, including distributed GPU computational frameworks. Experience or academic... 
    Full time
    Part time
    Summer work
    Internship
    Local area

    Capital One

    New York, NY
    18 days ago
  • $200k - $300k

     ...Trading is building a centralized AI function and is now hiring for...  ...specific tasks, operating the inference infrastructure to run them on-...  ...Partner with the agent engineering team to ensure the model layer...  ...equivalent) ~ Experience managing GPU infrastructure (provisioning,... 
    Immediate start
    Worldwide
    Flexible hours

    DV Trading

    New York, NY
    24 days ago
  • $120k - $160k

     ...CoreWeave is The Essential Cloud for AI™. Built for pioneers by...  ...About the role ASoftware Engineer contributes to the design, implementation...  ...hardware teams to evolve our GPU performance testing platform...  ...infrastructure and training / inference. Why CoreWeave At... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    New York, NY
    2 days ago
  • $225k - $275k

     ...AI Data EngineerNew York, New York, United StatesSchonfeld Strategic...  ...an experienced AI Data Engineer to join our Data Engineering team...  ...for AI training and inference pipelines.Optimize data delivery...  ...language (e.g. Java, Scala, Go, Rust)Data Engineering: 5+ years of... 

    Schonfeld

    New York, NY
    16 hours ago
  • Job Description: GPU Programming Software Engineer Role Type: Contractor Location: Remote In this role, you will apply your expertise as a GPU...  ...implement, and optimize GPU-based software solutions using CUDA, WebGPU or GLSL. Profile and fine-tune GPU kernels and... 
    For contractors
    Remote work

    ESR Healthcare

    New York, NY
    21 hours ago
  • $160k - $240k

    Senior Software Engineer - AI Inference Location New York Business Area Engineering and CTO Ref # 10050779...  ...or memory-aware serving. Familiarity with PyTorch and GPU software stacks such as CUDA and NCCL. Exposure to high-performance interconnects... 
    Temporary work
    For contractors
    Work experience placement

    Bloomberg

    New York, NY
    4 days ago
  •  ...FMX is seeking a Senior AI Research Engineer to help design, build, and scale Skynapse, FMX's agentic...  ...fine-tuning, governed deployment, and inference optimization.Bachelor's degree in...  ...Hugging Face TGI, ONNX, batching, caching, GPU utilization, and latency/cost optimization... 

    Cantor Fitzgerald

    New York, NY
    2 days ago
  • Senior AI Engineer Location: Summit, NJ (Candidate has to go onsite only on need bases like once in a month or twice in a month) Role...  ...experience, including feature stores, model deployment, and GPU inference. Experience with AI evaluation, observability, and cost optimization... 

    E-Solutions

    New York, NY
    3 days ago
  • $229.9k - $262.4k

     ...Overview AI Engineer 5 ((AI Foundations, LLM Core and Agentic AI)...  ...training, large language model inference, agents and multi-agent workflows...  ...across AI systems, tracking GPU utilization, model throughput...  ...with Python, Go, Scala, CUDA, or Java Preferred Qualifications... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    6 days ago
  • AI/ML Software Engineer At Gallatin, we are rebuilding logistics infrastructure for the national security...  ...scale ML pipelines and real-time inference systems—while collaborating with cross...  ...frameworks, including: PyTorch, TensorFlow, CUDA, Jupyter Notebooks Large Language... 
    Local area

    Gallatin AI, Inc.

    New York, NY
    21 hours ago
  • $152k - $241.5k

    We are seeking a Senior AI/ML Performance and Efficiency Engineer, GPU Clusters at NVIDIA to join our AI Efficiency...  ...investigating, and resolving, training & inference performance end to endDebugging...  ...:Background with NVIDIA GPUs, CUDA Programming, NCCL and MLPerf benchmarkingExperience... 
    Full time
    Remote work

    Nvidia

    New York, NY
    4 days ago
  • $269.1k - $335.1k

    Staff AI Engineer At Capital One, we are creating responsible and reliable AI systems, changing...  ...model training, large language model inference, agents and multi-agent workflows,...  ...experience programming with Python, Go, Scala, CUDA, or Java Preferred Qualifications:... 
    Full time
    Part time
    Local area

    Capital One

    New York, NY
    4 days ago
  • Staff AI Engineer This role will be based in Sunnyvale, San Francisco, or New York City. At...  ...experiments that prove efficacy, and managing the GPU fleets that run them at scale in...  ...efficiency or quality improvement (i.e. inference/training efficiency, engineer velocity,... 
    For contractors
    Work at office
    Immediate start
    Flexible hours

    LinkedIn

    New York, NY
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI Inference Engineer (GPU/Rust/CUDA). Be the first to apply!