Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Remote Inference Runtime Engineer for LLM & Multimodal AI

Inferact

New York, NY
  • Remote job

Inferact is seeking an inference runtime engineer to advance LLM and diffusion model serving. You will optimize how models execute across diverse hardware, shaping the core of vLLM and enabling faster AI inference. This remote role embraces flexible timezones with Pacific overlap for critical syncs, and compensation includes salary plus equity. The ideal candidate will have deep knowledge of transformer models, strong Python/PyTorch skills, and hands-on experience with LLM inference systems. #J-18808-Ljbffr Inferact

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Remote Inference Runtime Engineer for LLM & Multimodal AI in New York, NY vacancy
  •  ...are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte,...  ....Role Overview We are seeking an AI Infrastructure Runtime Engineer to build and maintain large...  .... While many positions offer remote or hybrid work options, these arrangements... 
    Remote work
    Work at office
    Flexible hours

    NTT DATA

    Charlotte, NC
    3 days ago
  • OpenAI is seeking a systems-focused engineer to design and implement the LLM inference runtime for frontier models on our custom silicon. You'll bridge model execution...  ...and to deliver reliable, production-grade performance on OpenAI's AI #J-18808-Ljbffr AI Chopping Block
    Suggested

    AI Chopping Block

    Eastern, KY
    3 days ago
  • $236k - $330k

     ...usher in this new era, we seek AI-native thinkers across every...  ...the state of the art in LLM inference systems and optimization.Our...  ...from distributed serving and runtime systems to GPU kernels and model...  ...tuning. We embrace AI-native engineering, using AI not only as the workload... 
    Suggested

    Snowflake

    Bellevue, WA
    3 days ago
  • $160k - $250k

     ...for large-language-model inference and training, with HW/...  ...from the others. The runtime owns the host-side stack...  ...downstream consumersBuild the LLM inference serving stack...  ...the Python surfaces ML engineers actually use — and hit...  ...+ up to 3 weeks remote workHealth: Company-subsidized... 
    Remote work
    Daily paid
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    4 days ago
  • $150k - $220k

     ...Runpod is the AI Developer Cloud. More than...  ...than 20 billion inference requests. We closed...  ...on. We're a small, remote-first team. We take...  ...looking for a ML Systems Engineer, Inference. We want...  ...the world to run LLM inference, meaning...  ...production-ready runtimes, configurations, and... 
    Remote work
    Full time

    Runpod

    Remote
    14 days ago
  •  ...Job Title: LLM Engineer (Large Language Model Engineer...  ...applications and AI-powered solutions. The...  ...optimization, and inference strategies. Strong...  ...Location: Remote / Hybrid / On-site...  ...assistants. Knowledge of multimodal AI models (text, image... 
    Remote work
    Full time

    Ova Technologies

    New York, NY
    1 day ago
  • $175k - $200k

     ...the quality layer that takes generative AI from experiment to enterprise reality at...  ...one of the fastest growing areas in AI engineering. LLM evaluation, fine-tuning, red-teaming, prompt...  ...production. \n \n This is a fully remote role \n \n Ready to stop hoping... 
    Remote work

    Saragossa

    West Virginia
    4 days ago
  • $87.95k - $203.95k

     ...apply now.We are currently seeking a API LLM Integration / ReactJS Engineer - Hybrid to join our team in Santa...  ...Engineer will work with our AI team, software engineers, and business...  ...roles. The starting pay range for this remote role is $$87,952 - $203,954 This range... 
    Remote work
    Temporary work
    Work at office
    Flexible hours

    NTT DATA

    Santa Clara, CA
    1 day ago
  • $90 - $120 per hour

    Gridnaut Recruiting is hiring a remote MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) contractor (pay $90-$120/hr). Contribute to frontier AI research and evaluation work. Ideal candidates: Have 2+ years of hands-on professional experience in ML systems,... 
    Remote work
    Temporary work
    For contractors
    Weekday work

    Gridnaut Recruiting

    Remote
    24 days ago
  • $190k - $260k

     ...leading security-first enterprise AI company. We build cutting-edge...  ...is a team of researchers, engineers, designers, and more, who are...  ...influence latency and throughput of inference.Strong understanding or...  ...traveling to other offices if you are remote, plus an annual company... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Local area
    Home office

    Cohere

    New York, NY
    1 day ago
  • $184k - $287.5k

     ...unlimited potential of AI to define the next era...  ...seeking an AI Compiler Engineer with deep expertise in...  ...efficiency, and advancing LLM-enabled workflows for...  ...measurable outcomes such as runtime gains, compile-time...  ...US, TX, Austin; US, TX, Remote; US, CA, Remote; US, WA... 
    Remote work
    Full time

    Nvidia

    Austin, TX
    1 day ago
  • $94.58k - $110k

     ...opportunity to join our team as an Engineer II.In this role, the AI Engineer will design,...  ...NYU Langone Healths Remote Patient Monitoring (RPM) initiatives...  ..., embedding models, and LLM providers, balancing...  ...performance, and cost.Optimize inference performance and cost efficiency... 
    Remote work
    Full time

    NYU Langone Medical Center

    New York, NY
    3 days ago
  • $90 - $120 per hour

     ...MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) is a remote engineering review track for evaluating production code...  ...debugging traces, and developer-facing AI outputs against real-world...  ...failure modes (compile error, runtime crash, off-by-one, security... 
    Remote job
    For contractors
    10 hours per week

    AuraOne Human Data

    Remote
    1 day ago
  •  ...mass-producible, multimodal missile seeker that...  ...supply-chain security engineered in from day one....  .... This is an AI-native team. Through...  ..., TensorRT inference, and shared-memory...  ...heterogeneous-compute runtimes beyond CUDA. Direct...  ...Angeles, CA HQ; remote work is not... 
    Remote work
    Weekend work

    Furientis

    Los Angeles, CA
    18 days ago
  • $100k - $150k

     ...LLM Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States....  ...ML venues. Experience with multimodal model fine-tuning. Familiarity... 
    Remote work
    Full time
    H1b
    Local area
    Visa sponsorship

    Bright Vision Technologies

    Edison, NJ
    4 days ago
  • Snowflake is seeking AI-native thinkers for the AI Research team to advance LLM inference systems and optimization. You will work on high-performance, adaptive inference across distributed serving, GPU kernels, and system co-design, collaborating with researchers and product... 

    Snowflake Computing

    Bellevue, WA
    4 days ago
  • $184k - $287.5k

     ...motivated Deep Learning engineer to bring advanced...  ...technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX,...  ...up to 100K GPUs to inference down at microsecond...  ...one communication runtime (NCCL, NVSHMEM, MPI...  ..., TX, Austin; US, Remote; US, NC, DurhamType... 
    Remote work
    Full time

    Nvidia

    Westford, MA
    1 day ago
  • BairesDev is seeking a Senior AI Engineer to design and ship pipelines that connect to LLM providers, driving AI-enabled features and platform capabilities across products. You will lead AI-driven initiatives from concept to production, build agentic systems with tool... 
    Remote job

    BairesDev

    New York, NY
    2 days ago
  • $130.33k - $195.5k

    ## AI Automation EngineerApply: Fully Remote: Anywhere in the U.S.: Full time: Posted...  ...& Automation Engineer is a hands-on technical...  ...data, as well as multimodal data such as...  ...mind. Treat cost per inference and infrastructure...  ...Transformers; experience with LLM-based frameworks... 
    Remote work
    Full time
    Work experience placement

    Alignment Health

    Eastern, KY
    4 days ago
  • A tech company is seeking a skilled Prompt Engineer for a remote position in the European Union. The ideal candidate will design, test, and optimize prompts to enhance AI model performance. Responsibilities include collaborating with data scientists, ensuring compliance... 
    Remote work
    Flexible hours

    Codertal

    United States
    5 days ago
  • $193.3k - $261.5k

     ...looking for a Senior Inference Engineer to own inference for real...  ...AI. This is a full-stack...  ...building thereal-time runtime that serves it within...  ...path for large-scale multimodal models — attentionand...  ...fall outside standard LLM serving patterns — sustained... 
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    2 days ago
  • $170k - $245k

     ...have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale, we’re building...  ...$250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the... 
    Work at office

    Anyscale

    San Francisco, CA
    4 days ago
  • $100k - $150k

     ...AI Risk Engineer – Remote Bright Vision Technologies is a technology consulting...  ...specifically targeting LLM and AI-powered application...  ...model endpoints. Implement runtime detection and response capabilities...  ...weights, datasets, and inference dependencies.... 
    Remote work
    Full time
    H1b
    Local area
    Immediate start
    Visa sponsorship

    Bright Vision Technologies

    New Albany, OH
    a month ago
  • $160k - $275k

     ...BuildingMatX is building next-generation AI compute infrastructure for large-scale LLM training and inference. We are looking for a hands-on Mechanical Engineer to design, develop, and validate...  ...2 company Holidays + up to 3 weeks remote workHealth: Company-subsidized... 
    Remote work
    Daily paid
    Full time
    Work experience placement
    Work at office
    Local area
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    4 days ago
  • $160k - $275k

     ...BuildingMatX is building next-generation AI compute infrastructure for large-scale LLM training and inference. We’re looking for an...  ..., hands-on NPI Manufacturing Engineer to lead our AI chip, compute...  ...company Holidays + up to 3 weeks remote workHealth: Company-subsidized... 
    Remote work
    Daily paid
    Full time
    Work experience placement
    Work at office
    Local area
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    4 days ago
  • $145k - $200k

     ...an experienced Software Engineer to join a newly-formed...  ...and deploying advanced AI systems to the tactical...  ...evaluating, and deploying LLM agents to edge hardware...  ...AI models and runtime performance for edge hardwareFamiliarity...  ...roles that allow for “Remote” work on an exceptional... 
    Remote work
    Full time
    Work experience placement
    Work at office
    Work from home
    Relocation package

    Palantir Technologies

    Seattle, WA
    3 days ago
  • $220k - $280k

     ...is a fast-scaling AI infrastructure company...  ..., audio, and multimodal workloads. Freshly...  ...forward-deployed engineering role where you're...  ...SGLang, and TensorRT-LLM (SFT baseline, DPO...  ...with open-model LLM inference and/or fine-tuning...  ...Location ~ Remote-friendly with hubs... 
    Remote job
    Full time
    H1b
    Visa sponsorship
    Flexible hours
    Day shift

    Lavendo

    San Mateo, CA
    7 days ago
  • AI/ML Ops EngineerLocation: Remote / Hybrid (Client-Facing Consulting Engagement)Employment...  ...an experienced AI/ML Engineer to design, deploy, and operate...  ...real-time APIs, batch inference pipelines, and feature stores...  ...or Large Language Model (LLM) solutions in production... 
    Remote work
    Full time
    Contract work
    Local area
    Flexible hours

    Slalom

    Houston, TX
    1 day ago
  •  ...areMoveworks is the Agentic AI Assistant platform that...  ...Moveworks’ Reasoning Engine and natural language...  ...for building and serving LLM’s at Moveworks. This...  ...distributed training and inference pipeline for large...  ...Work personas (flexible, remote, or required in office)... 
    Remote work
    Permanent employment
    Work at office
    Flexible hours

    Moveworks

    Mountain View, CA
    3 days ago
  • $145k - $165k

     ...company delivering cloud, AI, data, and enterprise solutions...  ...Job Title: ML Systems Engineer Location: 100% Remote (U.S.) Position Type: Full...  ...performance, highly reliable inference platforms for serving large...  ...Hands-on experience with LLM or large model inference frameworks... 
    Remote work
    Full time
    H1b
    Local area
    Immediate start
    Visa sponsorship

    Bright Vision Technologies

    Austin, TX
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Remote Inference Runtime Engineer for LLM & Multimodal AI. Be the first to apply!