Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference Engineer - Scale AI at Speed

Anyscale

Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open-source technologies and contributing to community projects. Candidates should have a solid understanding of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale offers competitive compensation and extensive benefits, including healthcare coverage and stock options. #J-18808-Ljbffr Anyscale

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference Engineer - Scale AI at Speed in San Francisco, CA vacancy
  • $170k - $245k

     ...on a mission to democratize distributed computing and make it accessible...  ...accelerate the progress of AI applications out into the...  ...developer or data scientist can scale an ML application from their...  ...the roleAs a Distributed LLM Inference Engineer, you will help systems and... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    15 hours ago
  •  ...mission-critical inference for the world's most dynamic AI companies, like...  ...build the platform engineers turn to to ship...  ...system for distributed, heterogeneous AI...  ...believe that as LLM and multi-modal workloads scale, the network is...  ...operates at wire-speed. In this role... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  • $160k - $230k

     ...the RoleAt Together.ai, we are building...  ...efficient and scalable inference for large language...  ...and Optimization Engineer to design, develop, and optimize distributed inference engines that...  ...language models at scale. This role will...  ...shape the future of LLM inference infrastructure... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  •  ...San Francisco is seeking a talented engineer to design and implement robust systems...  ...that ensure fast and cost-efficient AI inference at global scale. You will be responsible for...  ...candidate has a strong background in distributed systems and is eager to engage in complex... 
    Suggested

    Sail Research

    San Francisco, CA
    4 days ago
  • $200.8k - $251k

    A leading AI technology company in San Francisco seeks a team member to build and optimize...  ...experience and solid software engineering skills, particularly in tools like CUDA and...  ...range of $200,800 - $251,000, along with comprehensive benefits. #J-18808-Ljbffr Scale AI
    Suggested
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • $160k - $194k

     ...What is Verse? The race to AI has become the race to...  ...built by pioneers in grid-scale batteries, energy markets,...  ...The Role As a Software Engineer focusing on Distributed Systems at Verse, you will...  ...Balance & Precision: We believe speed and perseverance must be accompanied... 
    Full time
    Remote work
    Flexible hours

    Ad Verse

    San Francisco, CA
    15 hours ago
  •  ...company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing...  ...infrastructure for large-scale multimodal models, focusing on high-...  ...product teams to push the boundaries of AI technology, ensuring reliable production... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    2 days ago
  •  ...the Team OpenAI’s Inference team powers the...  ...fast-moving team of engineers focused on delivering...  ...of what AI can do. We’re expanding...  ...models at scale. You’ll be part of...  ...span networking, distributed compute, and high-...  ...like vLLM, TensorRT-LLM, or custom model parallel... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,...  ...and help build the platform engineers turn to to ship AI products....  ...Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    15 hours ago
  • $300k

     ...interpretable, and steerable AI systems. We want...  ...researchers, engineers, policy experts,...  ...role Our Inference team is responsible...  ...tackle complex, distributed systems challenges...  ...performance, large-scale distributed systems...  ...systems LLM inference optimization... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    15 hours ago
  •  ...state-of-the-art AI models - unlocking...  ...performance model inference and accelerating research...  ...optimization, and scaling of our inference...  ...role, you’ll lead engineering efforts to ensure...  ...development, and distributed inference best...  ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron... 
    Full time

    OpenAI

    San Francisco, CA
    15 hours ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control...  ..., and explore KV caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep observability, tracing... 

    Sail

    San Francisco, CA
    3 days ago
  •  ...NEAR AI was started by Illia Polosukhin, co-author of the...  ...open‑source AI at a global scale. We are specifically seeking...  ...expert in high‑performance LLM serving systems and inference optimization. In this role,...  ...optimizing major inference engines such as SGLang, vLLM, or... 

    NEAR.AI

    San Francisco, CA
    2 days ago
  • $190k - $265k

     ...enabling data and AI teams to solve the...  ...business. Founded by engineers — and customer-...  ...interfacing with data to scaling our services and...  ...Foundation Model Inference team is the...  ...you will have:Build LLM infrastructure powering...  ...and efficiency of distributed AI... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...AI/ML Engineer (RL & Physical Systems) FLUIX is...  ...systems to power distribution, where milliseconds...  ...and real megawatt-scale infrastructure....  ...Support integration of LLM-based tools and...  ...knowledge distillation, inference orchestration, etc...  ...at startup speed. Bonus Points... 
    Weekend work

    Fluix AI

    San Francisco, CA
    1 day ago
  • $180k - $275k

     ...About the role You'll build and scale the application and data...  ...shipping velocity. As Software Engineer on the Platform team, you'll...  ...and implement scalable APIs, distributed systems, and data infrastructure...  ...systems, event pipelines, or AI-powered applications (Nice to... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    15 hours ago
  •  ...powers web browsing capabilities for AI agents and applications. We manage...  ...team keeps our browsers running at scale, solving massive distributed systems challenges and making sure our...  ...APIs. Work closely with the rest of Engineering, gathering input and providing great... 
    Full time
    Immediate start
    Relocation

    Browserbase

    San Francisco, CA
    15 hours ago
  •  ...robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale...  ...and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    15 hours ago
  • $166k - $225k

     ...running the world's best data and AI infrastructure platform so...  ...their business. Founded by engineers — and customer obsessed — we...  ...for interfacing with data to scaling our services and infrastructure...  ...building the next generation distributed data storage and processing... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $300k

     ...interpretable, and steerable AI systems. We want AI...  ...researchers, engineers, policy experts, and...  ...Role The Cloud Inference team scales and optimizes Claude...  ...performance, large-scale distributed systems serving millions...  ...familiarity with LLM inference optimization... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    15 hours ago
  •  ...out into multiple AI inference requests running in...  ..., our inference engineers and researchers build...  ...Keep long-running distributed training jobs...  ...managed GPU clusters at scale: NVIDIA hardware,...  ..., or TensorRT-LLM. Slurm or other HPC...  ...but notable. High-speed interconnects:... 
    Shift work

    Neura Market

    San Francisco, CA
    3 days ago
  • $176k - $209k

     ...DialpadDialpad is the AI platform for customer...  ...fostering an AI-native engineering culture.This position...  ...observability systems.Build & Scale: Design and deploy...  ...agent frameworks, LLM inference optimization, advanced...  ...foundations in scaling distributed systems and production... 
    Work at office

    Dialpad

    San Francisco, CA
    2 days ago
  • $153k - $376k

     ...designs into code, or iterating with AI. From idea to product, Figma...  ...everything we build. As a Software Engineer on our Infrastructure team, you’...  ...millions of people worldwide. We’re scaling fast, and we’re looking for experienced distributed systems engineers across a... 
    Minimum wage
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    Figma

    San Francisco, CA
    2 days ago
  • Crusoe is seeking a Senior Software Engineer to help scale Crusoe Cloud's container registry. You will own...  ...scalable cloud infrastructure. You’ll work across teams to balance speed and reliability, shape system trade-offs, and impact how customers’ AI #J-18808-Ljbffr Crusoe
    Full time

    Crusoe

    San Francisco, CA
    1 day ago
  • Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimization, and software... 

    Together AI

    San Francisco, CA
    2 days ago
  •  ...Data Integration Engineers build the algorithms...  ...transform large-scale geospatial datasets...  ...parallel computing or distributed systems A...  ...harnesses — orchestrating LLM-driven workflows...  ...ML training and inference. Familiar with...  ...design with AI-powered geospatial... 
    Full time
    Work at office
    Work from home

    Mach9

    San Francisco, CA
    15 hours ago
  • $216.2k - $270.25k

    Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative...  ...knowledge retrieval, inference, evaluation, and more...  ...Senior Full-Stack Engineer to help us build, scale...  ...Python, working with distributed systems, data pipelines, and ML/LLM components.Integrate... 
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $325k

    A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate...  ...with ML architectures, and experience with distributed systems. This role involves collaboration with researchers... 

    Jobleads-US

    San Francisco, CA
    4 days ago
  •  ...to build best-in-class AI agents with flexible and...  ...a world-class team of engineers, designers, marketers,...  ...for our customers as we scale. You will also be a steward...  ...Have deep intuition on distributed systems, databases,...  ...about trade-offs between speed, scalability, and... 
    Work at office
    Visa sponsorship
    Flexible hours

    Parallel

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference Engineer - Scale AI at Speed. Be the first to apply!