Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference & Optimization Engineer

Together AI

Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale. This role focuses on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for scalable LLM and vision model deployment. #J-18808-Ljbffr Together AI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference & Optimization Engineer in San Francisco, CA vacancy
  • $160k - $230k

     ...efficient and scalable inference for large language models...  ...LLMs). Our mission is to optimize inference frameworks,...  ...Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that...  ...to shape the future of LLM inference infrastructure... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    5 days ago
  • $170k - $245k

     ...AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software...  ...raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    3 days ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open... 
    Suggested

    Anyscale

    San Francisco, CA
    2 days ago
  • $190.9k - $232.8k

     ...leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong... 
    Suggested

    Jobleads-US

    San Francisco, CA
    5 days ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global...  ...caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep... 
    Suggested

    Sail

    San Francisco, CA
    1 day ago
  •  ...Francisco is seeking a talented engineer to design and implement...  ...fast and cost-efficient AI inference at global scale. You will be...  ...-performance schedulers and optimizing global routing while focusing...  ...has a strong background in distributed systems and is eager to engage... 

    Sail Research

    San Francisco, CA
    2 days ago
  •  ...AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical... 

    TensorScale AI

    San Francisco, CA
    2 days ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...operating system for distributed, heterogeneous AI hardware...  .... We believe that as LLM and multi-modal workloads...  ...distributed inference optimizations. THE OPPORTUNITY... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $200.8k - $251k

     ...technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models....  ...should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • $227.2k - $417k

     ...the Role:As a Software Engineer on the ML...  ...class machine learning inference platforms. These platforms...  ...support Deep Learning, LLM, and Search models. This...  ...throughput, and low latency distributed systems using...  ...approach to identifying & optimizing latency, cost, and efficiency... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    5 days ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more...  ...(Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch... 

    Inception

    San Francisco, CA
    2 days ago
  • Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in...  ...build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability. This role... 

    Inception

    San Francisco, CA
    2 days ago
  •  ...Team Our team analyzes inference stack performance across the...  ...understanding into performance optimizations and models that project...  ...from first principles about distributed systems, model inference,...  ...Enjoy collaborating with engineering and research teams to improve... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $298k - $368k

     ...with downstream teams on the optimization and integration into the Waymo...  ...set of sensors, enabling engineers like you to (1) develop methods...  ...You will: Design VLM/LLM model architecture and drive...  ...expertise in low-latency on-device inference techniques and a deep... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative...  ...will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm.... 

    Kindredventures

    San Francisco, CA
    2 days ago
  • $150k - $300k

    Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over... 
    Relocation package

    Prime Intellect

    San Francisco, CA
    1 day ago
  • $225k

    Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate... 

    Dormont Manufacturing Co

    San Francisco, CA
    1 day ago
  •  ...San Francisco, CA, is seeking a Member of Technical Staff for distributed systems to design, build, and operate the platform that schedules...  ...-growing environment. You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie Labs while shaping... 

    Acceler8 Talent

    San Francisco, CA
    1 day ago
  • Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from... 

    Jobot

    San Francisco, CA
    2 days ago
  •  ...seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation...  ...improve latency and throughput, optimize the inference stack to exhaust...  ...extend Kubernetes, Ray, and Slurm for distributed inference, ensuring reliability... 

    Causal Labs

    San Francisco, CA
    5 days ago
  • $300k

     ...of committed researchers, engineers, policy experts, and business...  ...About the role Our Inference team is responsible for...  ...models. We tackle complex, distributed systems challenges across...  ...traffic management systems LLM inference optimization, batching, and caching... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...About the Team OpenAI’s Inference team powers the...  ..., fast-moving team of engineers focused on delivering...  ...interaction. You'll build and optimize the systems that let...  ...that span networking, distributed compute, and high-...  ...tooling like vLLM, TensorRT-LLM, or custom model... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI products...  ...Stack team builds the distributed runtime that powers large-scale LLM inference across our...  ...to make new inference optimizations broadly available to customers... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...high-performance model inference and accelerating...  ...to drive the design, optimization, and scaling of our inference...  ...role, you’ll lead engineering efforts to ensure our...  ...development, and distributed inference best practices...  ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron,... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $249.5k - $273.5k

     ...Applied Research, Design, and Engineering leadership, you will lead a...  ...emerging agent frameworks, inference optimization techniques, retrieval...  ...Experience in AI, ML platforms, LLM systems, or agentic...  ...Strong foundations in scaling distributed systems and production-grade... 
    Work at office

    Dialpad

    San Francisco, CA
    4 days ago
  • $300k

     ...of committed researchers, engineers, policy experts, and business...  ...the Role The Cloud Inference team scales and optimizes Claude to serve the...  ...-performance, large-scale distributed systems serving millions of...  ...Strong familiarity with LLM inference optimization, batching... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...browsers running at scale, solving massive distributed systems challenges and making sure our...  .... Work closely with the rest of Engineering, gathering input and providing great...  ...databases, automated testing, performance optimization, and zero-downtime multi-region... 
    Full time
    Immediate start
    Relocation

    Browserbase

    San Francisco, CA
    1 day ago
  • $180k - $275k

     ...rapid shipping velocity. As Software Engineer on the Platform team, you'll...  ...Design and implement scalable APIs, distributed systems, and data infrastructure that serve...  ...traffic production systems and performance optimization ~ Track record shipping high-quality,... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    1 day ago
  •  ...Sensor Data Integration Engineers build the algorithms...  ...consistency of our data Optimize the performance of our...  ...parallel computing or distributed systems A bachelor'...  ...— orchestrating LLM-driven workflows for triage...  ...feed ML training and inference. Familiar with C++.... 
    Full time
    Work at office
    Work from home

    Mach9

    San Francisco, CA
    1 day ago
  •  ...constraints of robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure...  ...translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference & Optimization Engineer. Be the first to apply!