Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Frameworks Engineer High-Performance Serving

Together AI

Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimization, and software-hardware co-design to enable efficient deployment of LLMs and vision models. Join a research-driven team shaping AI infrastructure, collaborating with researchers to craft end-to-end serving pipelines and push the boundaries #J-18808-Ljbffr Together AI

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the LLM Inference Frameworks Engineer High-Performance Serving in San Francisco, CA vacancy
  • $160k - $230k

     ...efficient and scalable inference for large...  ...inference frameworks, algorithms, and...  ...boundaries of performance, scalability, and...  ...and Optimization Engineer to design,...  ...on low-latency, high-throughput inference...  ...the future of LLM inference...  ...high-performance serving.Apply CUDA graph... 
    Performance
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  •  ...owned AI. Our mission is to build highly scalable and efficient infrastructure...  ...seeking an expert in high‑performance LLM serving systems and inference optimization. In this role, you will...  ...debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.... 
    Performance

    NEAR.AI

    San Francisco, CA
    2 days ago
  • $170k - $245k

     ...About the roleAs a Distributed LLM Inference Engineer, you will help systems and...  ...push the boundaries of performance for inference at large scale...  ...Batch and Online inference at high scale which will be used by...  ...learning and deep learning frameworks (e.g. PyTorch)Solid... 
    Performance
    Work at office

    Anyscale

    San Francisco, CA
    20 hours ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal...  ...involves pushing the boundaries of performance for ML inference at scale. You'll work...  ...and familiarity with deep learning frameworks, ideally with experience in PyTorch and... 
    Performance

    Anyscale

    San Francisco, CA
    4 days ago
  •  ...build and operate the inference systems that serve our models in production...  .... This is an engineering role, not a research role...  ...serving large models at high throughput Own the performance characteristics of those...  ...inference / serving frameworks Experience with mixed... 
    Performance

    MakerMaker.AI

    San Francisco, CA
    1 day ago
  •  ...unicorn founders and senior engineers with deep expertise in 3D,...  ...a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This is a...  ...'ll work across the model-serving stack, designing novel inference frameworks, optimizing inference performance... 
    Performance
    Relocation
    Visa sponsorship
    Relocation package

    Reactor.am

    San Francisco, CA
    20 hours ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying...  ...for large-scale LLM training. Design distributed...  ...hardware configurations support high-performance training. Investigate and... 
    Performance
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    2 days ago
  • $180k - $270k

     ...elevate productivity and performance through note-taking...  ...building and deploying high-throughput, ultra-low-latency inference engines for large language...  ...experience with: Frontier Serving Frameworks: Deep, under-the-hood...  ...with modern LLM serving frameworks like... 
    Performance
    Full time
    Work at office
    Worldwide

    Plaud

    San Francisco, CA
    20 hours ago
  • Mirai Labs in San Francisco seeks engineers to join a senior team building the full on...  .... You will primarily work on uzu, our inference engine, and focus on supporting new modalities...  ...work, and experience in writing high-performance GPU kernels or Rust systems... 
    Performance
    Local area

    Mirai Labs

    San Francisco, CA
    3 days ago
  • Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize... 
    Performance

    Causal Labs

    San Francisco, CA
    2 days ago
  •  ...and production systems. You will own the bridge between research checkpoints and production-ready inference services, designing scalable APIs and high-performance serving. You will optimize GPU workloads, manage distributed systems, and collaborate with researchers to... 
    Performance

    Black Forest Labs Inc.

    San Francisco, CA
    20 hours ago
  • $200k - $300k

     ...research to systems engineering to product design...  ...looking for a Performance Engineer to make...  ...models train and serve as fast as the hardware...  ...Optimize inference and serving end to...  ...and low-latency, high-throughput sampling...  ...knowledge of ML framework internals (PyTorch... 
    Performance
    Full time

    World Labs

    San Francisco, CA
    2 days ago
  •  ...are forming small, highly capable product...  ...management, design, and engineering together with the...  ...quality issues, performance bottlenecks, and...  ...create evaluation frameworks, and establish monitoring...  ...— configuring LLM-based products,...  ...approximately 791,000 people serving clients in more... 
    Performance
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    San Francisco, CA
    4 days ago
  •  ...NEAR AI in San Francisco is seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries...  ...requires deep hands-on experience with inference engines such as SGLang, vLLM or TensorRT, GPU architectures,... 
    Performance

    NEAR.AI

    San Francisco, CA
    2 days ago
  • $300 per month

     ...strategies, and be part of a high-performing team that believes in each...  ...At Crusoe, our Production Engineering team ensures the...  ...services with a focus on serving and scaling LLM workloadsBuild automation...  ...distributed AI pipelines and inference servicesDefine, measure, and... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...worldwide.We’re a team of engineers, clinicians, and...  ...work helps care teams perform with greater...  ...improving and integrating high performance robotic AI...  ...to real-time onboard inference—while serving as a core contributor...  ...learning models, and frameworks such as PyTorch, TensorFlow... 
    Performance
    Local area
    Worldwide
    Flexible hours

    Intuitive Surgical

    San Francisco, CA
    2 days ago
  • $165k - $206k

     ...the world’s best engineers, scientists, designers...  ...the heavy lifting.Serve as a cross-...  ...supportLeverage Cursor with LLM-pair programming (...  ...time.Produce high-fidelity technical...  ..., integration framework (EIBs, Workday Studio...  ...procedures, CTEs, performance tuning) for data profiling... 
    Performance

    Juul

    San Francisco, CA
    4 days ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over...  ...ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation of $220,000 to $320... 
    Performance
    Local area

    Inference

    San Francisco, CA
    3 days ago
  • $264.8k - $331k

     ...Systems Research Engineer, Agent Post-training...  ...to reach the performance necessary for complex...  ...resources that serve all of our...  ...our training and inference framework. Post-train state...  ...least 1-3 years of LLM training in a production...  ...provide the high-quality data and... 
    Performance
    Full time
    Contract work
    For contractors
    For subcontractor
    Work at office

    Scale LLP

    San Francisco, CA
    3 days ago
  • $315k

     ...committed researchers, engineers, policy experts,...  ...and addressing performance issues across many...  ...research, training, and inference. A significant...  ...experience with: High performance, large...  ...architecture ML framework internals Language...  ...means a veteran who served on active duty in... 
    Performance
    Contract work
    For contractors
    For subcontractor
    Work at office
    Relocation
    Visa sponsorship
    Work visa
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $180k - $237.5k

     ...Senior Software Engineer – Site Controller, Energy...  ...batteries into a single, high-performance energy asset. We are...  ...fault-management frameworks, designing the state...  ...our "Pack Manager" to serve as a universal translator...  ...information, and inferences drawn from your PI. We... 
    Performance
    Full time
    Local area

    Redwood Materials

    San Francisco, CA
    20 hours ago
  • $138.68k - $174.43k

    APPLIED AI ENGINEER (1042) - Department of Technology...  ...process and shall serve at the discretion...  ..., integrating LLM APIs, and ensuring...  ...on, creative, and highly collaborative —...  ...logging, alerting, and performance metrics to ensure...  .... Use frameworks like LangChain, LlamaIndex... 
    Performance
    Permanent employment
    Full time
    Traineeship
    Work at office
    Remote work
    Work from home
    Flexible hours
    Night shift
    1 day per week

    City and County of San Francisco

    San Francisco, CA
    3 days ago
  • Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will...  ...have advanced C++, experience with parallel frameworks, and a strong track record in high-performance systems. This role is full-... 
    Performance
    Full time

    Vast.ai

    San Francisco, CA
    2 days ago
  • $166.6k - $208.3k

     ...deployment, real-time inference, observability, and...  ...themselves — and we serve low-latency, highly available scores to the decision engine that depends on them....  ...versioning, CI/CD with performance, bias, and consistency...  ...in Python, with API frameworks like FastAPI or FlaskExperience... 
    Performance

    Mercury

    San Francisco, CA
    3 days ago
  • $300 per month

     ...strategies, and be part of a high-performing team that believes in...  ...Hardware Systems Engineer to strengthen Crusoe’...  ...across training and inference - dense, MoE, long-...  ..., or data-analysis frameworks using Python, Shell,...  ...Experience with inference serving frameworks, training... 
    Performance
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...mission-critical inference for the world's most...  ...the platform engineers turn to to ship AI...  ...We believe that as LLM and multi-modal workloads...  ...for Disaggregated Serving, Wide Expert...  ...validate networking performance on bleeding-edge clusters...  ...experience with high-performance... 
    Performance
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    20 hours ago
  •  ...and more. We focus on high-performance model inference and accelerating...  ...this role, you’ll lead engineering efforts to ensure...  ...infrastructure for serving frontier AI models in...  ...Experience with inference frameworks like TensorRT, vLLM,...  ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron... 
    Performance
    Full time

    OpenAI

    San Francisco, CA
    20 hours ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across...  ...KV caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep... 
    Performance

    Sail

    San Francisco, CA
    3 days ago
  • $286.2k - $326.7k

     ...Senior Distinguished Engineer, AI Compute (...  ...experiences and scalable, high-performance AI infrastructure...  ...reimagine how we serve our customers and...  ...training, model inference and feature...  ...compute frameworks including Spark /...  ...diverse workloads from LLM pre-training and... 
    Performance
    Full time
    Part time
    Local area
    Remote work

    Capital One

    San Francisco, CA
    5 days ago
  •  ...looking for a Founding Engineer & CTO to be the...  ...from data pipelines to LLM orchestration to client...  ...CD pipelines, testing frameworks, and infrastructure-as...  ...engineering team, fostering a high-performance, collaborative culture...  ...tenant SaaS platforms serving enterprise clients.... 
    Performance
    Flexible hours

    NovusMinds AI

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Frameworks Engineer High-Performance Serving. Be the first to apply!