Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

LLM Inference Frameworks and Optimization Engineer

$160k - $230k

Togetherai

About the Role At Together.ai, we are building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). Our mission is to optimize inference frameworks, algorithms, and infrastructure, pushing the boundaries of performance, scalability, and cost‑efficiency. We are seeking an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale. This role will focus on low‑latency, high‑throughput inference, GPU/accelerator optimizations, and software‑hardware co‑design, ensuring efficient large‑scale deployment of LLMs and vision models. This role offers a unique opportunity to shape the future of LLM inference infrastructure, ensuring scalable, high‑performance AI deployment across a diverse range of applications. If you’re passionate about pushing the boundaries of AI inference, we’d love to hear from you! Responsibilities Inference Framework Development and Optimization Design and develop fault‑tolerant, high‑concurrency distributed inference engine for text, image, and multimodal generation models. Implement and optimize distributed inference strategies, including Mixture of Experts (MoE) parallelism, tensor parallelism, pipeline parallelism for high‑performance serving. Apply CUDA graph optimizations, TensorRT/TRT‑LLM graph optimizations, and PyTorch‑based compilation (torch.compile), and speculative decoding to enhance efficiency and scalability. Software‑Hardware Co‑Design and AI Infrastructure Collaborate with hardware teams on performance bottleneck analysis, co‑optimize inference performance for GPUs, TPUs, or custom accelerators. Work closely with AI researchers and infrastructure engineers to develop efficient model execution plans and optimize E2E model serving pipelines. Requirements Must‑Have: Experience: 3+ years of experience in deep learning inference frameworks, distributed systems, or high‑performance computing. Technical Skills: Familiar with at least one LLM inference framework (e.g., TensorRT‑LLM, vLLM, SGLang, TGI (Text Generation Inference)). Background knowledge and experience in at least one of the following: GPU programming (CUDA/Triton/TensorRT), compiler, model quantization, and GPU cluster scheduling. Deep understanding of KV cache systems like Mooncake, PagedAttention, or custom in‑house variants. Programming: Proficient in Python and C++/CUDA for high‑performance deep learning inference. Optimization Techniques: Deep understanding of Transformer architectures and LLM/VLM/Diffusion model optimization. Knowledge of inference optimization, such as workload scheduling, CUDA graph, compiled, efficient kernels. Soft Skills: Strong analytical problem‑solving skills with a performance‑driven mindset. Excellent collaboration and communication skills across teams. Nice‑to‑Have: Experience in developing software systems for large‑scale data center networks with RDMA/RoCE. Familiar with distributed filesystem (e.g., 3FS, HDFS, Ceph). Familiar with open source distributed scheduling/orchestration frameworks, such as Kubernetes (K8S). Contributions to open-source deep learning inference projects. About Together AI Together AI is a research‑driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co‑designing software, hardware, algorithms, and models. We have contributed to leading open‑source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure. Compensation We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full‑time position is: $160,000 - $230,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job‑related knowledge. Equal Opportunity Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. #J-18808-Ljbffr Togetherai

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the LLM Inference Frameworks and Optimization Engineer in San Francisco, CA vacancy
  • $160k - $230k

     ...enable efficient and scalable inference for large language models (LLMs). Our mission is to optimize inference frameworks, algorithms, and...  ...Frameworks and Optimization Engineer to design, develop, and optimize...  ...opportunity to shape the future of LLM inference infrastructure,... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    4 days ago
  • $170k - $245k

     ...million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for...  ...with deep learning and deep learning frameworks (e.g. PyTorch)Solid understanding of distributed... 
    Suggested
    Work at office

    Anyscale

    San Francisco, CA
    2 days ago
  • $190.9k - $232.8k

    A leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a... 
    Suggested

    Jobleads-US

    San Francisco, CA
    4 days ago
  • $298k - $368k

     ...downstream teams on the optimization and integration into...  ...of sensors, enabling engineers like you to (1) develop...  ...will: Design VLM/LLM model architecture and...  ...low-latency on-device inference techniques and a deep...  ...experience with deep learning frameworks (e.g. PyTorch, JAX)... 
    Suggested
    Full time
    Remote work

    Waymo

    San Francisco, CA
    1 day ago
  • Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance...  ...of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale... 
    Suggested

    Anyscale

    San Francisco, CA
    1 day ago
  • Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse...  ...targets someone who understands how LLM inference works under the hood...  ..., CUDA/Rocm, and model-serving frameworks, with strong Python and C++/Rust... 

    Jobot

    San Francisco, CA
    1 day ago
  • Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more...  .... You will extend orchestration frameworks (Kubernetes, Ray, SLURM) for distributed... 

    Inception

    San Francisco, CA
    1 day ago
  • Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in...  ...build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability. This role... 

    Inception

    San Francisco, CA
    1 day ago
  • LeoForce is seeking a Global Inference Library Engineer to design and optimize a high‑performance inference library for modern AI models. You...  ..., and low-level kernels, with exposure to frameworks like vLLM and TensorRT‑LLM. Join a technically focused startup building... 

    Leoforce

    San Francisco, CA
    1 day ago
  • $200.8k - $251k

     ...technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This... 
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  • $150k - $300k

    Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over... 
    Relocation package

    Prime Intellect

    San Francisco, CA
    5 days ago
  • Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global...  ...caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep... 

    Sail

    San Francisco, CA
    5 days ago
  • Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-...  ...high-throughput systems, latency-optimized serving, and distributed...  ...aware optimization, and deep learning frameworks like PyTorch or JAX. Collaboration... 

    Kindredventures

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...We're looking for an engineer to help build and maintain...  ...a high-performance inference library designed to support...  ...how modern LLM inference systems work...  ...Experience with LLM inference frameworks and model-serving...  ...developing, integrating, or optimizing performance-critical... 
    Local area

    Jobot

    San Francisco, CA
    1 day ago
  • $220k - $320k

    A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over 2 years of experience...  ..., fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation of $220,0... 
    Local area

    Inference

    San Francisco, CA
    5 days ago
  • $175k - $250k

    Global Inference Library Engineer Experience: Senior Level Salary: $175,000 - $2...  ...who understands how modern LLM inference systems work under...  ...with LLM inference frameworks and model-serving infrastructure...  ...developing, integrating, or optimizing performance-critical compute... 

    LeoForce

    San Francisco, CA
    1 day ago
  • Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This... 

    Vast.ai Inc.

    San Francisco, CA
    4 days ago
  • $249.5k - $273.5k

     ...Applied Research, Design, and Engineering leadership, you will lead a...  ...multi-agent orchestration frameworks capable of real-time...  ...emerging agent frameworks, inference optimization techniques, retrieval approaches...  ...Experience in AI, ML platforms, LLM systems, or agentic... 
    Full time
    Work at office

    Dialpad

    San Francisco, CA
    3 days ago
  • $250k - $300k

     ...That means owning the inference stack end to end: profiling...  ...go, bringing modern optimization techniques into real...  ...with customer engineering teams to tailor deployments...  ...the serving stack, from frameworks like vLLM and SGLang to...  ....Comfort with modern LLM serving frameworks such... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  •  ...powers mission-critical inference for the world's most...  ...help build the platform engineers turn to to ship AI...  ...that powers large-scale LLM inference across our platform...  ...to make new inference optimizations broadly available to...  ...learn new languages, frameworks, and systems as needed... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...high-performance model inference and accelerating...  ...to drive the design, optimization, and scaling of our inference...  ...role, you’ll lead engineering efforts to ensure our...  ...Experience with inference frameworks like TensorRT, vLLM, SGLang...  ..., ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron,... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • $197.3k - $225.1k

    Lead AI Engineer (FM Hosting, LLM Inference) Overview At Capital One, we are creating responsible and reliable AI systems, changing banking for good...  ...and more. Invent and introduce state-of-the-art LLM optimization techniques to improve the performance — scalability,... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    2 days ago
  •  ...Sensor Data Integration Engineers build the algorithms...  ...consistency of our data Optimize the performance of our...  ...— orchestrating LLM-driven workflows for triage...  ...scale data processing frameworks and cloud platforms (e...  ...feed ML training and inference. Familiar with C++.... 
    Full time
    Work at office
    Work from home

    Mach9

    San Francisco, CA
    1 day ago
  • $229.9k - $262.4k

    Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform) Overview: At Capital One, we are creating responsible and reliable AI systems...  ...PyTorch, and more. ~ Invent and introduce state-of-the-art LLM optimization techniques to improve the performance —... 
    Full time
    Part time
    Local area

    Capital One

    San Francisco, CA
    5 days ago
  •  ...About the Team Our team analyzes inference stack performance across the application,...  ...turn that understanding into performance optimizations and models that project performance and...  .... Enjoy collaborating with engineering and research teams to improve real production... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  • Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence...  ...framework responsible for large-scale LLM training. Design distributed training...  ...caches). Experience with data pipeline optimization, sharded datasets, or caching... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    4 days ago
  •  ...their business. Founded by engineers — and customer obsessed...  ...customers.Data Platform Optimization: Own the operational...  ...traditional sys-admin; you build frameworks, automation, and tooling...  .... Familiarity with LLM infrastructure, training/inference pipelines, or agentic... 
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $189.6k - $237k

     ...our internal distributed framework for large language model training and inference. The platform has been...  ...and evaluation of LLM's, as well as evaluation...  ...You will be building and optimizing the platform to enable...  ...systemsStrong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    4 days ago
  • $225k

    Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate... 

    Dormont Manufacturing Co

    San Francisco, CA
    5 days ago
  •  ...in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be responsible...  ...building high-performance schedulers and optimizing global routing while focusing on deep observability... 

    Sail Research

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to LLM Inference Frameworks and Optimization Engineer. Be the first to apply!