Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference Engineer

$170k - $245k

Anyscale

About AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.As part of this role, you willIterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of AnyscaleWork across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open sourceFollow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you haveFamiliarity with running ML inference at large scale with high throughput and low latencyFamiliarity with deep learning and deep learning frameworks (e.g. PyTorch)Solid understanding of distributed systems, ML inference challengesBonus points!ML Systems knowledgeExperience using Ray Work closely with community on LLM engines like vLLM, TensorRT-LLMContributions to deep learning frameworks (PyTorch, TensorFlow)Contributions to deep learning compilers (Triton, TVM, MLIR)Prior experience working on GPUs / CUDACompensationAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted. This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents401k Retirement PlanEducation & Wellbeing StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in-office meals coveredAnyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and SpanishCompensation Range: $170K - $245KLocationSan Francisco; Palo AltoEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensationTarget Base Salary:$170K – $245K • Offers EquityAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99%401k Retirement PlanWellness & Education StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in office meals covered

Vacancy posted 17 hours ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference Engineer in San Francisco, CA vacancy
  • $160k - $230k

     ...to enable efficient and scalable inference for large language models (LLMs)....  ...anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support...  ...to shape the future of LLM inference infrastructure, ensuring... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    2 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI...  ...us and help build the platform engineers turn to to ship AI products....  ...the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    21 hours ago
  •  ...We are specifically seeking an expert in high‑performance LLM serving systems and inference optimization. In this role, you will push the boundaries...  ...with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT. Deep knowledge of state... 
    Suggested

    NEAR.AI

    San Francisco, CA
    2 days ago
  • $227.2k - $417k

     ...About the Role:As a Software Engineer on the ML Infrastructure team...  ...world-class machine learning inference platforms. These platforms power...  ...that support Deep Learning, LLM, and Search models. This...  ...throughput, and low latency distributed systems using ScalaBuild reusable... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    2 days ago
  • $180k - $250k

     ...unified platform where high-performance inference, orchestration, and observability come...  ...role:  You are an experienced software engineer who thrives on building large-scale...  ...You have deep expertise in large scale distributed systems that deal with high complexity,... 
    Suggested
    Currently hiring
    Remote work
    Relocation package

    features and labels

    San Francisco, CA
    2 days ago
  • $249.5k - $273.5k

     ...Applied Research, Design, and Engineering leadership, you will lead a...  ...Evaluate emerging agent frameworks, inference optimization techniques,...  ...in AI, ML platforms, LLM systems, or agentic architectures...  ...Strong foundations in scaling distributed systems and production-grade... 
    Work at office

    Dialpad

    San Francisco, CA
    1 day ago
  • $117.2k - $223.9k

     ...unparalleled customer experiences. Join our team of talented engineers and help us advance the integration of Salesforce applications...  ...vulnerability management, infrastructure security, and the security of distributed and scalable distributed systems. This role requires hands-on... 
    Full time

    Salesforce

    San Francisco, CA
    2 days ago
  • $142.2k - $204.6k

    P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize...  ..., ensuring our large language model (LLM) serving systems are fast, scalable,...  ...versioningIntegrate with federated, distributed inference infrastructure -... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $153k - $376k

     ...infrastructure is at the heart of everything we build. As a Software Engineer on our Infrastructure team, you’ll help design, build, and...  .... We’re scaling fast, and we’re looking for experienced distributed systems engineers across a variety of teams. Whether you’re passionate... 
    Minimum wage
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    Figma

    San Francisco, CA
    2 days ago
  • $300k

     ...group of committed researchers, engineers, policy experts, and business...  ...About the role Our Inference team is responsible for building...  ...models. We tackle complex, distributed systems challenges across...  ...traffic management systems LLM inference optimization, batching... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    21 hours ago
  •  ...improve their business. Founded by engineers — and customer obsessed — we leap...  ...Software Engineer or SRE in highly distributed, multi-cloud environments....  ...infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks... 
    Worldwide

    DataBricks

    San Francisco, CA
    4 days ago
  • $166k - $225k

     ...use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve...  ...team at Databricks, you will be building the next generation distributed data storage and processing systems that can outperform specialized... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies...  ...us and help build the platform engineers turn to to ship AI products....  ...Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    21 hours ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our...  ...a small, fast-moving team of engineers focused on delivering a world...  ...systems that span networking, distributed compute, and high-throughput...  ...tooling like vLLM, TensorRT-LLM, or custom model parallel systems... 
    Full time

    OpenAI

    San Francisco, CA
    21 hours ago
  •  ...focus on high-performance model inference and accelerating research...  ...systems. In this role, you’ll lead engineering efforts to ensure our largest...  ..., CUDA development, and distributed inference best practices....  ...AMD GPUs, ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron, MPI, or... 
    Full time

    OpenAI

    San Francisco, CA
    21 hours ago
  • $180k - $275k

     ...technical investments with rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate across frontend,...  ...What you'll do Design and implement scalable APIs, distributed systems, and data infrastructure that serve millions of users... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    21 hours ago
  •  ...This team keeps our browsers running at scale, solving massive distributed systems challenges and making sure our platform is fast,...  ...with developer-friendly APIs. Work closely with the rest of Engineering, gathering input and providing great support so every team can... 
    Full time
    Immediate start
    Relocation

    Browserbase

    San Francisco, CA
    21 hours ago
  • $300k

     ...group of committed researchers, engineers, policy experts, and business...  ...the Role The Cloud Inference team scales and optimizes Claude...  ...-performance, large-scale distributed systems serving millions of users...  ...Strong familiarity with LLM inference optimization, batching... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    21 hours ago
  • $189.6k - $237k

     ...RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been...  ...automatic training and evaluation of LLM's, as well as evaluation of data...  ...ML systemsStrong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  •  ...Mach9, Sensor Data Integration Engineers build the algorithms and...  ...foundation in parallel computing or distributed systems A bachelor's...  ...agent harnesses — orchestrating LLM-driven workflows for triage,...  ...pipelines that feed ML training and inference. Familiar with C++.... 
    Full time
    Work at office
    Work from home

    Mach9

    San Francisco, CA
    21 hours ago
  • $300 per month

     ...RoleAt Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and cloud services to...  ...focus on serving and scaling LLM workloadsDefine, measure, and...  ...optimize large-scale training and inference clustersAutomate... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  •  ...software architecture for our core platform. This person will be the technical conscience for large‑scale, distributed systems, and will collaborate closely with engineering leads, product owners, and infrastructure teams. You will design, evolve, and enforce architectural... 
    Remote work
    Home office
    Flexible hours

    Alteryx, Inc.

    San Francisco, CA
    1 day ago
  • $300 per month

     ...Role:At Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and hands-on experience...  ...focus on serving and scaling LLM workloadsBuild automation and...  ...distributed AI pipelines and inference servicesDefine, measure, and... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $172.5k - $260.1k

     ...Description: Salesforce has immediate opportunities for Lead software engineers who want their lines of code to have significant and...  ...Envision and Build new and exciting components/frameworks in distributed filesystems in an ever-growing and evolving market technology... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    1 day ago
  •  ...combined with seamless physical skills under the practical constraints of robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    21 hours ago
  • Electrical Engineer - Power Distribution and AnalysisHED is looking to add an experienced team member to our group of talented Electrical Engineers. The candidate should have broad electrical engineer expertise, with a focus on power distribution system and power system... 
    Work at office
    Work from home
    Flexible hours

    HED

    San Francisco, CA
    3 days ago
  • B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a... 

    B Capital

    San Francisco, CA
    3 days ago
  • $160k - $200k

     ...success and ability to scale. This role reports to the Senior Engineering Manager of Realtime Infrastructure. What You'll Be Doing...  ...Build and operate large-scale, reliable and performant distributed systems. Collaborate with product teams to create new features... 
    Full time
    Relocation
    Relocation package

    Discord

    San Francisco, CA
    3 days ago
  • $206.4k - $379.1k

     ...Services team is seeking a Principal Service Engineer to serve as the technical lead for our...  ...flagship products.Design and architect inference infrastructure for enterprise-scale...  ....Hands-on expertise with Kubernetes, distributed systems, and MLOps platforms.Preferred... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    2 days ago
  • $190k - $265k

     ...improve their business. Founded by engineers — and customer-obsessed — we...  ...across real-time and batch inference, powering model inference at...  ...impact you will have:Build LLM infrastructure powering large...  ..., latency, and efficiency of distributed AI workloadsCollaborate with... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    17 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!