Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference Engineer

$170k - $245k

Anyscale

About AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.As part of this role, you willIterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of AnyscaleWork across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open sourceFollow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you haveFamiliarity with running ML inference at large scale with high throughput and low latencyFamiliarity with deep learning and deep learning frameworks (e.g. PyTorch)Solid understanding of distributed systems, ML inference challengesBonus points!ML Systems knowledgeExperience using Ray Work closely with community on LLM engines like vLLM, TensorRT-LLMContributions to deep learning frameworks (PyTorch, TensorFlow)Contributions to deep learning compilers (Triton, TVM, MLIR)Prior experience working on GPUs / CUDACompensationAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted. This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents401k Retirement PlanEducation & Wellbeing StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in-office meals coveredAnyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and SpanishCompensation Range: $170K - $245KLocationSan Francisco; Palo AltoEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensationTarget Base Salary:$170K – $245K • Offers EquityAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99%401k Retirement PlanWellness & Education StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in office meals covered

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference Engineer in San Francisco, CA vacancy
  • $160k - $230k

     ...to enable efficient and scalable inference for large language models (LLMs)....  ...anInference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support...  ...to shape the future of LLM inference infrastructure, ensuring... 
    Suggested
    Full time

    Together AI

    San Francisco, CA
    5 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI...  ...us and help build the platform engineers turn to to ship AI products....  ...the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...We are specifically seeking an expert in high‑performance LLM serving systems and inference optimization. In this role, you will push the boundaries...  ...with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT. Deep knowledge of state... 
    Suggested

    NEAR.AI

    San Francisco, CA
    5 days ago
  • $227.2k - $417k

     ...About the Role:As a Software Engineer on the ML Infrastructure team...  ...world-class machine learning inference platforms. These platforms power...  ...that support Deep Learning, LLM, and Search models. This...  ...throughput, and low latency distributed systems using ScalaBuild reusable... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    5 days ago
  • $200k - $300k

     ..., governance, and trust become the real engineering challenge. Arcade is the MCP runtime...  ...vector database team at Redis, shipped 100+ LLM applications, and is a contributor to...  ...assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis... 
    Suggested
    Work at office
    Shift work

    Arcade AI, Inc

    San Francisco, CA
    3 days ago
  •  ...the frontier forward. The Production Engineering Team Examples of key exciting problems...  ..., actual state inspection, and distributed command execution. One interface for the...  ...move on. You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks... 
    Local area

    FluidStack

    San Francisco, CA
    4 days ago
  • $249.5k - $273.5k

     ...Applied Research, Design, and Engineering leadership, you will lead a...  ...Evaluate emerging agent frameworks, inference optimization techniques,...  ...in AI, ML platforms, LLM systems, or agentic architectures...  ...Strong foundations in scaling distributed systems and production-grade... 
    Work at office

    Dialpad

    San Francisco, CA
    4 days ago
  • $117.2k - $223.9k

     ...unparalleled customer experiences. Join our team of talented engineers and help us advance the integration of Salesforce applications...  ...vulnerability management, infrastructure security, and the security of distributed and scalable distributed systems. This role requires hands-on... 
    Full time

    Salesforce

    San Francisco, CA
    5 days ago
  • $142.2k - $204.6k

    P-1284About This RoleAs a software engineer for GenAI inference, you will help design, develop, and optimize...  ..., ensuring our large language model (LLM) serving systems are fast, scalable,...  ...versioningIntegrate with federated, distributed inference infrastructure -... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    5 days ago
  • $153k - $376k

     ...infrastructure is at the heart of everything we build. As a Software Engineer on our Infrastructure team, you’ll help design, build, and...  .... We’re scaling fast, and we’re looking for experienced distributed systems engineers across a variety of teams. Whether you’re passionate... 
    Minimum wage
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    Figma

    San Francisco, CA
    5 days ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our...  ...a small, fast-moving team of engineers focused on delivering a world...  ...systems that span networking, distributed compute, and high-throughput...  ...tooling like vLLM, TensorRT-LLM, or custom model parallel systems... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies...  ...us and help build the platform engineers turn to to ship AI products....  ...Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $300k

     ...group of committed researchers, engineers, policy experts, and business...  ...About the role Our Inference team is responsible for building...  ...models. We tackle complex, distributed systems challenges across...  ...traffic management systems LLM inference optimization, batching... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...focus on high-performance model inference and accelerating research...  ...systems. In this role, you’ll lead engineering efforts to ensure our largest...  ..., CUDA development, and distributed inference best practices....  ...AMD GPUs, ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron, MPI, or... 
    Full time

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...improve their business. Founded by engineers — and customer obsessed — we leap...  ...Software Engineer or SRE in highly distributed, multi-cloud environments....  ...infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks... 
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  • $166k - $225k

     ...use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve...  ...team at Databricks, you will be building the next generation distributed data storage and processing systems that can outperform specialized... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    22 hours ago
  •  ...This team keeps our browsers running at scale, solving massive distributed systems challenges and making sure our platform is fast,...  ...with developer-friendly APIs. Work closely with the rest of Engineering, gathering input and providing great support so every team can... 
    Full time
    Immediate start
    Relocation

    Browserbase

    San Francisco, CA
    1 day ago
  • $300k

     ...group of committed researchers, engineers, policy experts, and business...  ...the Role The Cloud Inference team scales and optimizes Claude...  ...-performance, large-scale distributed systems serving millions of users...  ...Strong familiarity with LLM inference optimization, batching... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    1 day ago
  • $180k - $275k

     ...technical investments with rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate across frontend,...  ...What you'll do Design and implement scalable APIs, distributed systems, and data infrastructure that serve millions of users... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    1 day ago
  • $189.6k - $237k

     ...RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been...  ...automatic training and evaluation of LLM's, as well as evaluation of data...  ...ML systemsStrong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    5 days ago
  •  ...Mach9, Sensor Data Integration Engineers build the algorithms and...  ...foundation in parallel computing or distributed systems A bachelor's...  ...agent harnesses — orchestrating LLM-driven workflows for triage,...  ...pipelines that feed ML training and inference. Familiar with C++.... 
    Full time
    Work at office
    Work from home

    Mach9

    San Francisco, CA
    1 day ago
  •  ...Distributed Systems Engineer As a distributed systems engineer, you'll work across the stack to solve problems as they come up and help build Archil volumes. You'll have significant influence over the technical and product direction. We'll expect you to be able... 
    Flexible hours

    Archil

    San Francisco, CA
    2 days ago
  •  ...more time putting knowledge into action. We're looking for engineers who want to build the operating system for AI Data Applications...  .... About the role We're looking for experienced distributed systems engineers to build the core infrastructure for our durable... 

    Tensorlake, Inc.

    San Francisco, CA
    1 day ago
  • $300 per month

     ...RoleAt Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and cloud services to...  ...focus on serving and scaling LLM workloadsDefine, measure, and...  ...optimize large-scale training and inference clustersAutomate... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $172.5k - $260.1k

     ...Description: Salesforce has immediate opportunities for Lead software engineers who want their lines of code to have significant and...  ...Envision and Build new and exciting components/frameworks in distributed filesystems in an ever-growing and evolving market technology... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    4 days ago
  • $300 per month

     ...Role:At Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and hands-on experience...  ...focus on serving and scaling LLM workloadsBuild automation and...  ...distributed AI pipelines and inference servicesDefine, measure, and... 
    Temporary work

    Crusoe

    San Francisco, CA
    22 hours ago
  •  ...combined with seamless physical skills under the practical constraints of robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    1 day ago
  •  ...Reflection in San Francisco is looking for a data engineer dedicated to enhancing data quality for LLM pre-training. Your role will involve collaborating with world-class researchers to establish high standards for data collection and processing. You will design automated... 

    Reflection

    San Francisco, CA
    1 day ago
  • $120k - $150k

     ...A tech startup is seeking a Software Engineer New Grad to work in San Francisco. You will contribute to core products and an open-source engine, gaining hands-on experience in building distributed data systems. Ideal candidates are recent graduates with strong programming... 

    Eventual

    San Francisco, CA
    1 day ago
  • $59.15k - $106.93k

     ...to thrive, professionally and personally. For us, helping you grow your career is good business. Leidos is seeking Distribution Engineers in Maui, HI and Oahu, HI who are passionate about electric utility design engineering. We're looking for someone dedicated... 
    Local area
    Immediate start
    Relocation
    Relocation package
    Flexible hours
    Night shift

    Leidos

    San Francisco, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!