Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference Engineer

$170k - $245k

Anyscale

About AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.As part of this role, you willIterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of AnyscaleWork across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open sourceFollow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you haveFamiliarity with running ML inference at large scale with high throughput and low latencyFamiliarity with deep learning and deep learning frameworks (e.g. PyTorch)Solid understanding of distributed systems, ML inference challengesBonus points!ML Systems knowledgeExperience using Ray Work closely with community on LLM engines like vLLM, TensorRT-LLMContributions to deep learning frameworks (PyTorch, TensorFlow)Contributions to deep learning compilers (Triton, TVM, MLIR)Prior experience working on GPUs / CUDACompensationAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted. This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents401k Retirement PlanEducation & Wellbeing StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in-office meals coveredAnyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and SpanishCompensation Range: $170K - $245KLocationSan Francisco; Palo AltoEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensationTarget Base Salary:$170K – $245K • Offers EquityAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99%401k Retirement PlanWellness & Education StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in office meals covered

Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference Engineer in San Francisco, CA vacancy
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI...  ...us and help build the platform engineers turn to to ship AI products....  ...the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  • $227.2k - $324.5k

     ...the Role:As a Staff Software Engineer on the ML Infrastructure team...  ...world-class machine learning inference platforms. These platforms power...  ...that support Deep Learning, LLM, and Search models. This...  ...throughput, and low latency distributed systems using ScalaBuild reusable... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    2 days ago
  • $180k - $250k

     ...unified platform where high-performance inference, orchestration, and observability come...  ...role:  You are an experienced software engineer who thrives on building large-scale...  ...You have deep expertise in large scale distributed systems that deal with high complexity,... 
    Suggested
    Full time
    Currently hiring
    Remote work
    Relocation package

    Falò

    San Francisco, CA
    a month ago
  • $200k - $260k

     ...vector database team at Redis, shipped 100+ LLM applications, and is a contributor to...  ...assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis...  ...desire to ship. ~7+ years of software engineering experience comprising of: ~5+ years... 
    Suggested
    Work at office
    Shift work

    Arcade

    San Francisco, CA
    20 hours ago
  • $180k - $310k

     ...makes Gamma magical. This means designing distributed systems for real-time content scanning,...  ...velocity. You'll collaborate across engineering, product, and design to define how Gamma...  ...suspicious or malicious activity Leverage AI/LLM-based detection to stay ahead of AI-... 
    Suggested
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    a month ago
  • $146.5k

     ...preferences. About the team: The ML Data Engineering team powers metadata extraction,...  ...machine learning, data engineering, and distributed systems, collaborating closely with...  ...product teams to deploy scalable ML and LLM-powered solutions in production. Role... 
    Full time
    Local area
    Worldwide
    Home office
    Flexible hours

    Scribd, Inc.

    San Francisco, CA
    more than 2 months ago
  • $117.2k - $223.9k

     ...unparalleled customer experiences. Join our team of talented engineers and help us advance the integration of Salesforce applications...  ...vulnerability management, infrastructure security, and the security of distributed and scalable distributed systems. This role requires hands-on... 
    Full time

    Salesforce

    San Francisco, CA
    2 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI companies...  ...us and help build the platform engineers turn to to ship AI products....  ...Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our platform. We... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    more than 2 months ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our...  ...a small, fast-moving team of engineers focused on delivering a world...  ...systems that span networking, distributed compute, and high-throughput...  ...tooling like vLLM, TensorRT-LLM, or custom model parallel systems... 
    Full time

    OpenAI

    San Francisco, CA
    more than 2 months ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and business...  ...the role The Cloud Inference team scales and optimizes Claude...  ...-performance, large-scale distributed systems serving millions of users...  ...Are curious about LLM serving; prior inference or ML... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    more than 2 months ago
  •  ...This team keeps our browsers running at scale, solving massive distributed systems challenges and making sure our platform is fast,...  ...with developer-friendly APIs. Work closely with the rest of Engineering, gathering input and providing great support so every team can... 
    Full time
    Immediate start
    Relocation

    Browserbase

    San Francisco, CA
    more than 2 months ago
  • $224.5k - $251.5k

     ...datasets while fostering an AI-native engineering culture. This position reports directly...  ...implement emerging agent frameworks, LLM inference optimization, advanced retrieval...  ...Background: Strong foundations in scaling distributed systems and production-grade... 
    Full time
    Work at office

    Dialpad

    San Francisco, CA
    a month ago
  • $180k - $275k

     ...technical investments with rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate across frontend,...  ...What you'll do Design and implement scalable APIs, distributed systems, and data infrastructure that serve millions of users... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    more than 2 months ago
  •  ...About the Role Join a startup building an agentic data lakehouse platform. As a Senior Software Engineer, Distributed Data Systems, you'll work on a greenfield project to build scalable data infrastructure that transforms enterprise data into actionable insights at scale... 
    Full time

    Clera

    San Francisco, CA
    more than 2 months ago
  • $160k - $194k

     ...we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud... 
    Full time
    Remote work
    Flexible hours

    Ad Verse

    San Francisco, CA
    a month ago
  • $161.3k - $241.9k

     ...customer interaction, every model inference, and every production...  ...’re looking for a Production Engineer to help build and operate Harvey...  ...Experience building and operating distributed systems with strong...  ...Experience supporting AI/ML or LLM infrastructure at scale. Experience... 
    Full time

    Harvey, Inc.

    San Francisco, CA
    a month ago
  •  ...Distributed Systems Engineer @ Dedalus Labs Mission Dedalus Labs is an AI research neolab building infrastructure for AI agents. We're building the persistent compute layer that powers the next generation of autonomous software. Our platform spans distributed... 
    Work at office
    Visa sponsorship
    Relocation package

    Dedalus Labs

    San Francisco, CA
    3 days ago
  •  ...Distributed Systems EngineerAs a distributed systems engineer, you'll work across the stack to solve problems as they come up and help build Archil volumes. You'll have significant influence over the technical and product direction.We'll expect you to be able to:Be oncall... 
    Flexible hours

    Archil

    San Francisco, CA
    3 days ago
  •  ...combined with seamless physical skills under the practical constraints of robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale multimodal training and evaluation at... 
    Full time
    Work at office
    Relocation package

    OpenAI

    San Francisco, CA
    more than 2 months ago
  • $153k - $376k

     ...infrastructure is at the heart of everything we build. As a Software Engineer on our Infrastructure team, you’ll help design, build, and...  .... We’re scaling fast, and we’re looking for experienced distributed systems engineers across a variety of teams. Whether you’re passionate... 
    Minimum wage
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    Figma

    San Francisco, CA
    2 days ago
  • $166k - $225k

     ...use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve...  ...team at Databricks, you will be building the next generation distributed data storage and processing systems that can outperform specialized... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  • $230k - $288k

     ...Austin, NYC, San Francisco About the Dept  Cloudflare’s engineers build and operate the software that helps power 25+ million...  ...own the technical coherence of a platform that spans globally distributed key-value storage, progressive release and configuration delivery... 
    Full time
    Temporary work
    Local area
    Flexible hours

    Cloudflare

    San Francisco, CA
    more than 2 months ago
  •  ...Senior Systems Engineer San Francisco, California Onsite or Remote...  ..., database performance, AI inference infrastructure: you can cover...  ...disaster recovery. Own AI and LLM inference infrastructure end-...  ...for instrumentation and distributed tracing across complex environments... 
    Remote work
    Work from home

    Evidently

    San Francisco, CA
    1 day ago
  • $300 per month

     ...RoleAt Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and cloud services to...  ...focus on serving and scaling LLM workloadsDefine, measure, and...  ...optimize large-scale training and inference clustersAutomate... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $189.6k - $237k

     ...RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been...  ...automatic training and evaluation of LLM's, as well as evaluation of data...  ...ML systemsStrong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    2 days ago
  • $172.5k - $260.1k

     ...services that self-heal and self-optimize, we have to balance this with global scale and traffic distribution to enhance our end user experience.Edge is hiring a backend Java engineer with distributed systems, micro-services, and public cloud experience to support this... 
    Full time

    Salesforce

    San Francisco, CA
    2 days ago
  • $148.5k - $223.9k

     ...include: Build new and exciting components/frameworks in developing distributed filesystems in an ever-growing and evolving market technology...  ...filesystem environment, Code review, mentoring junior engineers, and providing technical guidance to the teamRequired Skills:... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    4 days ago
  • B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a... 

    B Capital

    San Francisco, CA
    3 days ago
  • $172.5k - $260.1k

     ...Description: Salesforce has immediate opportunities for Lead software engineers who want their lines of code to have significant and...  ...Envision and Build new and exciting components/frameworks in distributed filesystems in an ever-growing and evolving market technology... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    1 day ago
  • $300 per month

     ...Role:At Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and hands-on experience...  ...focus on serving and scaling LLM workloadsBuild automation and...  ...distributed AI pipelines and inference servicesDefine, measure, and... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!