Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference Engineer

$170k - $245k

Anyscale

About AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.As part of this role, you willIterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of AnyscaleWork across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open sourceFollow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you haveFamiliarity with running ML inference at large scale with high throughput and low latencyFamiliarity with deep learning and deep learning frameworks (e.g. PyTorch)Solid understanding of distributed systems, ML inference challengesBonus points!ML Systems knowledgeExperience using Ray Work closely with community on LLM engines like vLLM, TensorRT-LLMContributions to deep learning frameworks (PyTorch, TensorFlow)Contributions to deep learning compilers (Triton, TVM, MLIR)Prior experience working on GPUs / CUDACompensationAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted. This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents401k Retirement PlanEducation & Wellbeing StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in-office meals coveredAnyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and SpanishCompensation Range: $170K - $245KLocationSan Francisco; Palo AltoEmployment TypeFull timeLocation TypeHybridDepartmentEngineeringCompensationTarget Base Salary:$170K – $245K • Offers EquityAt Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:Stock OptionsHealthcare plans, with premiums covered by Anyscale at 99%401k Retirement PlanWellness & Education StipendPaid Parental LeaveFertility BenefitsPaid Time OffCommute reimbursement100% of in office meals covered

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference Engineer in San Francisco, CA vacancy
  • $227.2k - $324.5k

     ...the Role:As a Staff Software Engineer on the ML Infrastructure team...  ...world-class machine learning inference platforms. These platforms power...  ...that support Deep Learning, LLM, and Search models. This...  ...throughput, and low latency distributed systems using ScalaBuild reusable... 
    Suggested
    Full time
    Temporary work
    Local area
    Flexible hours

    Tubi TV

    San Francisco, CA
    1 day ago
  •  ...matters to the world. The Production Engineering Team Examples of key exciting problems...  ..., actual state inspection, and distributed command execution. One interface for the...  ...move on. You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks... 
    Suggested
    Local area

    Fluidstack

    San Francisco, CA
    5 days ago
  • $230k

     ..., governance, and trust become the real engineering challenge. Arcade is the MCP runtime...  ...vector database team at Redis, shipped 100+ LLM applications, and is a contributor to...  ...assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis... 
    Suggested
    Work at office
    Shift work

    Arcade AI, Inc

    San Francisco, CA
    4 days ago
  •  ...Baseten powers mission-critical inference for the world's most dynamic AI...  ...us and help build the platform engineers turn to to ship AI products....  ...the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  • $117.2k - $223.9k

     ...unparalleled customer experiences. Join our team of talented engineers and help us advance the integration of Salesforce applications...  ...vulnerability management, infrastructure security, and the security of distributed and scalable distributed systems. This role requires hands-on... 
    Suggested
    Full time

    Salesforce

    San Francisco, CA
    1 day ago
  • $175k - $250k

     ...vision, dental) Job Details We’re looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across...  ...ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing... 
    Local area

    Jobot

    San Francisco, CA
    1 day ago
  •  ...the infrastructure. Help design, implement, and monitor testnets Required Skills: Expert knowledge of peer-to-peer distributed system design and implementation (required) Ability to build and maintain high available infrastructure (required) Knowledge... 

    1872 Consulting

    San Francisco, CA
    5 days ago
  •  ...more time putting knowledge into action. We're looking for engineers who want to build the operating system for AI Data Applications...  .... About the role We're looking for experienced distributed systems engineers to build the core infrastructure for our durable... 

    Tensorlake, Inc.

    San Francisco, CA
    2 days ago
  • $153k - $376k

     ...infrastructure is at the heart of everything we build. As a Software Engineer on our Infrastructure team, you’ll help design, build, and...  .... We’re scaling fast, and we’re looking for experienced distributed systems engineers across a variety of teams. Whether you’re passionate... 
    Minimum wage
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    Figma

    San Francisco, CA
    1 day ago
  • $166k - $225k

     ...use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve...  ...team at Databricks, you will be building the next generation distributed data storage and processing systems that can outperform specialized... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...Distributed Systems Engineer @ Dedalus Labs Mission Dedalus Labs is an AI research neolab building infrastructure for AI agents. We’re building the persistent compute layer that powers the next generation of autonomous software. Our platform spans distributed... 
    Work at office
    Visa sponsorship
    Relocation package

    Dedalus Labs

    San Francisco, CA
    1 day ago
  • $300 per month

     ...RoleAt Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and cloud services to...  ...focus on serving and scaling LLM workloadsDefine, measure, and...  ...optimize large-scale training and inference clustersAutomate... 
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • $148.5k - $223.9k

     ...include: Build new and exciting components/frameworks in developing distributed filesystems in an ever-growing and evolving market technology...  ...filesystem environment, Code review, mentoring junior engineers, and providing technical guidance to the teamRequired Skills:... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    3 days ago
  • $172.5k - $260.1k

     ...services that self-heal and self-optimize, we have to balance this with global scale and traffic distribution to enhance our end user experience.Edge is hiring a backend Java engineer with distributed systems, micro-services, and public cloud experience to support this... 
    Full time

    Salesforce

    San Francisco, CA
    1 day ago
  • $189.6k - $237k

     ...RLXF) team builds our internal distributed framework for large language model training and inference. The platform has been...  ...automatic training and evaluation of LLM's, as well as evaluation of data...  ...ML systemsStrong software engineering skills, proficient in frameworks... 
    Full time

    Scale AI

    San Francisco, CA
    1 day ago
  • B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a... 

    B Capital

    San Francisco, CA
    2 days ago
  • $230k - $385k

     ...benefits all of humanity, our team combines world-class researchers, engineers, designers, and operators who care deeply about creating...  ...Role As an Operating Systems Engineer focused on on-device inference, you will design, develop, and ship the OS stack that makes... 
    Full time

    OpenAI

    San Francisco, CA
    3 days ago
  • $172.5k - $260.1k

     ...Description: Salesforce has immediate opportunities for Lead software engineers who want their lines of code to have significant and...  ...Envision and Build new and exciting components/frameworks in distributed filesystems in an ever-growing and evolving market technology... 
    Full time
    Immediate start

    Salesforce

    San Francisco, CA
    14 hours ago
  • $300 per month

     ...Role:At Crusoe, our Production Engineering team ensures the reliability...  ...with a strong background in distributed systems and hands-on experience...  ...focus on serving and scaling LLM workloadsBuild automation and...  ...distributed AI pipelines and inference servicesDefine, measure, and... 
    Temporary work

    Crusoe

    San Francisco, CA
    2 days ago
  • $180k - $220k

     ...success and ability to scale. This role reports to the Senior Engineering Manager of Realtime Infrastructure. What You'll Be Doing...  ...Build and operate large-scale, reliable and performant distributed systems. Collaborate with product teams to create new features... 
    Full time
    Relocation
    Relocation package

    Discord

    San Francisco, CA
    9 days ago
  • Electrical Engineer - Power Distribution and AnalysisHED is looking to add an experienced team member to our group of talented Electrical Engineers. The candidate should have broad electrical engineer expertise, with a focus on power distribution system and power system... 
    Work at office
    Work from home
    Flexible hours

    HED

    San Francisco, CA
    2 days ago
  • $229.9k - $262.4k

     ...Senior Lead Software Engineer, Distributed Systems (Golang + Python on Kubernetes) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive, and iterative delivery environment... 
    Full time
    Part time
    Internship
    Local area

    Capital One

    San Francisco, CA
    4 days ago
  • $350k

     ...AI Systems Engineer Salary range: $350,000 - $600,000/year + benefits...  ...Interpretability: Inference stacks that are as performant...  ...Behavior elicitation: Distributed RL training and roll-outs allowing...  ...scale) Bonus: can set up LLM pipelines, e.g. multiple specialized... 
    Visa sponsorship
    Flexible hours

    Transluce

    San Francisco, CA
    5 days ago
  • $206.4k - $379.1k

     ...Services team is seeking a Principal Service Engineer to serve as the technical lead for our...  ...flagship products.Design and architect inference infrastructure for enterprise-scale...  ....Hands-on expertise with Kubernetes, distributed systems, and MLOps platforms.Preferred... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    1 day ago
  • $176k - $220k

     ...institutions Work together with engineers, scientists, operators, and...  ...Handshake is hiring a Senior LLM Platform Engineer to join our...  ..., and hosted or self-hosted inference. You’ll also contribute to...  ...Experience building observability for distributed systems and leading... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Handshake

    San Francisco, CA
    5 days ago
  • $190k - $265k

     ...improve their business. Founded by engineers — and customer-obsessed — we...  ...of AI.The Foundation Model Inference team is the backbone of...  ...The impact you will have:Build LLM infrastructure powering large...  ..., latency, and efficiency of distributed AI workloadsCollaborate with... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $180k - $310k

     ...technical investments with rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate across frontend,...  ...matters most. What You'll Do Design and implement scalable APIs, distributed systems, and data infrastructure that serve millions of users... 
    Full time
    Work at office
    Work from home

    Gamma

    San Francisco, CA
    2 days ago
  • $150k - $240k

     ...we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud... 
    Remote work
    Flexible hours

    Verse

    San Francisco, CA
    21 days ago
  • $180k - $250k

     ...unified platform where high-performance inference, orchestration, and observability come...  ...role:  You are an experienced software engineer who thrives on building large-scale...  ...You have deep expertise in large scale distributed systems that deal with high complexity,... 
    Full time
    Currently hiring
    Remote work
    Relocation package

    Falò

    San Francisco, CA
    1 day ago
  • $172.5k - $285.8k

     ...ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI...  ..., we have to balance this with global scale and traffic distribution to enhance our end user experience. Edge is hiring a backend... 
    Full time

    Salesforce

    San Francisco, CA
    8 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!