Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Distributed LLM Inference Engineer

Anyscale

Distributed LLM Inference Engineer

At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.

With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.

Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.

About the role

As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.

As part of this role, you will

  • Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale
  • Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference
  • Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source
  • Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices

We'd love to hear from you if you have

  • Familiarity with running ML inference at large scale with high throughput and low latency
  • Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)
  • Solid understanding of distributed systems, ML inference challenges

Bonus points!

  • ML Systems knowledge
  • Experience using Ray
  • Work closely with community on LLM engines like vLLM, TensorRT-LLM
  • Contributions to deep learning frameworks (PyTorch, TensorFlow)
  • Contributions to deep learning compilers (Triton, TVM, MLIR)
  • Prior experience working on GPUs / CUDA

Compensation

At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.

This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:

  • Stock Options
  • Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents
  • 401k Retirement Plan
  • Education & Wellbeing Stipend
  • Paid Parental Leave
  • Fertility Benefits
  • Paid Time Off
  • Commute reimbursement
  • 100% of in-office meals covered

Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.

Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish

Vacancy posted 5 days ago
Similar jobs that could be interesting for youBased on the Distributed LLM Inference Engineer in Palo Alto, CA vacancy
  • GMI Cloud, a fast-growing AI infrastructure company, is hiring a Machine Learning Engineer, LLM Optimization to build a world-leading inference optimization team. You will drive research, validation, and productionization of advanced optimization techniques to boost latency... 
    Suggested

    GMI Cloud

    Mountain View, CA
    5 days ago
  •  ...leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter...  ...fabric. We are an applied research and engineering team that moves fast, ships real systems...  ...— from kernel-level optimization to distributed orchestration to high-level serving APIs... 
    Suggested

    d-Matrix

    Santa Clara, CA
    4 days ago
  • $224k - $356.5k

    NVIDIA is seeking an Engineering Manager to lead the development of an...  ...visibility into model behavior, inference performance, reliability, and...  ...signals across large scale LLM and VLM deployments. It will...  ...serving, model optimization, distributed systems, GPU performance, and... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale...  ...architecture, parallel programming, distributed systems, deep learning theories.Knowledgeable...  ...building and optimizing LLM inference engines (e.g., vLLM, SGLang... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $174k - $252k

     ...and implement efficient, next-generation distributed caching and database architectures....  ...Experience developing accessible technologies or engineering highly reliable, planet-scale systems...  ...Generative AI applications or LLM-based orchestration frameworks.Google's... 
    Suggested

    Google

    Mountain View, CA
    4 days ago
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational...  ...models that fall outside standard LLM serving patterns — sustained low-latencyoutput...  ...-scale evaluation- Experience with distributed training and post-training pipelines (... 
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    2 days ago
  •  ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis...  ...(AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed...  ...inference networking: Profile and optimize distributed inference topologies including... 

    AMD

    Santa Clara, CA
    1 day ago
  • $160.36k - $240.54k

     ...connected future. About the Role We’re looking for senior engineers to build/scale Nuro's large-scale computing infrastructure in...  ...have proven experience in building and developing large-scale distributed applications (e.g. Kubernetes). You’re self-motivated to... 
    Full time

    Nuro

    Mountain View, CA
    1 day ago
  •  ...of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application...  ...systems end-to-end, owning everything from distributed infrastructure to deep low-level...  ...engines (e.g., SGLang, vLLM, TensorRT-LLM) Develop custom tools for tracing, replaying... 
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Palo Alto, CA
    4 days ago
  • $158k - $237k

     ...whether in the data center, at the edge, or in the cloud. It is a distributed, scale-out, fault tolerant, performant, deduplicated user-...  ...high bar for quality and velocity. We are looking for motivated engineers, who can complement this rockstar team!About RoleRubrik is... 
    Local area
    Shift work

    Rubrik

    Palo Alto, CA
    4 days ago
  • $90 - $121.86 per hour

     ...Job Description Job Description LLM Research Engineer Key Responsibilities: Design, train...  ...latency, memory footprint, and inference time for real-time applications. Collaborate...  .... ~ Strong understanding of distributed computing and GPU acceleration using... 
    Hourly pay

    Cypress HCM

    Mountain View, CA
    8 days ago
  •  ...automation with Moveworks’ Reasoning Engine and natural language capabilities, we...  ...infrastructure for building and serving LLM’s at Moveworks. This role will be...  ...variety of responsibilities including distributed training and inference pipeline for large language models(LLM... 
    Work at office
    Remote work
    Flexible hours

    Moveworks

    Mountain View, CA
    1 day ago
  • $158k - $237k

     ...About the team The Forge Team: Engineering the Backbone of Rubrik's Platform The Forge team is at the core of Rubrik's mission...  ...~ Strong fundamentals in data structures, algorithms, and distributed systems design ~ Strong background in Systems Programming... 
    Full time
    Local area

    Rubrik

    Palo Alto, CA
    1 day ago
  • $152k - $241.5k

     ...Infrastructure team supports 1000+ chip design engineers by building tools and platforms that...  ...a generalist role with an emphasis on distributed systems and operational excellence in a...  ...unfamiliar codebases in any language (including LLM-generated code) to implement high-... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $135.42k - $226.98k

    As an EDS Integration Engineer for the Advanced EV Team, you will be at the forefront of defining the "nervous system" for our next-generation...  ...EV Expertise: Experience with High Voltage (HV) DC power distribution and the challenges of integrating diverse voltage... 
    Immediate start
    Visa sponsorship
    Flexible hours

    Ford

    Palo Alto, CA
    1 day ago
  • $207k - $300k

    Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically...  ...in Computer Science, Computer Engineering, Electrical Engineering,...  ...qualifications:Experience with real world LLM inference serving environments... 

    Google

    Mountain View, CA
    3 days ago
  •  ...a Senior Principal Software Engineer at JPMorganChase within the Commercial...  ...and optimization using model inference servers such as Triton...  ...architecting and deploying LLM & GNN solutions on AWS (e.g.,...  ...in inference optimization and distributed systems for large models focused... 

    JP Morgan Chase

    Palo Alto, CA
    1 day ago
  • $198k - $326k

     ...power AI across LinkedIn. The LLM Serving team builds the...  ...for a Senior Staff Software Engineer with deep expertise at the intersection...  ..., and large-scale inference. This is a highly technical,...  ...experience in software engineering, distributed systems, infrastructure, or machine... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    1 day ago
  • $142.8k - $274.8k

     ...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft...  ...-leading AI training and inference. The Platform Systems...  ...compiler-generated kernels, distributed communication fabrics, and large...  ...including: HPL/HPC benchmarks LLM training workloads Transformer... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Mountain View, CA
    3 days ago
  •  ...the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training...  ....   The Mission We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-... 
    Visa sponsorship

    Institute of Foundation Models

    Sunnyvale, CA
    6 days ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $192k - $260k

    A leading data and AI infrastructure company in Mountain View, California, is seeking a Staff Software Engineer to work on distributed data systems. The ideal candidate will have 8+ years of experience in Java, Scala, or C++ and a strong foundation in algorithms and distributed... 

    Databricks

    Mountain View, CA
    1 day ago
  • $192k - $260k

    Staff Software Engineer - Distributed Data Systems P-186 Get AI-powered advice on this job and more exclusive features. At Databricks, we are obsessed with enabling data teams to solve the world’s toughest problems, from security threat detection to cancer drug development... 
    Work at office
    Local area

    Databricks

    Mountain View, CA
    1 day ago
  • $160k - $275k

    What MatX Is BuildingMatX is building next-generation AI compute infrastructure for large-scale LLM training and inference. We are looking for a hands-on Mechanical Engineer to design, develop, and validate mechanical systems for our rack-scale AI platform from compute... 
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    3 days ago
  • $250k - $350k

     ...the world's leading ML systems engineers, including leaders behind...  ...powering our large-scale training, inference, and reinforcement learning....  ...and serving systemsDesign distributed runtimes and scheduling systems...  ...with vLLM, TensorRT-LLM, or production LLM serving systems... 
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    1 day ago
  • $224k - $356.5k

     ...growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud...  ...Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo organization...  ...infrastructure, ML tooling, or distributed systems.3+ years of engineering... 
    Full time
    Local area
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    3 days ago
  • $224k - $356.5k

     ...exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI...  ...optimization of large-scale models for LLM, multimodal, and generative AI...  ...(NIXL, NCCL, NVSHMEM) and distributed inference architectures.Expertise... 
    Full time
    Remote work
    Worldwide

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who...  ...:Implement language and multimodal model inference as part of NVIDIA Inference Microservices...  ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference serving library... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training...  .... This role sits at the intersection of distributed training, GPU architecture, systems...  ...will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs,... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $174k - $252k

     ...components critical for Large Language Model (LLM) inference serving on GDC. This includes areas...  ..., and efficient model sharding across distributed hardware.Collaborate closely with...  ...orchestration (Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure, and... 

    Google

    Sunnyvale, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!