Distributed LLM Inference Engineer
Anyscale
Distributed LLM Inference Engineer
At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world.
With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.
About the role
As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure.
As part of this role, you will
- Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale
- Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference
- Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source
- Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices
We'd love to hear from you if you have
- Familiarity with running ML inference at large scale with high throughput and low latency
- Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)
- Solid understanding of distributed systems, ML inference challenges
Bonus points!
- ML Systems knowledge
- Experience using Ray
- Work closely with community on LLM engines like vLLM, TensorRT-LLM
- Contributions to deep learning frameworks (PyTorch, TensorFlow)
- Contributions to deep learning compilers (Triton, TVM, MLIR)
- Prior experience working on GPUs / CUDA
Compensation
At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted.
This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following:
- Stock Options
- Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents
- 401k Retirement Plan
- Education & Wellbeing Stipend
- Paid Parental Leave
- Fertility Benefits
- Paid Time Off
- Commute reimbursement
- 100% of in-office meals covered
Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law.
Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish
- GMI Cloud, a fast-growing AI infrastructure company, is hiring a Machine Learning Engineer, LLM Optimization to build a world-leading inference optimization team. You will drive research, validation, and productionization of advanced optimization techniques to boost latency...Suggested
- ...leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter... ...fabric. We are an applied research and engineering team that moves fast, ships real systems... ...— from kernel-level optimization to distributed orchestration to high-level serving APIs...Suggested
$224k - $356.5k
NVIDIA is seeking an Engineering Manager to lead the development of an... ...visibility into model behavior, inference performance, reliability, and... ...signals across large scale LLM and VLM deployments. It will... ...serving, model optimization, distributed systems, GPU performance, and...SuggestedFull time$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...architecture, parallel programming, distributed systems, deep learning theories.Knowledgeable... ...building and optimizing LLM inference engines (e.g., vLLM, SGLang...SuggestedFull time$174k - $252k
...and implement efficient, next-generation distributed caching and database architectures.... ...Experience developing accessible technologies or engineering highly reliable, planet-scale systems... ...Generative AI applications or LLM-based orchestration frameworks.Google's...Suggested$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time multimodalconversational... ...models that fall outside standard LLM serving patterns — sustained low-latencyoutput... ...-scale evaluation- Experience with distributed training and post-training pipelines (...InternshipLocal areaFlexible hours- ...ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis... ...(AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed... ...inference networking: Profile and optimize distributed inference topologies including...
$160.36k - $240.54k
...connected future. About the Role We’re looking for senior engineers to build/scale Nuro's large-scale computing infrastructure in... ...have proven experience in building and developing large-scale distributed applications (e.g. Kubernetes). You’re self-motivated to...Full time- ...of enabling human life on Mars.SOFTWARE ENGINEER, INFERENCE (AI DATA ENGINEERING)The application... ...systems end-to-end, owning everything from distributed infrastructure to deep low-level... ...engines (e.g., SGLang, vLLM, TensorRT-LLM) Develop custom tools for tracing, replaying...Permanent employmentTemporary workRemote workWorldwideWeekend work
$158k - $237k
...whether in the data center, at the edge, or in the cloud. It is a distributed, scale-out, fault tolerant, performant, deduplicated user-... ...high bar for quality and velocity. We are looking for motivated engineers, who can complement this rockstar team!About RoleRubrik is...Local areaShift work$90 - $121.86 per hour
...Job Description Job Description LLM Research Engineer Key Responsibilities: Design, train... ...latency, memory footprint, and inference time for real-time applications. Collaborate... .... ~ Strong understanding of distributed computing and GPU acceleration using...Hourly pay- ...automation with Moveworks’ Reasoning Engine and natural language capabilities, we... ...infrastructure for building and serving LLM’s at Moveworks. This role will be... ...variety of responsibilities including distributed training and inference pipeline for large language models(LLM...Work at officeRemote workFlexible hours
$158k - $237k
...About the team The Forge Team: Engineering the Backbone of Rubrik's Platform The Forge team is at the core of Rubrik's mission... ...~ Strong fundamentals in data structures, algorithms, and distributed systems design ~ Strong background in Systems Programming...Full timeLocal area$152k - $241.5k
...Infrastructure team supports 1000+ chip design engineers by building tools and platforms that... ...a generalist role with an emphasis on distributed systems and operational excellence in a... ...unfamiliar codebases in any language (including LLM-generated code) to implement high-...Full time$135.42k - $226.98k
As an EDS Integration Engineer for the Advanced EV Team, you will be at the forefront of defining the "nervous system" for our next-generation... ...EV Expertise: Experience with High Voltage (HV) DC power distribution and the challenges of integrating diverse voltage...Immediate startVisa sponsorshipFlexible hours$207k - $300k
Analyze and optimize AI inference workloads across the application, model, and distributed fleet infrastructure layers to methodically... ...in Computer Science, Computer Engineering, Electrical Engineering,... ...qualifications:Experience with real world LLM inference serving environments...- ...a Senior Principal Software Engineer at JPMorganChase within the Commercial... ...and optimization using model inference servers such as Triton... ...architecting and deploying LLM & GNN solutions on AWS (e.g.,... ...in inference optimization and distributed systems for large models focused...
$198k - $326k
...power AI across LinkedIn. The LLM Serving team builds the... ...for a Senior Staff Software Engineer with deep expertise at the intersection... ..., and large-scale inference. This is a highly technical,... ...experience in software engineering, distributed systems, infrastructure, or machine...For contractorsWork at officeFlexible hours$142.8k - $274.8k
...Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft... ...-leading AI training and inference. The Platform Systems... ...compiler-generated kernels, distributed communication fabrics, and large... ...including: HPL/HPC benchmarks LLM training workloads Transformer...Ongoing contractWork at officeLocal areaWorldwide3 days per week- ...the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training... .... The Mission We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-...Visa sponsorship
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...Full time$192k - $260k
A leading data and AI infrastructure company in Mountain View, California, is seeking a Staff Software Engineer to work on distributed data systems. The ideal candidate will have 8+ years of experience in Java, Scala, or C++ and a strong foundation in algorithms and distributed...$192k - $260k
Staff Software Engineer - Distributed Data Systems P-186 Get AI-powered advice on this job and more exclusive features. At Databricks, we are obsessed with enabling data teams to solve the world’s toughest problems, from security threat detection to cancer drug development...Work at officeLocal area$160k - $275k
What MatX Is BuildingMatX is building next-generation AI compute infrastructure for large-scale LLM training and inference. We are looking for a hands-on Mechanical Engineer to design, develop, and validate mechanical systems for our rack-scale AI platform from compute...Daily paidFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours$250k - $350k
...the world's leading ML systems engineers, including leaders behind... ...powering our large-scale training, inference, and reinforcement learning.... ...and serving systemsDesign distributed runtimes and scheduling systems... ...with vLLM, TensorRT-LLM, or production LLM serving systems...Visa sponsorship$224k - $356.5k
...growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud... ...Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo organization... ...infrastructure, ML tooling, or distributed systems.3+ years of engineering...Full timeLocal areaRemote workWorldwide$224k - $356.5k
...exceptional Manager, Deep Learning Inference Software, to lead a world-class engineering team advancing the state of AI... ...optimization of large-scale models for LLM, multimodal, and generative AI... ...(NIXL, NCCL, NVSHMEM) and distributed inference architectures.Expertise...Full timeRemote workWorldwide$184k - $287.5k
...are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who... ...:Implement language and multimodal model inference as part of NVIDIA Inference Microservices... ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference serving library...Full time$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training... .... This role sits at the intersection of distributed training, GPU architecture, systems... ...will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs,...Full time$174k - $252k
...components critical for Large Language Model (LLM) inference serving on GDC. This includes areas... ..., and efficient model sharding across distributed hardware.Collaborate closely with... ...orchestration (Kubernetes, Google Kubernetes Engine (GKE)), networking infrastructure, and...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!




