Distributed LLM Inference Engineer
Anyscale
Distributed LLM Inference Engineer At Anyscale, we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We're commercializing Ray, a popular open-source project that's creating an ecosystem of libraries for scalable machine learning. Companies like OpenAI, Uber, Spotify, Instacart, Cruise, and many more, have Ray in their tech stacks to accelerate the progress of AI applications out into the real world. With Anyscale, we're building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert. Proud to be backed by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date. About the role As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale. This is an incredibly critical role to Anyscale as it allows us to achieve a market leading position for AI infrastructure. As part of this role, you will Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices We'd love to hear from you if you have Familiarity with running ML inference at large scale with high throughput and low latency Familiarity with deep learning and deep learning frameworks (e.g. PyTorch) Solid understanding of distributed systems, ML inference challenges Bonus points! ML Systems knowledge Experience using Ray Work closely with community on LLM engines like vLLM, TensorRT-LLM Contributions to deep learning frameworks (PyTorch, TensorFlow) Contributions to deep learning compilers (Triton, TVM, MLIR) Prior experience working on GPUs / CUDA Compensation At Anyscale, we take a market-based approach to compensation. We are data-driven, transparent, and consistent. As the market data changes over time, the target salary for this role may be adjusted. This role is also eligible to participate in Anyscale's Equity and Benefits offerings, including the following: Stock Options Healthcare plans, with premiums covered by Anyscale at 99% for both employees and dependents 401k Retirement Plan Education & Wellbeing Stipend Paid Parental Leave Fertility Benefits Paid Time Off Commute reimbursement 100% of in-office meals covered Anyscale Inc. is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Anyscale Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish Anyscale
$200.8k - $251k
...a machine learning framework for large language models. Candidates should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full-time position offers a competitive salary range of $200,800 - $251...SuggestedFull time- ...matters to the world. The Production Engineering Team Examples of key exciting problems... ..., actual state inspection, and distributed command execution. One interface for the... ...move on. You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks...SuggestedLocal area
$200k - $260k
...vector database team at Redis, shipped 100+ LLM applications, and is a contributor to... ...assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis... ...desire to ship. ~7+ years of software engineering experience comprising of: ~5+ years...SuggestedWork at officeShift work- Mirai Labs in San Francisco seeks engineers to join a senior team building the full on-device stack for real-time local intelligence. You will primarily work on uzu, our inference engine, and focus on supporting new modalities and a variety of features. The ideal candidates...SuggestedLocal area
- ...demand grow Most infrastructure engineers joining an AI cloud inherit a... ...building a new kind of inference cloud, turning a heterogeneous... ...this if you’ve built serious distributed systems, cloud infrastructure... ...simply operating mature ones. LLM serving experience with tools...Suggested
- Member of Technical Staff, Inference Performance 200k base + equity I’m... ...full path from model and serving engine through kernels, accelerators,... ...PyTorch, vLLM, SGLang, TensorRT-LLM, or similar frameworks GPU, accelerator, HPC, or distributed-compute infrastructure Low-...
- ...Baseten powers mission-critical inference for the world's most dynamic AI... ...us and help build the platform engineers turn to to ship AI products.... ...the global operating system for distributed, heterogeneous AI hardware. We believe that as LLM and multi-modal workloads scale...Full timeFlexible hours
$227.2k - $324.5k
...the Role:As a Staff Software Engineer on the ML Infrastructure team... ...world-class machine learning inference platforms. These platforms power... ...that support Deep Learning, LLM, and Search models. This... ...throughput, and low latency distributed systems using ScalaBuild reusable...Full timeTemporary workLocal areaFlexible hours- ...Distributed Systems Engineer @ Dedalus Labs Mission Dedalus Labs is an AI research neolab building infrastructure for AI agents. We're building the persistent compute layer that powers the next generation of autonomous software. Our platform spans distributed...Work at officeVisa sponsorshipRelocation package
- ...Distributed Systems EngineerAs a distributed systems engineer, you'll work across the stack to solve problems as they come up and help build Archil volumes. You'll have significant influence over the technical and product direction.We'll expect you to be able to:Be oncall...Flexible hours
$175k - $250k
..., dental) Job Details We're looking for an engineer to help build and maintain a high-performance inference library designed to support modern AI models across... ...ideal for an engineer who understands how modern LLM inference systems work under the hood and enjoys squeezing...Local area- ...Senior Systems Engineer San Francisco, California Onsite or Remote... ..., database performance, AI inference infrastructure: you can cover... ...disaster recovery. Own AI and LLM inference infrastructure end-... ...for instrumentation and distributed tracing across complex environments...Remote workWork from home
- ...more time putting knowledge into action. We're looking for engineers who want to build the operating system for AI Data Applications... .... About the role We're looking for experienced distributed systems engineers to build the core infrastructure for our durable...
- ...Francisco is seeking a Member of Technical Staff to design and build distributed systems for AI workloads. The role involves developing... ...-grade APIs. Ideal candidates should have strong software engineering skills and experience with distributed systems. This role is...
- B Capital is seeking a data engineer to ensure high data quality for training AI models. You will own the upstream data quality for LLM post-training and design automated QA methods in a collaborative environment. Ideal candidates will have strong engineering skills, a...
$176k - $220k
...institutions Work together with engineers, scientists, operators, and... ...Handshake is hiring a Senior LLM Platform Engineer to join our... ..., and hosted or self-hosted inference. You’ll also contribute to... ...Experience building observability for distributed systems and leading...Full timeWork at officeRemote workFlexible hours$150k - $240k
...we're redefining how the world's most ambitious organizations access and manage energy. The Role As a Software Engineer focusing on Distributed Systems at Verse, you will work in collaboration with some of the brightest industry experts in the field building cloud...Remote workFlexible hours- ...ABOUT THE ROLE You build and operate the inference systems that serve our models in... ...with running real workloads. This is an engineering role, not a research role. You'll measure... ...large‑scale serving infrastructure Strong distributed systems experience; you've been on‑call...
- ...is the neocloud for alternative chips. Inference is fragmenting: purpose-built silicon from... ...roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in... ...serving system. Direct experience with LLM inference serving — request batching, KV-...
$165k - $310k
Senior Research Engineer, LLM Training & Post-Training New York, New York, United States; Remote... ..., training, and production inference, with security, observability, and control... ...model training, post‑training, PyTorch, distributed systems, and AI systems engineering to...For contractorsFor subcontractorWork at officeRemote workWork from homeFlexible hours2 days per week$176k - $209k
Software Engineer (Agentic Systems) Bay Area, US Dialpad is the AI platform for customer... ...implement emerging agent frameworks, LLM inference optimization, advanced retrieval systems... ...Background: Strong foundations in scaling distributed systems and production-grade...Work at office$153k - $376k
...infrastructure is at the heart of everything we build. As a Software Engineer on our Infrastructure team, you'll help design, build, and... .... We're scaling fast, and we're looking for experienced distributed systems engineers across a variety of teams. Whether you're passionate...Minimum wageFull timeLocal areaRemote workWorldwideFlexible hours$180k - $310k
...technical investments with rapid shipping velocity. As Software Engineer on the Platform team, you'll collaborate across frontend,... ...matters most. What You'll Do Design and implement scalable APIs, distributed systems, and data infrastructure that serve millions of users...Full timeWork at officeWork from home$180k - $250k
...unified platform where high-performance inference, orchestration, and observability come... ...role: You are an experienced software engineer who thrives on building large-scale... ...You have deep expertise in large scale distributed systems that deal with high complexity,...Full timeCurrently hiringRemote workRelocation package$170k - $260k
| Software Engineer, Distributed Systems (Core) | Title of Role: | Software Engineer, Distributed Systems (Core) | Location: San Francisco, CA, remote Company Stage of Funding: Series C - Software Development Office Type: Remote Salary: $170K-$260K Company Description...Work at officeRemote workVisa sponsorship$300 per month
...At Crusoe, our Production Engineering team ensures the reliability... ...with a strong background in distributed systems and cloud services to... ...focus on serving and scaling LLM workloads Define, measure,... ...optimize large-scale training and inference clusters Automate observability...Temporary work$180k - $310k
...makes Gamma magical. This means designing distributed systems for real-time content scanning,... ...velocity. You'll collaborate across engineering, product, and design to define how Gamma... ...suspicious or malicious activity Leverage AI/LLM-based detection to stay ahead of AI-...Full timeWork at officeWork from home$172.5k - $260.1k
...ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforce is the #1 AI CRM,... ...optimize, we have to balance this with global scale and traffic distribution to enhance our end user experience. Edge is hiring a backend...$146.5k
...preferences. About the team: The ML Data Engineering team powers metadata extraction,... ...machine learning, data engineering, and distributed systems, collaborating closely with... ...product teams to deploy scalable ML and LLM-powered solutions in production. Role...Full timeLocal areaWorldwideHome officeFlexible hours$350k
...for an exceptional AI systems engineer to lead the design and... ...checkpoints Interpretability: Inference stacks that are as performant... ...models Behavior elicitation: Distributed RL training and roll-outs allowing... ...and scale) Bonus: can set up LLM pipelines, e.g. multiple...Visa sponsorshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Distributed LLM Inference Engineer. Be the first to apply!


