Distributed LLM Inference Engineer - Scale AI at Speed
Anyscale
Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open-source technologies and contributing to community projects. Candidates should have a solid understanding of distributed systems and familiarity with deep learning frameworks, ideally with experience in PyTorch and Ray. Anyscale offers competitive compensation and extensive benefits, including healthcare coverage and stock options. #J-18808-Ljbffr Anyscale
$170k - $245k
...on a mission to democratize distributed computing and make it accessible... ...accelerate the progress of AI applications out into the... ...developer or data scientist can scale an ML application from their... ...the roleAs a Distributed LLM Inference Engineer, you will help systems and...SuggestedWork at office- ...mission-critical inference for the world's most dynamic AI companies, like... ...build the platform engineers turn to to ship... ...system for distributed, heterogeneous AI... ...believe that as LLM and multi-modal workloads scale, the network is... ...operates at wire-speed. In this role...SuggestedFull timeFlexible hours
$160k - $230k
...the RoleAt Together.ai, we are building... ...efficient and scalable inference for large language... ...and Optimization Engineer to design, develop, and optimize distributed inference engines that... ...language models at scale. This role will... ...shape the future of LLM inference infrastructure...SuggestedFull time- ...San Francisco is seeking a talented engineer to design and implement robust systems... ...that ensure fast and cost-efficient AI inference at global scale. You will be responsible for... ...candidate has a strong background in distributed systems and is eager to engage in complex...Suggested
$200.8k - $251k
A leading AI technology company in San Francisco seeks a team member to build and optimize... ...experience and solid software engineering skills, particularly in tools like CUDA and... ...range of $200,800 - $251,000, along with comprehensive benefits. #J-18808-Ljbffr Scale AISuggestedFull time$160k - $194k
...What is Verse? The race to AI has become the race to... ...built by pioneers in grid-scale batteries, energy markets,... ...The Role As a Software Engineer focusing on Distributed Systems at Verse, you will... ...Balance & Precision: We believe speed and perseverance must be accompanied...Full timeRemote workFlexible hours- ...company is seeking a talented software engineer to join their dynamic Inference team. This role involves designing... ...infrastructure for large-scale multimodal models, focusing on high-... ...product teams to push the boundaries of AI technology, ensuring reliable production...
- Vast.ai Inc. is seeking a systems engineer with HPC or parallel programming experience to help scale AI inference. You will design and optimize GPU kernels and tensor libraries, leveraging CUDA/C++ and related frameworks to push the bleeding edge of AI performance. This...
- ...the Team OpenAI’s Inference team powers the... ...fast-moving team of engineers focused on delivering... ...of what AI can do. We’re expanding... ...models at scale. You’ll be part of... ...span networking, distributed compute, and high-... ...like vLLM, TensorRT-LLM, or custom model parallel...Full time
- ...Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor,... ...and help build the platform engineers turn to to ship AI products.... ...Inference Stack team builds the distributed runtime that powers large-scale LLM inference across our...Full timeFlexible hours
$300k
...interpretable, and steerable AI systems. We want... ...researchers, engineers, policy experts,... ...role Our Inference team is responsible... ...tackle complex, distributed systems challenges... ...performance, large-scale distributed systems... ...systems LLM inference optimization...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...state-of-the-art AI models - unlocking... ...performance model inference and accelerating research... ...optimization, and scaling of our inference... ...role, you’ll lead engineering efforts to ensure... ...development, and distributed inference best... ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron...Full time
- Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control... ..., and explore KV caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep observability, tracing...
- ...NEAR AI was started by Illia Polosukhin, co-author of the... ...open‑source AI at a global scale. We are specifically seeking... ...expert in high‑performance LLM serving systems and inference optimization. In this role,... ...optimizing major inference engines such as SGLang, vLLM, or...
$190k - $265k
...enabling data and AI teams to solve the... ...business. Founded by engineers — and customer-... ...interfacing with data to scaling our services and... ...Foundation Model Inference team is the... ...you will have:Build LLM infrastructure powering... ...and efficiency of distributed AI...Local areaWorldwide- ...AI/ML Engineer (RL & Physical Systems) FLUIX is... ...systems to power distribution, where milliseconds... ...and real megawatt-scale infrastructure.... ...Support integration of LLM-based tools and... ...knowledge distillation, inference orchestration, etc... ...at startup speed. Bonus Points...Weekend work
$180k - $275k
...About the role You'll build and scale the application and data... ...shipping velocity. As Software Engineer on the Platform team, you'll... ...and implement scalable APIs, distributed systems, and data infrastructure... ...systems, event pipelines, or AI-powered applications (Nice to...Full timeWork at officeWork from home- ...powers web browsing capabilities for AI agents and applications. We manage... ...team keeps our browsers running at scale, solving massive distributed systems challenges and making sure our... ...APIs. Work closely with the rest of Engineering, gathering input and providing great...Full timeImmediate startRelocation
- ...robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure that powers large-scale... ...and rapid change About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring...Full timeWork at officeRelocation package
$166k - $225k
...running the world's best data and AI infrastructure platform so... ...their business. Founded by engineers — and customer obsessed — we... ...for interfacing with data to scaling our services and infrastructure... ...building the next generation distributed data storage and processing...Local areaWorldwide$300k
...interpretable, and steerable AI systems. We want AI... ...researchers, engineers, policy experts, and... ...Role The Cloud Inference team scales and optimizes Claude... ...performance, large-scale distributed systems serving millions... ...familiarity with LLM inference optimization...Full timeWork at officeVisa sponsorshipFlexible hours- ...out into multiple AI inference requests running in... ..., our inference engineers and researchers build... ...Keep long-running distributed training jobs... ...managed GPU clusters at scale: NVIDIA hardware,... ..., or TensorRT-LLM. Slurm or other HPC... ...but notable. High-speed interconnects:...Shift work
$176k - $209k
...DialpadDialpad is the AI platform for customer... ...fostering an AI-native engineering culture.This position... ...observability systems.Build & Scale: Design and deploy... ...agent frameworks, LLM inference optimization, advanced... ...foundations in scaling distributed systems and production...Work at office$153k - $376k
...designs into code, or iterating with AI. From idea to product, Figma... ...everything we build. As a Software Engineer on our Infrastructure team, you’... ...millions of people worldwide. We’re scaling fast, and we’re looking for experienced distributed systems engineers across a...Minimum wageFull timeLocal areaRemote workWorldwideFlexible hours- Crusoe is seeking a Senior Software Engineer to help scale Crusoe Cloud's container registry. You will own... ...scalable cloud infrastructure. You’ll work across teams to balance speed and reliability, shape system trade-offs, and impact how customers’ AI #J-18808-Ljbffr CrusoeFull time
- Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimization, and software...
- ...Data Integration Engineers build the algorithms... ...transform large-scale geospatial datasets... ...parallel computing or distributed systems A... ...harnesses — orchestrating LLM-driven workflows... ...ML training and inference. Familiar with... ...design with AI-powered geospatial...Full timeWork at officeWork from home
$216.2k - $270.25k
Scale GP (Scale Generative AI Platform) is an enterprise-grade Generative... ...knowledge retrieval, inference, evaluation, and more... ...Senior Full-Stack Engineer to help us build, scale... ...Python, working with distributed systems, data pipelines, and ML/LLM components.Integrate...Full time$325k
A leading AI research company in San Francisco seeks an engineer to optimize their powerful AI models for high-volume production environments. The ideal candidate... ...with ML architectures, and experience with distributed systems. This role involves collaboration with researchers...- ...to build best-in-class AI agents with flexible and... ...a world-class team of engineers, designers, marketers,... ...for our customers as we scale. You will also be a steward... ...Have deep intuition on distributed systems, databases,... ...about trade-offs between speed, scalability, and...Work at officeVisa sponsorshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Distributed LLM Inference Engineer - Scale AI at Speed. Be the first to apply!


