Distributed LLM Inference & Optimization Engineer
Together AI
Together AI is building state-of-the-art infrastructure to enable efficient and scalable inference for large language models (LLMs). We seek an Inference Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that support multimodal and language models at scale. This role focuses on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for scalable LLM and vision model deployment. #J-18808-Ljbffr Together AI
$160k - $230k
...efficient and scalable inference for large language models... ...LLMs). Our mission is to optimize inference frameworks,... ...Frameworks and Optimization Engineer to design, develop, and optimize distributed inference engines that... ...to shape the future of LLM inference infrastructure...SuggestedFull time$170k - $245k
...AnyscaleAt Anyscale, we're on a mission to democratize distributed computing and make it accessible to software... ...raised to date.About the roleAs a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for...SuggestedWork at office- Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal role involves pushing the boundaries of performance for ML inference at scale. You'll work closely with product teams to deliver end-to-end solutions while leveraging open...Suggested
$190.9k - $232.8k
...leading data and AI company is seeking a Staff Software Engineer for GenAI inference to lead the architecture and optimization of the inference engine. The role requires expertise in CUDA, GPU programming, and distributed systems design. Ideal candidates will have a strong...Suggested- Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across a global... ...caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep...Suggested
- ...Francisco is seeking a talented engineer to design and implement... ...fast and cost-efficient AI inference at global scale. You will be... ...-performance schedulers and optimizing global routing while focusing... ...has a strong background in distributed systems and is eager to engage...
- ...AI in San Francisco seeks a hardware‑aware software engineer to optimize GPU systems for training and inference across image, video, and world‑model workloads. You will push performance from kernel code to distributed engines, profiling bottlenecks and implementing practical...
- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI... ...operating system for distributed, heterogeneous AI hardware... .... We believe that as LLM and multi-modal workloads... ...distributed inference optimizations. THE OPPORTUNITY...Full timeFlexible hours
$200.8k - $251k
...technology company in San Francisco seeks a team member to build and optimize a machine learning framework for large language models.... ...should have system optimization experience and solid software engineering skills, particularly in tools like CUDA and Pytorch. This full...Full time$227.2k - $417k
...the Role:As a Software Engineer on the ML... ...class machine learning inference platforms. These platforms... ...support Deep Learning, LLM, and Search models. This... ...throughput, and low latency distributed systems using... ...approach to identifying & optimizing latency, cost, and efficiency...Full timeTemporary workLocal areaFlexible hours- Inception is seeking engineers and scientists to design, optimize, and scale the diffusion LLM serving systems powering production inference. Your work will help make inference faster, more... ...(Kubernetes, Ray, SLURM) for distributed inference, evaluation, and large-batch...
- Inception in San Francisco is seeking experienced backend engineers to own the systems that serve our diffusion LLMs in... ...build and operate infrastructure that handles billions of inference requests, optimizing for latency, throughput, cost, and reliability. This role...
- ...Team Our team analyzes inference stack performance across the... ...understanding into performance optimizations and models that project... ...from first principles about distributed systems, model inference,... ...Enjoy collaborating with engineering and research teams to improve...Full time
$298k - $368k
...with downstream teams on the optimization and integration into the Waymo... ...set of sensors, enabling engineers like you to (1) develop methods... ...You will: Design VLM/LLM model architecture and drive... ...expertise in low-latency on-device inference techniques and a deep...Full timeRemote work- Kindredventures is recruiting infrastructure engineers to scale large-scale inference and evaluation around a physics-based LPM initiative... ...will work on high-throughput systems, latency-optimized serving, and distributed orchestration across Kubernetes, Ray, and Slurm....
$150k - $300k
Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over...Relocation package$225k
Dormont Manufacturing Co is looking for a Software Engineer on the Inference & RL Systems team in San Francisco. The role involves designing distributed systems, optimizing performance, and ensuring high reliability for RL and post-training workflows. The ideal candidate...- ...San Francisco, CA, is seeking a Member of Technical Staff for distributed systems to design, build, and operate the platform that schedules... ...-growing environment. You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie Labs while shaping...
- Jobot is seeking an engineer to build and maintain a high-performance inference library for modern AI models across diverse compute architectures. The role targets someone who understands how LLM inference works under the hood and aims to squeeze maximum performance from...
- ...seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation... ...improve latency and throughput, optimize the inference stack to exhaust... ...extend Kubernetes, Ray, and Slurm for distributed inference, ensuring reliability...
$300k
...of committed researchers, engineers, policy experts, and business... ...About the role Our Inference team is responsible for... ...models. We tackle complex, distributed systems challenges across... ...traffic management systems LLM inference optimization, batching, and caching...Full timeWork at officeWorldwideVisa sponsorshipFlexible hours- ...About the Team OpenAI’s Inference team powers the... ..., fast-moving team of engineers focused on delivering... ...interaction. You'll build and optimize the systems that let... ...that span networking, distributed compute, and high-... ...tooling like vLLM, TensorRT-LLM, or custom model...Full time
- ...powers mission-critical inference for the world's most... ...help build the platform engineers turn to to ship AI products... ...Stack team builds the distributed runtime that powers large-scale LLM inference across our... ...to make new inference optimizations broadly available to customers...Full timeFlexible hours
- ...high-performance model inference and accelerating... ...to drive the design, optimization, and scaling of our inference... ...role, you’ll lead engineering efforts to ensure our... ...development, and distributed inference best practices... ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron,...Full time
$249.5k - $273.5k
...Applied Research, Design, and Engineering leadership, you will lead a... ...emerging agent frameworks, inference optimization techniques, retrieval... ...Experience in AI, ML platforms, LLM systems, or agentic... ...Strong foundations in scaling distributed systems and production-grade...Work at office$300k
...of committed researchers, engineers, policy experts, and business... ...the Role The Cloud Inference team scales and optimizes Claude to serve the... ...-performance, large-scale distributed systems serving millions of... ...Strong familiarity with LLM inference optimization, batching...Full timeWork at officeVisa sponsorshipFlexible hours- ...browsers running at scale, solving massive distributed systems challenges and making sure our... .... Work closely with the rest of Engineering, gathering input and providing great... ...databases, automated testing, performance optimization, and zero-downtime multi-region...Full timeImmediate startRelocation
$180k - $275k
...rapid shipping velocity. As Software Engineer on the Platform team, you'll... ...Design and implement scalable APIs, distributed systems, and data infrastructure that serve... ...traffic production systems and performance optimization ~ Track record shipping high-quality,...Full timeWork at officeWork from home- ...Sensor Data Integration Engineers build the algorithms... ...consistency of our data Optimize the performance of our... ...parallel computing or distributed systems A bachelor'... ...— orchestrating LLM-driven workflows for triage... ...feed ML training and inference. Familiar with C++....Full timeWork at officeWork from home
- ...constraints of robotic platforms. About the Role As a Software Engineer, Distributed Data Systems, you will design and scale the infrastructure... ...translate them into production-ready systems. Harden, optimize, and maintain critical data infrastructure systems that...Full timeWork at officeRelocation package
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Distributed LLM Inference & Optimization Engineer. Be the first to apply!

