LLM Inference Frameworks Engineer High-Performance Serving
Together AI
Together AI is recruiting a highly skilled Inference Frameworks and Optimization Engineer to design and optimize distributed inference engines for multimodal models at scale. You will focus on low-latency, high-throughput inference, GPU/accelerator optimization, and software-hardware co-design to enable efficient deployment of LLMs and vision models. Join a research-driven team shaping AI infrastructure, collaborating with researchers to craft end-to-end serving pipelines and push the boundaries #J-18808-Ljbffr Together AI
$160k - $230k
...efficient and scalable inference for large... ...inference frameworks, algorithms, and... ...boundaries of performance, scalability, and... ...and Optimization Engineer to design,... ...on low-latency, high-throughput inference... ...the future of LLM inference... ...high-performance serving.Apply CUDA graph...PerformanceFull time- ...owned AI. Our mission is to build highly scalable and efficient infrastructure... ...seeking an expert in high‑performance LLM serving systems and inference optimization. In this role, you will... ...debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT....Performance
$170k - $245k
...About the roleAs a Distributed LLM Inference Engineer, you will help systems and... ...push the boundaries of performance for inference at large scale... ...Batch and Online inference at high scale which will be used by... ...learning and deep learning frameworks (e.g. PyTorch)Solid...PerformanceWork at office- Anyscale is seeking a Distributed LLM Inference Engineer in San Francisco, California. This pivotal... ...involves pushing the boundaries of performance for ML inference at scale. You'll work... ...and familiarity with deep learning frameworks, ideally with experience in PyTorch and...Performance
- ...build and operate the inference systems that serve our models in production... .... This is an engineering role, not a research role... ...serving large models at high throughput Own the performance characteristics of those... ...inference / serving frameworks Experience with mixed...Performance
- ...unicorn founders and senior engineers with deep expertise in 3D,... ...a Founding Engineer, ML Inference with deep expertise in high-performance ML engineering. This is a... ...'ll work across the model-serving stack, designing novel inference frameworks, optimizing inference performance...PerformanceRelocationVisa sponsorshipRelocation package
- Senior ML Systems Engineer, Frameworks & Tooling at Cohere Our mission is to scale intelligence to serve humanity. We’re training and deploying... ...for large-scale LLM training. Design distributed... ...hardware configurations support high-performance training. Investigate and...PerformanceFull timeWork at officeRemote workFlexible hours
$180k - $270k
...elevate productivity and performance through note-taking... ...building and deploying high-throughput, ultra-low-latency inference engines for large language... ...experience with: Frontier Serving Frameworks: Deep, under-the-hood... ...with modern LLM serving frameworks like...PerformanceFull timeWork at officeWorldwide- Mirai Labs in San Francisco seeks engineers to join a senior team building the full on... .... You will primarily work on uzu, our inference engine, and focus on supporting new modalities... ...work, and experience in writing high-performance GPU kernels or Rust systems...PerformanceLocal area
- Causal Labs in San Francisco is seeking an experienced Infrastructure Engineer to build high-throughput inference systems for large-scale evaluation and backtesting against historical physical observations. You will design techniques to improve latency and throughput, optimize...Performance
- ...and production systems. You will own the bridge between research checkpoints and production-ready inference services, designing scalable APIs and high-performance serving. You will optimize GPU workloads, manage distributed systems, and collaborate with researchers to...Performance
$200k - $300k
...research to systems engineering to product design... ...looking for a Performance Engineer to make... ...models train and serve as fast as the hardware... ...Optimize inference and serving end to... ...and low-latency, high-throughput sampling... ...knowledge of ML framework internals (PyTorch...PerformanceFull time- ...are forming small, highly capable product... ...management, design, and engineering together with the... ...quality issues, performance bottlenecks, and... ...create evaluation frameworks, and establish monitoring... ...— configuring LLM-based products,... ...approximately 791,000 people serving clients in more...PerformanceFull timeWork experience placementLive inWork at officeLocal area
- ...NEAR AI in San Francisco is seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries... ...requires deep hands-on experience with inference engines such as SGLang, vLLM or TensorRT, GPU architectures,...Performance
$300 per month
...strategies, and be part of a high-performing team that believes in each... ...At Crusoe, our Production Engineering team ensures the... ...services with a focus on serving and scaling LLM workloadsBuild automation... ...distributed AI pipelines and inference servicesDefine, measure, and...PerformanceTemporary work- ...worldwide.We’re a team of engineers, clinicians, and... ...work helps care teams perform with greater... ...improving and integrating high performance robotic AI... ...to real-time onboard inference—while serving as a core contributor... ...learning models, and frameworks such as PyTorch, TensorFlow...PerformanceLocal areaWorldwideFlexible hours
$165k - $206k
...the world’s best engineers, scientists, designers... ...the heavy lifting.Serve as a cross-... ...supportLeverage Cursor with LLM-pair programming (... ...time.Produce high-fidelity technical... ..., integration framework (EIBs, Workday Studio... ...procedures, CTEs, performance tuning) for data profiling...Performance$220k - $320k
A tech startup specializing in AI inference seeks a skilled professional to optimize their inference stack. Candidates should have over... ...ML systems, fluency in Python, and hands-on experience with LLM frameworks. The role offers competitive compensation of $220,000 to $320...PerformanceLocal area$264.8k - $331k
...Systems Research Engineer, Agent Post-training... ...to reach the performance necessary for complex... ...resources that serve all of our... ...our training and inference framework. Post-train state... ...least 1-3 years of LLM training in a production... ...provide the high-quality data and...PerformanceFull timeContract workFor contractorsFor subcontractorWork at office$315k
...committed researchers, engineers, policy experts,... ...and addressing performance issues across many... ...research, training, and inference. A significant... ...experience with: High performance, large... ...architecture ML framework internals Language... ...means a veteran who served on active duty in...PerformanceContract workFor contractorsFor subcontractorWork at officeRelocationVisa sponsorshipWork visaFlexible hours$180k - $237.5k
...Senior Software Engineer – Site Controller, Energy... ...batteries into a single, high-performance energy asset. We are... ...fault-management frameworks, designing the state... ...our "Pack Manager" to serve as a universal translator... ...information, and inferences drawn from your PI. We...PerformanceFull timeLocal area$138.68k - $174.43k
APPLIED AI ENGINEER (1042) - Department of Technology... ...process and shall serve at the discretion... ..., integrating LLM APIs, and ensuring... ...on, creative, and highly collaborative —... ...logging, alerting, and performance metrics to ensure... .... Use frameworks like LangChain, LlamaIndex...PerformancePermanent employmentFull timeTraineeshipWork at officeRemote workWork from homeFlexible hoursNight shift1 day per week- Vast.ai is seeking a systems engineer to scale AI inference and optimize GPU performance at our San Francisco or Los Angeles offices. You will... ...have advanced C++, experience with parallel frameworks, and a strong track record in high-performance systems. This role is full-...PerformanceFull time
$166.6k - $208.3k
...deployment, real-time inference, observability, and... ...themselves — and we serve low-latency, highly available scores to the decision engine that depends on them.... ...versioning, CI/CD with performance, bias, and consistency... ...in Python, with API frameworks like FastAPI or FlaskExperience...Performance$300 per month
...strategies, and be part of a high-performing team that believes in... ...Hardware Systems Engineer to strengthen Crusoe’... ...across training and inference - dense, MoE, long-... ..., or data-analysis frameworks using Python, Shell,... ...Experience with inference serving frameworks, training...PerformanceTemporary work- ...mission-critical inference for the world's most... ...the platform engineers turn to to ship AI... ...We believe that as LLM and multi-modal workloads... ...for Disaggregated Serving, Wide Expert... ...validate networking performance on bleeding-edge clusters... ...experience with high-performance...PerformanceFull timeFlexible hours
- ...and more. We focus on high-performance model inference and accelerating... ...this role, you’ll lead engineering efforts to ensure... ...infrastructure for serving frontier AI models in... ...Experience with inference frameworks like TensorRT, vLLM,... ...ROCm, HIP, TensorRT-LLM, Ray Serve, Megatron...PerformanceFull time
- Sail is hiring for an engineering role in San Francisco to design and implement high-performance schedulers that optimize admission control, queuing, and fairness across... ...KV caching for memory/compute trade-offs in LLM inference stacks. You will contribute to deep...Performance
$286.2k - $326.7k
...Senior Distinguished Engineer, AI Compute (... ...experiences and scalable, high-performance AI infrastructure... ...reimagine how we serve our customers and... ...training, model inference and feature... ...compute frameworks including Spark /... ...diverse workloads from LLM pre-training and...PerformanceFull timePart timeLocal areaRemote work- ...looking for a Founding Engineer & CTO to be the... ...from data pipelines to LLM orchestration to client... ...CD pipelines, testing frameworks, and infrastructure-as... ...engineering team, fostering a high-performance, collaborative culture... ...tenant SaaS platforms serving enterprise clients....PerformanceFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to LLM Inference Frameworks Engineer High-Performance Serving. Be the first to apply!
- performance test architect San Francisco, CA
- performance food service San Francisco, CA
- performance improvement consultant San Francisco, CA
- senior performance engineer San Francisco, CA
- performance nutrition San Francisco, CA
- system performance engineer San Francisco, CA
- IT performance management San Francisco, CA
- acting performance San Francisco, CA
- application performance engineer San Francisco, CA
- performance testing San Francisco, CA


