Runtime Engineer — LLM Inference & Host Stack
MatX Inc.
MatX Inc. seeks a systems programmer to build the host-side interface library and manage the compiler→runtime contract. You will design the custom-kernel ABI, and implement Python bindings to move tensors from Python to accelerator hardware. The role involves working with CUDA/ROCm-style accelerators, memory models, and high-performance computing stacks, delivering efficient runtime and serving throughput. #J-18808-Ljbffr MatX Inc.
$193.3k - $261.5k
We are looking for a Senior Inference Engineer to own inference for real-time... ...multimodalconversational AI. This is a full-stack inference role: you will work... ..., building thereal-time runtime that serves it within hard... ...that fall outside standard LLM serving patterns — sustained...SuggestedInternshipLocal areaFlexible hours$120k - $250k
...for large-language-model inference and training, with HW/SW co... ...benefits from the others. The runtime owns the host-side stack and the contracts that... ...downstream consumersBuild the LLM inference serving stack —... ...the Python surfaces ML engineers actually use — and hit measurable...SuggestedFull timeContract workWork experience placementLocal areaRemote workMonday to FridayFlexible hours- ...industry-leading training and inference speeds; over 10 times faster than... ...hiring a Senior Performance Engineer to join our Product team. You... ...resident expert on how Cerebras stacks up against alternative... ...stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization...SuggestedContract workShift work
- ...are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance... ...explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive... ...AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data...Suggested
$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization!... ...What does it take to push every LLM inference operation to its... ...accelerating NVIDIA's LLM inference stack. The first is GPU kernel... ..., compiler decisions, and runtime scheduling.Direct experience...SuggestedFull time$195.2k - $361.2k
...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments... ...Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling...Full timeInternshipLocal areaImmediate startShift work- ...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments... ...Python; comfortable reading systems‑level code Experience with LLM inference. (attention, KV cache, decoding) Experience profiling...Local areaShift work
$184k - $287.5k
...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior... ...layers of the hardware/software stack from GPU architecture to Deep... ...language and multimodal model inference as part of NVIDIA Inference... ...deliver production code to TRT-LLM, NVIDIA’s open-source inference...Full time$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large... ...’s full hardware and software stack. This role sits at the... ...analyze and optimize frontier-scale LLM workloads running on thousands... ...behavior to framework/runtime internals, CUDA libraries, communication...Full time- NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal candidate will have a PhD and 3+ years... ...in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks...
- ...industry-leading training and inference speeds; over 10 times faster... ...inference. About The Role The Host and Network IO Team develops... ...down to the custom RoCE network stack implemented in Cerebras'... ...PhD in Computer or Electrical Engineering + 3 years industry experience...
$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...implement high-performance inference stacks, optimize GPU kernels and compilers... ...crowdExperience building and optimizing LLM inference engines (e.g., vLLM,...Full time$152k - $241.5k
...TensorRT team as a Senior Software Engineer, and be at the forefront of... ...enabling high-performance AI inference solutions for automotive... ...into TensorRT's compiler and runtime for specialized and constrained... ...the hardware and software stack to understand and leverage new...Full time$166k - $244k
...the Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-... ...usage, improving data streaming between host and device, and tuning execution... ...routing multi-sensor inputs into model inference loops.Collaborate with machine learning...Full time$207k - $300k
...friction points in Google’s AI stack, converting them into... ...product feature requests for the Engineering teams.Co-build with customer... ...systems, navigating real-time inference constraints, and implementing... ...experimentation.Knowledge of "LLM-native" metrics (tokens/sec,...Local area$152k - $241.5k
...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...problems for AI workloads (both inference and training) and successfully transition... ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language...Full time$184k - $287.5k
...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis... ...on optimizing NVIDIA’s high-performance LLM software stack in frameworks like PyTorch and JAX for high-performance training...Full timeWork experience placement$138k - $206k
...hardware and software engineers to identify and address... ...optimize the software stack required to maximize performance... ...modern and emerging LLM workloads.We are... ...and serving systems to runtime software, networking, memory... ..., disaggregated inference, and Mixture-of-Experts...Work experience placementWork at officeFlexible hours$198k - $326k
...AI across LinkedIn. The LLM Serving team builds the... ...Senior Staff Software Engineer with deep expertise at... ...infrastructure, and large-scale inference. This is a highly... ...models interact with runtimes, compilers, and hardware... ...across the full stack, including model architecture...For contractorsWork at officeFlexible hours- ...integration of hardware and software performance engineering. We seek a senior engineer who will shape core... ...the team as they implement high-throughput inference systems. You will own the inference engine, optimize runtimes, and push the boundaries of latency and throughput...
- ...week.The role: Senior toStaff Runtime Systems EngineerWhat You Will... ...focusing on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture,... ..., and systems software that hosts this SoC.In this role, you...3 days per week
$152k - $241.5k
...a motivated Deep Learning engineer to bring advanced communication... ...technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You... ...scales up to 100K GPUs to inference down at microsecond... ...at least one communication runtime (NCCL, NVSHMEM, MPI). Good...Full timeRemote work$138k - $206k
Senior LLM Systems Performance Engineer The AGI Computing Lab’s STG group is looking... ...agentic workflows, distributed inference systems, disaggregated... ...benchmarks. Analyze the impact of runtime, memory hierarchy,... ...hardware and software stack. Collaborate with hardware...Work at officeFlexible hours$184k - $287.5k
We are now looking for a Senior High-Performance LLM Training Engineer! NVIDIA is seeking experienced engineers specializing in performance analysis... ...on optimizing NVIDIA’s high‑performance LLM software stack in frameworks like PyTorch and JAX for high‑performance training...Work experience placement- Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce...
- ...looking for a creative, skilled, and motivated engineers to join our founding team in advancing... ...Qualifications ~5+ years full-stack development experience ~ Experience in... ...platform experience (AWS, Azure, GCP) VLM and LLM application development experience...Full time
$180k - $250k
A leading AI infrastructure firm is seeking a TPU Systems Engineer to develop high-performance systems using JAX, XLA, and Pallas. This... ...workloads on TPU hardware and optimizing performance across the stack. Candidates should have at least 3 years of experience in...$250k - $350k
...world's leading ML systems engineers, including leaders... ...large-scale training, inference, and reinforcement learning... ...of the ML systems stack to maximize performance... ...Design distributed runtimes and scheduling systems... ...Familiarity with vLLM, TensorRT-LLM, or production LLM...Visa sponsorship$218.8k - $335.3k
...perception, planning, and controls stack that keeps the vehicle... ...looking for a Staff Software Engineer to provide technical leadership... ...robustness, and predictable runtime behavior under tight latency... ...with GPU/accelerator‑based ML inference, model deployment, and performance...Full timeLocal areaRemote workWork from homeFlexible hours$152k - $241.5k
...Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-... ...including vLLM, SGLang, and TensorRT-LLM as well as with external... ...and remove bottlenecks across runtimes, kernels, networking, routing,...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Runtime Engineer — LLM Inference & Host Stack. Be the first to apply!


