Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Runtime Engineer — LLM Inference & Host Stack

MatX Inc.

MatX Inc. seeks a systems programmer to build the host-side interface library and manage the compiler→runtime contract. You will design the custom-kernel ABI, and implement Python bindings to move tensors from Python to accelerator hardware. The role involves working with CUDA/ROCm-style accelerators, memory models, and high-performance computing stacks, delivering efficient runtime and serving throughput. #J-18808-Ljbffr MatX Inc.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Runtime Engineer — LLM Inference & Host Stack in Mountain View, CA vacancy
  • $193.3k - $261.5k

    We are looking for a Senior Inference Engineer to own inference for real-time...  ...multimodalconversational AI. This is a full-stack inference role: you will work...  ..., building thereal-time runtime that serves it within hard...  ...that fall outside standard LLM serving patterns — sustained... 
    Suggested
    Internship
    Local area
    Flexible hours

    Amazon

    Sunnyvale, CA
    18 hours ago
  • $120k - $250k

     ...for large-language-model inference and training, with HW/SW co...  ...benefits from the others. The runtime owns the host-side stack and the contracts that...  ...downstream consumersBuild the LLM inference serving stack —...  ...the Python surfaces ML engineers actually use — and hit measurable... 
    Suggested
    Full time
    Contract work
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    4 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster than...  ...hiring a Senior Performance Engineer to join our Product team. You...  ...resident expert on how Cerebras stacks up against alternative...  ...stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization... 
    Suggested
    Contract work
    Shift work

    Cerebras Systems

    Sunnyvale, CA
    18 hours ago
  •  ...are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance...  ...explain performance across the full stack, from GPU silicon through the software runtime, and drive competitive...  ...AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data... 
    Suggested

    AMD

    Santa Clara, CA
    18 hours ago
  • $184k - $287.5k

    We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization!...  ...What does it take to push every LLM inference operation to its...  ...accelerating NVIDIA's LLM inference stack. The first is GPU kernel...  ..., compiler decisions, and runtime scheduling.Direct experience... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $195.2k - $361.2k

     ...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments...  ...Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling... 
    Full time
    Internship
    Local area
    Immediate start
    Shift work

    Intel

    Santa Clara, CA
    18 hours ago
  •  ...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments...  ...Python; comfortable reading systems‑level code Experience with LLM inference. (attention, KV cache, decoding) Experience profiling... 
    Local area
    Shift work

    PVH (Tommy Hilfiger/Calvin Klein)

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior...  ...layers of the hardware/software stack from GPU architecture to Deep...  ...language and multimodal model inference as part of NVIDIA Inference...  ...deliver production code to TRT-LLM, NVIDIA’s open-source inference... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Principal Engineer to drive the performance of large...  ...’s full hardware and software stack. This role sits at the...  ...analyze and optimize frontier-scale LLM workloads running on thousands...  ...behavior to framework/runtime internals, CUDA libraries, communication... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • NVIDIA is seeking a Senior DL Algorithms Engineer to optimize LLM/Omni models and enhance performance across its software stack. The ideal candidate will have a PhD and 3+ years...  ...in deep learning, specifically in inference. This role involves profiling, analyzing bottlenecks... 

    NVIDIA

    Santa Clara, CA
    4 days ago
  •  ...industry-leading training and inference speeds; over 10 times faster...  ...inference. About The Role The Host and Network IO Team develops...  ...down to the custom RoCE network stack implemented in Cerebras'...  ...PhD in Computer or Electrical Engineering + 3 years industry experience... 

    Cerebras

    Sunnyvale, CA
    2 days ago
  • $184k - $287.5k

     ...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale...  ...implement high-performance inference stacks, optimize GPU kernels and compilers...  ...crowdExperience building and optimizing LLM inference engines (e.g., vLLM,... 
    Full time

    Nvidia

    Santa Clara, CA
    18 hours ago
  • $152k - $241.5k

     ...TensorRT team as a Senior Software Engineer, and be at the forefront of...  ...enabling high-performance AI inference solutions for automotive...  ...into TensorRT's compiler and runtime for specialized and constrained...  ...the hardware and software stack to understand and leverage new... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $166k - $244k

     ...the Role:We are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing on high-...  ...usage, improving data streaming between host and device, and tuning execution...  ...routing multi-sensor inputs into model inference loops.Collaborate with machine learning... 
    Full time

    X Company

    Mountain View, CA
    1 day ago
  • $207k - $300k

     ...friction points in Google’s AI stack, converting them into...  ...product feature requests for the Engineering teams.Co-build with customer...  ...systems, navigating real-time inference constraints, and implementing...  ...experimentation.Knowledge of "LLM-native" metrics (tokens/sec,... 
    Local area

    Google

    Sunnyvale, CA
    18 hours ago
  • $152k - $241.5k

     ...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class...  ...problems for AI workloads (both inference and training) and successfully transition...  ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $184k - $287.5k

     ...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis...  ...on optimizing NVIDIA’s high-performance LLM software stack in frameworks like PyTorch and JAX for high-performance training... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    2 days ago
  • $138k - $206k

     ...hardware and software engineers to identify and address...  ...optimize the software stack required to maximize performance...  ...modern and emerging LLM workloads.We are...  ...and serving systems to runtime software, networking, memory...  ..., disaggregated inference, and Mixture-of-Experts... 
    Work experience placement
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    3 days ago
  • $198k - $326k

     ...AI across LinkedIn. The LLM Serving team builds the...  ...Senior Staff Software Engineer with deep expertise at...  ...infrastructure, and large-scale inference. This is a highly...  ...models interact with runtimes, compilers, and hardware...  ...across the full stack, including model architecture... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    4 days ago
  •  ...integration of hardware and software performance engineering. We seek a senior engineer who will shape core...  ...the team as they implement high-throughput inference systems. You will own the inference engine, optimize runtimes, and push the boundaries of latency and throughput... 

    Sanas

    Palo Alto, CA
    3 days ago
  •  ...week.The role: Senior toStaff Runtime Systems EngineerWhat You Will...  ...focusing on in-memory compute for AI inference in datacenters.This position is for runtime software engineering, working on the architecture,...  ..., and systems software that hosts this SoC.In this role, you... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

     ...a motivated Deep Learning engineer to bring advanced communication...  ...technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You...  ...scales up to 100K GPUs to inference down at microsecond...  ...at least one communication runtime (NCCL, NVSHMEM, MPI). Good... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    18 hours ago
  • $138k - $206k

    Senior LLM Systems Performance Engineer The AGI Computing Lab’s STG group is looking...  ...agentic workflows, distributed inference systems, disaggregated...  ...benchmarks. Analyze the impact of runtime, memory hierarchy,...  ...hardware and software stack. Collaborate with hardware... 
    Work at office
    Flexible hours

    Samsung Semiconductor Inc.

    San Jose, CA
    4 days ago
  • $184k - $287.5k

    We are now looking for a Senior High-Performance LLM Training Engineer! NVIDIA is seeking experienced engineers specializing in performance analysis...  ...on optimizing NVIDIA’s high‑performance LLM software stack in frameworks like PyTorch and JAX for high‑performance training... 
    Work experience placement

    NVIDIA

    Santa Clara, CA
    4 days ago
  • Tilde Research is seeking a Kernel Engineer in Palo Alto to design, implement, and optimize high-performance GPU kernels for training and inference workloads. You will collaborate with ML researchers and engineers to push hardware limits, improve throughput, and reduce... 

    Tilde Research

    Palo Alto, CA
    4 days ago
  •  ...looking for a creative, skilled, and motivated engineers to join our founding team in advancing...  ...Qualifications ~5+ years full-stack development experience ~ Experience in...  ...platform experience (AWS, Azure, GCP) VLM and LLM application development experience... 
    Full time

    Dexmate

    Santa Clara, CA
    18 hours ago
  • $180k - $250k

    A leading AI infrastructure firm is seeking a TPU Systems Engineer to develop high-performance systems using JAX, XLA, and Pallas. This...  ...workloads on TPU hardware and optimizing performance across the stack. Candidates should have at least 3 years of experience in... 

    RadixArk

    Palo Alto, CA
    3 days ago
  • $250k - $350k

     ...world's leading ML systems engineers, including leaders...  ...large-scale training, inference, and reinforcement learning...  ...of the ML systems stack to maximize performance...  ...Design distributed runtimes and scheduling systems...  ...Familiarity with vLLM, TensorRT-LLM, or production LLM... 
    Visa sponsorship

    Periodic Labs

    Menlo Park, CA
    18 hours ago
  • $218.8k - $335.3k

     ...perception, planning, and controls stack that keeps the vehicle...  ...looking for a Staff Software Engineer to provide technical leadership...  ...robustness, and predictable runtime behavior under tight latency...  ...with GPU/accelerator‑based ML inference, model deployment, and performance... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $152k - $241.5k

     ...Senior Deep Learning Algorithms Engineer to advance Dynamo, our open-source distributed inference platform for large-scale, low-...  ...including vLLM, SGLang, and TensorRT-LLM as well as with external...  ...and remove bottlenecks across runtimes, kernels, networking, routing,... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Runtime Engineer — LLM Inference & Host Stack. Be the first to apply!