Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Runtime Systems Engineer - Host Stack & AI Inference

MatX

MatX is building custom silicon for LLM inference and training. You will develop the host-side interface library, manage device memory, DMA, streams and events, and extend the executable format to enable safe evolution of compiler-runtime contracts. You will design the kernel ABI and Python bindings (PyO3) to move tensors to the device, and contribute to the LLM inference serving stack with scheduler and memory management features. #J-18808-Ljbffr MatX

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Runtime Systems Engineer - Host Stack & AI Inference in Mountain View, CA vacancy
  • MatX Inc. seeks a systems programmer to build the host-side interface library and manage the compiler→runtime contract. You will design the custom-kernel ABI, and implement Python...  ...models, and high-performance computing stacks, delivering efficient runtime and serving... 
    Suggested
    Contract work

    MatX Inc.

    Mountain View, CA
    1 day ago
  • NVIDIA in Santa Clara, CA is seeking outstanding AI systems engineers to advance the inference software stack. You will design and optimize kernels, build new abstractions...  ...engines, and contribute to accelerators and runtimes that power large language models and AI workloads.... 
    Suggested

    NVIDIA

    Santa Clara, CA
    3 days ago
  •  ...elite, research-backed AI infrastructure startup...  ...exact gap by applying deep systems programming, software-...  ...and implement full-stack ML systems for multi-agent...  ...applied directly to runtime AI operations. What...  ...production experience engineering ML systems, OR a PhD from... 
    Suggested
    Shift work

    Success Matcher Recruitment

    Sunnyvale, CA
    21 days ago
  • $200k - $300k

     .... Machine Learning Systems Engineer Location - Palo Alto...  ...Learning, Generative AI, AI Infrastructure,...  ...large language model inference. The company has pioneered...  ...vLLM, TensorRT, ONNX Runtime, and SGLang Develop...  ...the ML infrastructure stack Operate with high... 
    Suggested
    H1b
    Work at office
    Remote work

    Recruiting from Scratch

    Palo Alto, CA
    3 days ago
  • $166k - $244k

     ...centers process plastics using AI and robotics, to make...  ...are looking for a Software Engineer, Edge Systems & Runtime to join our team, focusing...  ...improving data streaming between host and device, and tuning...  ...multi-sensor inputs into model inference loops.Collaborate with... 
    Suggested
    Full time

    X Company

    Mountain View, CA
    4 days ago
  • Senior AI Systems Performance Engineer Palo Alto, California, United States The era...  ...Suite™ is the first full‑stack, generative AI platform, from...  ...across compiler, runtime, and hardware layers to deliver...  ...performance for large‑scale AI inference. Responsibilities Bring... 
    Full time
    Temporary work
    Local area
    Flexible hours

    SambaNova

    Palo Alto, CA
    2 days ago
  • NVIDIA seeks a Senior Systems Software Engineer to tackle client-side AI challenges on Windows and Linux PCs with limited resources. You will collaborate...  ..., while optimizing AI models, data pipelines, and inference runtimes for performance on next-generation GPUs. The role... 
    Local area

    NVIDIA

    Santa Clara, CA
    4 days ago
  • About the Role A1 is building a proactive AI chat app for everyday users to bring...  ...context, and real-world task completion. The system must handle multi-step reasoning,...  ...model behavior. We are looking for a Full Stack Engineer - AI Systems to build the product layer... 
    Immediate start

    Bjak

    Palo Alto, CA
    1 day ago
  • $148.7k - $201.2k

     ...backbone of Generative AI at AWS? Do you want...  ...for AI training and inference, delivering continuous...  ...us. We are seeking a Systems Development Engineer to develop automation...  ...Troubleshoot Linux boot and runtime failures across x86...  .../GPU software stacks in Linux-based environments... 
    Internship
    Local area
    Worldwide
    Flexible hours

    Amazon

    Cupertino, CA
    1 day ago
  • xAI is seeking an engineer for the RL infrastructure team to help with low precision RL training and inference. You will design and optimize the inference stack for RL workloads, profile performance, and...  ...large-scale distributed systems and proficiency in Python, C++... 

    xAI

    Palo Alto, CA
    2 days ago
  • $184k - $287.5k

    We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme...  ...architect and implement high-performance inference stacks, optimize GPU kernels and compilers, drive industry... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...d-Matrix, headquartered in Santa Clara, CA, seeks a Principal System Software Engineer for AI Inference Execution. You will join the software team to productize the AI compute engine's SW stack, developing deployment software and collaborating with ML, compiler, and hardware... 

    Jobleads-US

    Santa Clara, CA
    5 days ago
  • $218.8k - $335.3k

     ...Motors, our Embodied AI teams are...  ...learning to build systems that are both intelligent...  ..., and controls stack that keeps the vehicle...  ...a Staff Software Engineer to provide...  ..., and predictable runtime behavior under tight...  ...accelerator‑based ML inference, model deployment,... 
    Full time
    Local area
    Remote work
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    1 day ago
  • $198k - $326k

     ...team. LinkedIn’s AI Infrastructure organization...  ...Staff Software Engineer with deep...  ...the intersection of systems, machine learning,...  ..., and large-scale inference. This is a highly...  ...models interact with runtimes, compilers, and hardware...  ...across the full stack, including model... 
    For contractors
    Work at office
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  •  ...computing, cloud, and AI. Whether you’re designing...  ...is seeking an AI Systems Engineer to help develop and optimize...  ...industry-leading AI inference performance across AMD...  ...with compiler, runtime, silicon, and architecture...  ...This role offers full-stack visibility from kernel... 
    Worldwide

    AMD

    San Jose, CA
    13 hours ago
  • $192k - $278k

    Partner with engineering leads to define the strategy...  ..., delivering host networking libraries...  ...for training and inference (e.g., NCCL, NIXL...  ...compute, TPU and GPU systems, aligning with...  ...networking software stacks.Define success...  ...next generation of AI training and inference... 
    Worldwide

    Google

    Sunnyvale, CA
    3 days ago
  •  ...data centers, delivering low-latency, high-throughput AI for multi-node GPU workloads. As a Senior Engineer, you will shape core infrastructure and...  ...decisions, lead performance optimizations, and own the inference engine to scale research and production workloads.... 

    Sanas

    Palo Alto, CA
    1 day ago
  •  ...potential of generative AI to power the...  ...role: Senior toStaff Runtime Systems EngineerWhat You Will...  ...memory compute for AI inference in datacenters.This position...  ...for runtime software engineering, working on the architecture...  ...systems software that hosts this SoC.In this role,... 
    3 days per week

    d-Matrix

    Santa Clara, CA
    4 days ago
  • A leader in AI technology in Palo Alto is seeking a Senior AI Systems Performance Engineer to optimize the latest foundation models on their innovative platform. This role involves collaborating with cross-functional teams to push the performance limits of AI systems.... 

    SambaNova

    Palo Alto, CA
    2 days ago
  • $90.1k - $191.8k

     ...understand the world!The Data Labeling Engineering team designs, builds, and...  ...features.We own a modern full‑stack architecture including...  ...leadership, and work directly on systems that unblock the next generation...  ..., and latency goals.Champion AI‑assisted engineeringUse and advocate... 
    Full time
    Work experience placement
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $174.9k - $261.3k

     ...understand the world!The Data Labeling Engineering team designs, builds, and...  ..., data engineering, and AI/ML, defining the strategies, tooling...  ....We own a modern full‑stack architecture including TypeScript...  ...leadership, and direct impact on systems that unblock the next... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    3 days ago
  • $112k - $184k

     ...a test course for mobility; and Cloud & AI, the digital infrastructure powering our...  ...are seeking a versatile and self-driven Systems Engineer to lead the architecture of our enterprise...  ...unique role, you will act as a full-stack owner for our Moveworks ServiceNow Platform... 
    Temporary work
    Work at office
    Immediate start
    Flexible hours

    Woven by Toyota

    Palo Alto, CA
    2 days ago
  • A leading AI infrastructure company in California is seeking a Member of Technical Staff — Inference to design and optimize large-scale AI inference systems. The role demands 5+ years in systems engineering and expertise in large-scale inference systems. Successful candidates... 
    Flexible hours

    RadixArk

    Palo Alto, CA
    1 day ago
  • $180k - $250k

    A leading AI infrastructure firm is seeking a TPU Systems Engineer to develop high-performance systems using JAX, XLA, and Pallas. This role involves pushing large...  ...TPU hardware and optimizing performance across the stack. Candidates should have at least 3 years of... 

    RadixArk

    Palo Alto, CA
    1 day ago
  • $174k - $252k

     ...sensors and Context Hub Runtime Environment (CHRE)...  ...with embedded operating systems.Preferred qualifications...  ...development.Google's software engineers develop the next-...  ...problems across the full-stack as we continue to push...  ...combine the best of Google AI, software, and hardware... 
    Remote work

    Google

    Mountain View, CA
    3 days ago
  • $124k - $195.5k

     ...and Deep Learning Systems. As the complexity...  ...performance for emerging AI workloads. You...  ...both training and inference pipelines....  ...exploratory tools and runtime systems to profile...  ...Science, Computer Engineering, Electrical Engineering...  ...across the stack.Ways to stand out... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ...creative, skilled, and motivated engineers to join our founding team in...  ...Develop frontend for AI model interaction Use cloud...  ...Qualifications ~5+ years full-stack development experience ~ Experience...  ...with user management systems (Auth0, Stack Auth, Firebase Auth... 
    Full time

    Dexmate

    Santa Clara, CA
    1 day ago
  • $150k - $350k

     ...verification with agentic AI workflows. Our platform...  ...AI to assist engineers in RTL design, simulation...  ...Overview We are seeking an ML Systems Engineer to optimize...  ...of large language model inference powering our agentic AI...  ...multiple layers of the stack from CUDA kernels to application... 

    ChipAgents

    San Jose, CA
    4 days ago
  • d-Matrix is seeking a Senior to Staff Runtime Systems Engineer to design and implement runtime firmware and software for our AI compute platform. You will work on multi-core SoC, drivers, and systems software, ensuring performance and reliability while collaborating with... 

    Entrada Ventures

    Santa Clara, CA
    3 days ago
  • $100k

     ...leading the industry on cutting-edge AI technology, revolutionizing performance...  ...contributors of all seniorities.As a Software Engineer on the Metal Runtime team at Tenstorrent, you’ll work on...  ...and optimize high-performance runtime systems that execute directly on the hardware,... 
    Permanent employment

    Tenstorrent

    Santa Clara, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Runtime Systems Engineer - Host Stack & AI Inference. Be the first to apply!