Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal AI Systems Architect, MoE Runtime & Memory Hierarchy

Jobleads-US

A leader in memory solutions we deliver best-in-class products. For consumers and businesses alike, we have committed ourselves to quality, performance, and customer service. It is our mission to design a broad range of memory solutions focused on high performance, exceptional reliability, and outstanding customer service. We provide our partners with top-notch services, from instant technical support and professional services to streamlined and efficient delivery processes that leverage our many years of experience. Our cutting-edge products drive higher performance, increase efficiency, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities.

About the Role

Own the end-to-end technical architecture for sparse MoE inference and model-aware memory and storage management. This person turns the strategic direction into concrete object models, APIs, data paths, cache policies, runtime integrations, performance models, and prototype specifications.

The role should bridge model architecture, runtime software, operating-system memory management, accelerator execution, storage firmware, and performance engineering. It is the central technical owner for the AI-HLC architecture.

Primary responsibilities

  • Architect the model-aware residency system. Define experts, expert bundles, shared experts, KV blocks, recurrent state, activations, routing metadata, and agent state as first-class objects with explicit lifecycle, identity, placement, quality, deadline, and security attributes.
  • Design HLC policies and mechanisms. Own hot, warm, cold, compressed, resident, prefetched, and evicted states; admission and eviction policies; routing-frequency telemetry; reuse prediction; prefetching; compression decisions; and recovery behavior.
  • Develop quantitative workload models. Model bytes per token, expert reuse distance, cache hit rate, cold-miss bursts, decode throughput, UFS or NVMe bandwidth, memory-write bandwidth, interconnect latency, thermal limits, and cost per token.
  • Own both deployment tracks. For mobile and edge, design compressed expert storage, SoC-side decode, pinned residency, heterogeneous CPU/GPU/NPU execution, and UFS-aware scheduling. For enterprise and near-data systems, evaluate host-only execution, CXL memory, computational SSDs, in-drive acceleration, multi-drive sharding, and activation-oriented host-device protocols.
  • Integrate with inference runtimes. Work with low-level and production runtimes such as llama.cpp, vLLM, SGLang, ExecuTorch, LiteRT, vendor accelerator stacks, and custom runtime components. Determine what can be implemented above the operating system and what requires kernel, firmware, compiler, or hardware changes.
  • Architect agentic-state storage. Define how agent sessions, reusable prefixes, KV state, checkpoints, tool results, files, embeddings, intermediate artifacts, and shared multi-agent state move across accelerator memory, DRAM, CXL, local SSD, and remote storage.
  • Design the runtime-to-system interface. Specify APIs for expert requests, prefetch deadlines, routing hints, cache state, quality constraints, memory pressure, compression eligibility, placement, cancellation, and telemetry.
  • Lead performance and correctness reviews. Ensure that optimizations improve p95 and p99 latency, energy, and total system cost without silently changing routing behavior, model quality, security isolation, or data correctness.
  • Guide implementation. Review C++, Python, kernel, firmware, and simulator designs; mentor the senior engineer; and translate prototype findings into production architecture.
  • 15+ years in systems software, storage software, high-performance computing, operating systems, ML infrastructure, embedded systems, or related engineering.
  • BS/MS degree in Computer Science or equivalent experience
  • At least 5 years in AI/ML infrastructure, transformer inference, model serving, or accelerator software.
  • Demonstrated experience architecting complex systems across multiple software or hardware layers.
  • Strong understanding of transformer execution, MoE routing, expert FFNs, KV caches, quantization, model formats, batching, scheduling, and inference performance.
  • Strong C++ and Python skills, with hands-on experience debugging concurrency, memory, I/O, and performance problems.
  • Deep Linux systems knowledge, including virtual memory, mmap, page cache, faults, NUMA, DMA, asynchronous I/O, process telemetry, and memory-pressure behavior.
  • Experience with at least one major storage interface or stack such as NVMe, UFS, PCIe, block I/O, SSD firmware, object storage, or distributed caching.
  • Ability to translate architecture into testable requirements, interface specifications, and implementation milestones.

Preferred qualifications

  • Android or embedded Linux experience, including Perfetto, cgroups, DMA buffers, vendor delegates, or mobile thermal and power analysis.
  • Experience with NPU, DSP, GPU, FPGA, or custom accelerator programming.
  • Familiarity with entropy coding, model compression, INT4 or mixed-precision inference, tensor packing, and accelerator-native weight layouts.
  • Experience with CXL, computational storage, SmartSSD-class platforms, remote memory, RDMA, GPU-direct I/O, or disaggregated inference.
  • Familiarity with agent frameworks, context engineering, MCP/A2A integration, workflow persistence, and multi-agent execution.
  • Experience designing secure multi-tenant protocols or executing AI workloads in trusted hardware and firmware environments.

About Lexar Enterprise

Building upon the foundation and the credibility of the well-established Lexar brand, Lexar Enterprise designs and manufactures memory and storage solutions that include embedded storage, mobile memory, solid-state drives, and memory modules in commercial, industrial, and automotive grades. Our cutting-edge products are deployed globally to drive higher performance, increase efficiencies, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities

Lexar Enterprise. 1737 N First Street, Suite 680, San Jose, CA USA

Unsolicited resumes sent to Lexar from third-party recruiters and placement agencies do not:

(a) constitute any type of relationship with Lexar or

(b) obligate Lexar to pay fees should we elect to hire from those resumes. Please do not contact or present candidates directly to any Lexar personnel.act or present candidates directly to any Lexar personnel.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 6 days ago
Similar jobs that could be interesting for youBased on the Principal AI Systems Architect, MoE Runtime & Memory Hierarchy in San Jose, CA vacancy
  •  ...Lexar Enterprise seeks a senior systems architect to own the end-to-end technical architecture for sparse MoE inference and memory management. You will convert strategic direction...  ...policies, and performance models across runtimes and hardware. You will bridge model architecture... 
    Suggested

    Jobleads-US

    San Jose, CA
    6 days ago
  •  ...Axiado is an AI-enhanced security processor company redefining...  ...management of every digital system. The company was founded in...  ...Summary  As a Lead System Architect at Axiado, you will define...  ...chip interconnects (AXI/AHB), memory hierarchies, and standard I/O interfaces... 
    Principal
    Full time

    Axiado

    San Jose, CA
    4 days ago
  • $152k - $241.5k

    NVIDIA is seeking a motivated architect to work with a team in...  ...silicon processes. This GPU memory architecture team creates new...  ..., autonomous vehicles, AI, gaming, mobile systems. What you will be doing: Developing...  ...interconnects, QoS. Memory hierarchy, Memory model & ordering,... 
    Suggested
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    a month ago
  •  ...innovations in Flash and advanced memory technologies, our solutions...  ...moving forward. Job Description A System Architect focusing on advanced memory technology and AI Inference solution...  ...integration of memory subsystems, cache hierarchies, and system interconnects Advanced... 
    Principal
    Temporary work
    Remote work
    Flexible hours
    Shift work

    Sandisk

    Milpitas, CA
    3 days ago
  •  ...Principal AI Systems Architect – Agentic AI & Engineering Automation Location: San Jose, CA (remote is okay) A leading global technology company...  ...environments. Define architectures for context, memory, retrieval, knowledge access, provenance, and human-in-the... 
    Principal
    Remote work

    Jobleads-US

    San Jose, CA
    4 days ago
  •  ...Summary Celestica is seeking a Principal Hardware Systems Architect for Compute Platforms to provide senior...  ...architecture and development of next-generation AI, accelerated computing, and high-...  ...CPUs, GPUs and AI accelerators, memory, storage, PCIe/CXL, scale-up... 
    Principal
    Local area

    Celestica International LP

    San Jose, CA
    19 days ago
  • $200k - $300k

     ...Job Title: Principal ASIC Architect - Memory Systems & AI InterconnectsJob Location: Santa Clara, CA or Boston, MACompensation: $200K - $300K plus bonus...  ...coherent fabrics, memory pooling, and tiered memory hierarchies.Lead specification and micro-architecture of... 
    Principal

    CyberCoders

    Santa Clara, CA
    29 days ago
  • $220k - $280k

     ...About the role Credo is looking for a Principal AI System Architect to join our team in San Jose, CA, reporting to AVP, XPU system AI interface...  ...PCIe, Ethernet, RDMA). Responsibilities Define the memory and interconnection architecture linking XPUs to memory and... 
    Principal
    Work experience placement

    Jobleads-US

    San Jose, CA
    6 days ago
  • Sandisk is seeking a System Architect to drive system-level solutions for advanced memory technologies and AI inference infrastructure in Milpitas, CA. You will optimize compute-memory bottlenecks, latency, and power while mapping workloads to hardware across chiplet,... 

    Sandisk

    Milpitas, CA
    3 days ago
  • $220k - $275k

     ...Job Description Job Description Principal System Architect - AI, SOCs At the forefront of the latest semiconductor technology & security...  ...environment.  ~ Deep SoC knowledge: interconnects, memory hierarchies, and standard interfaces (PCIe, Ethernet, DRAM, USB).... 
    Principal
    Ongoing contract
    Local area
    Work visa

    CyberCoders

    San Jose, CA
    2 days ago
  • Principal / Lead DSP & Systems Architect Omni Design Technologies provides high-performance, ultra-low power IP solutions...  ...SoC architectures for AI/ML accelerators, hyperscale datacenter...  ...on micro-architecture, latency, and memory trade-offs. Silicon Bring-up & Validation... 
    Principal

    Omni Design Technologies Inc

    Milpitas, CA
    5 days ago
  • $272k - $431.25k

     ...serving generative AI and reasoning...  ...feel like a single system at datacenter scale...  ...rapidly outgrow the memory and compute budget...  ....We are seeking a Principal Systems Engineer to...  ...LLM inference.Architect and implement deep...  ...understanding of memory hierarchies (GPU HBM, host... 
    Principal
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    a month ago
  • $272k - $431.25k

     ...NVIDIA’s Networking Systems & Software Architecture...  ...is solving some of AI’s hardest infrastructure...  .... This Principal Architect role leads the research...  ...bottlenecks in communication runtimes for large-scale AI...  ...computer architecture, memory hierarchies, DMA engines, and OS-... 
    Principal

    NVIDIA Gruppe

    Santa Clara, CA
    3 days ago
  •  ...experiences—from AI and data centers,...  ...gaming and embedded systems. Grounded in a culture...  ...are seeking a Principal GenAI Inference Optimization...  ...GPU architecture, memory systems, and...  ...—from kernels and runtimes to frameworks and...  ...(compute, memory hierarchy, interconnects... 
    Principal

    AMD

    San Jose, CA
    a month ago
  • $152k - $241.5k

     ...candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture...  ...6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...You Are You are PCIe/CXL System Architecture Specialist who can work directly with customer architects to close real system-level design...  ...partner for next-generation AI, HPC, storage, networking, and...  ...coherency, virtualization, RAS, memory expansion, or CXL-based... 
    Principal

    Synopsys Inc

    Sunnyvale, CA
    a month ago
  • $141k - $319k

     ...Technology is a world leader in innovating memory and storage solutions that accelerate the...  ...and advance faster than ever.The Principal Product Marketing Manager for the Cloud Memory...  ...the value of Micron's memory portfolio for AI, cloud, and data center markets. This... 
    Principal
    Full time
    Local area
    Immediate start

    Micron

    San Jose, CA
    6 days ago
  • $220.92k - $361.48k

     ...serves as the authoritative systems architect owning the end-to-end hierarchical...  ...co-optimization of advanced AI models with physical hardware...  ...power, thermal, latency, and memory-bandwidth constraints.Key...  ...Define quantization, sparsity/MoE, context lengths, and reference... 
    Full time
    Local area
    Immediate start
    Remote work
    Shift work

    Intel

    Santa Clara, CA
    6 days ago
  •  ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving...  ...workloads.Develop compute system standards and design...  ...experience (7+ years) architecting large-scale 10k-100k+ GPU HPC...  ...knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies.... 
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  •  ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture...  ...seeking a Robotics AI Architect to define and scale...  ...roadmap for our AI SDKs, runtime, and reference architectures...  ...models Memory and dataflow efficiency... 
    Principal

    Jobleads-US

    San Jose, CA
    3 days ago
  • $323k

     ...are seeking an outstanding and visionary Lead Architect to drive the definition and exploration of next-generation system architectures and technologies that craft the...  ...innovative solutions—from 3D integration and advance memory systems to emerging system-level paradigms.... 
    Principal
    Work at office
    Local area
    Remote work

    Jobleads-US

    San Jose, CA
    5 days ago
  •  ...Micron Technology is seeking a Senior / Principal Layout Engineer to drive pathfinding and architecture exploration for next-generation memory modules and form factors. This role sits in the Module Architecture Group, using layout-driven insights to influence design decisions... 
    Principal

    Jobleads-US

    San Jose, CA
    3 days ago
  • $159.05k - $199.3k

     ...powering the future of physical AI. Founded in 2017 and now...  ...and infrastructure, operating systems, and autonomy. Eighteen of the...  ...on production-grade embedded runtime environments. You’ll work across...  ...quantization, and support deployment on memory constrained platforms... 
    Full time
    For contractors
    For subcontractor
    Casual work
    Work at office
    Remote work
    Day shift

    Applied Intuition

    Sunnyvale, CA
    4 days ago
  • $216k - $345k

    We are seeking a Staff/Principal Solutions Architect to define and lead the technical strategy for semiconductor...  ..., and debug-abilityDrive adoption of AI/ML techniques for yield learning,...  ...0, Teradyne UltraFLEX/ETS, or similar systems.Strong background in DFT architectures... 
    Principal
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $272k - $431.25k

     ...and hands-on delivery across system software, drivers, and CUDA to...  ...in C/C++, including IPC/shared memory, and bounded CPU/memory budgets...  ...integrate with existing ML/AI workflows (e.g., PyTorch/XLA)...  ...and GPU architecture, including runtime/driver APIs, CUDA streams/graphs... 
    Principal
    Full time

    Nvidia

    Santa Clara, CA
    a month ago
  • $210k - $250k

     ...business leaders to join us.Job Summary:The Principal Solution Architect leads a virtual team of product engineers, design engineers, system engineers and FAEs to develop best in...  ...integrated systems that scale into large scale AI/HPC clusters, Scale-out Storage clusters,... 
    Principal
    Work experience placement
    Worldwide

    Super Micro Computer

    San Jose, CA
    10 days ago
  •  ...accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and...  ...and IP architecture with a focus on SerDes, fabrics, memory subsystems, I/O, clocking, or securityExperience analyzing... 

    AMD

    San Jose, CA
    a month ago
  • $190k - $245k

     ...flight. The company works across air taxis, UAS, AI and powertrain development — building real-world systems that combine software intelligence with...  ...supports and celebrates all of our team members.The Principal Solution Architect – Supply Chain Systems will lead the end-to-... 
    Principal
    Local area
    Visa sponsorship
    Night shift

    Archer Aviation

    San Jose, CA
    a month ago
  • $184k - $287.5k

     ...GPUs excel.We are seeking a GPU System Architect who will architect and design...  ...datacenter platforms for AI and HPC. The architect in this...  ...GPU compute, high-bandwidth memory, in-package interconnects and...  ...software interaction, drivers and runtimes, and performance tuning for... 
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    6 hours ago
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System...  ...libraries, drivers, and runtime components that workloads...  ...everything above it, we treat memory safety, explicit ownership...  ...how we engineer. We treat AI coding agents as a primary... 
    Principal
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal AI Systems Architect, MoE Runtime & Memory Hierarchy. Be the first to apply!