Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal AI Systems Architect, MoE Runtime & Memory Hierarchy

Jobleads-US

A leader in memory solutions we deliver best-in-class products. For consumers and businesses alike, we have committed ourselves to quality, performance, and customer service. It is our mission to design a broad range of memory solutions focused on high performance, exceptional reliability, and outstanding customer service. We provide our partners with top-notch services, from instant technical support and professional services to streamlined and efficient delivery processes that leverage our many years of experience. Our cutting-edge products drive higher performance, increase efficiency, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities.

About the Role

Own the end-to-end technical architecture for sparse MoE inference and model-aware memory and storage management. This person turns the strategic direction into concrete object models, APIs, data paths, cache policies, runtime integrations, performance models, and prototype specifications.

The role should bridge model architecture, runtime software, operating-system memory management, accelerator execution, storage firmware, and performance engineering. It is the central technical owner for the AI-HLC architecture.

Primary responsibilities

  • Architect the model-aware residency system. Define experts, expert bundles, shared experts, KV blocks, recurrent state, activations, routing metadata, and agent state as first-class objects with explicit lifecycle, identity, placement, quality, deadline, and security attributes.
  • Design HLC policies and mechanisms. Own hot, warm, cold, compressed, resident, prefetched, and evicted states; admission and eviction policies; routing-frequency telemetry; reuse prediction; prefetching; compression decisions; and recovery behavior.
  • Develop quantitative workload models. Model bytes per token, expert reuse distance, cache hit rate, cold-miss bursts, decode throughput, UFS or NVMe bandwidth, memory-write bandwidth, interconnect latency, thermal limits, and cost per token.
  • Own both deployment tracks. For mobile and edge, design compressed expert storage, SoC-side decode, pinned residency, heterogeneous CPU/GPU/NPU execution, and UFS-aware scheduling. For enterprise and near-data systems, evaluate host-only execution, CXL memory, computational SSDs, in-drive acceleration, multi-drive sharding, and activation-oriented host-device protocols.
  • Integrate with inference runtimes. Work with low-level and production runtimes such as llama.cpp, vLLM, SGLang, ExecuTorch, LiteRT, vendor accelerator stacks, and custom runtime components. Determine what can be implemented above the operating system and what requires kernel, firmware, compiler, or hardware changes.
  • Architect agentic-state storage. Define how agent sessions, reusable prefixes, KV state, checkpoints, tool results, files, embeddings, intermediate artifacts, and shared multi-agent state move across accelerator memory, DRAM, CXL, local SSD, and remote storage.
  • Design the runtime-to-system interface. Specify APIs for expert requests, prefetch deadlines, routing hints, cache state, quality constraints, memory pressure, compression eligibility, placement, cancellation, and telemetry.
  • Lead performance and correctness reviews. Ensure that optimizations improve p95 and p99 latency, energy, and total system cost without silently changing routing behavior, model quality, security isolation, or data correctness.
  • Guide implementation. Review C++, Python, kernel, firmware, and simulator designs; mentor the senior engineer; and translate prototype findings into production architecture.
  • 15+ years in systems software, storage software, high-performance computing, operating systems, ML infrastructure, embedded systems, or related engineering.
  • BS/MS degree in Computer Science or equivalent experience
  • At least 5 years in AI/ML infrastructure, transformer inference, model serving, or accelerator software.
  • Demonstrated experience architecting complex systems across multiple software or hardware layers.
  • Strong understanding of transformer execution, MoE routing, expert FFNs, KV caches, quantization, model formats, batching, scheduling, and inference performance.
  • Strong C++ and Python skills, with hands-on experience debugging concurrency, memory, I/O, and performance problems.
  • Deep Linux systems knowledge, including virtual memory, mmap, page cache, faults, NUMA, DMA, asynchronous I/O, process telemetry, and memory-pressure behavior.
  • Experience with at least one major storage interface or stack such as NVMe, UFS, PCIe, block I/O, SSD firmware, object storage, or distributed caching.
  • Ability to translate architecture into testable requirements, interface specifications, and implementation milestones.

Preferred qualifications

  • Android or embedded Linux experience, including Perfetto, cgroups, DMA buffers, vendor delegates, or mobile thermal and power analysis.
  • Experience with NPU, DSP, GPU, FPGA, or custom accelerator programming.
  • Familiarity with entropy coding, model compression, INT4 or mixed-precision inference, tensor packing, and accelerator-native weight layouts.
  • Experience with CXL, computational storage, SmartSSD-class platforms, remote memory, RDMA, GPU-direct I/O, or disaggregated inference.
  • Familiarity with agent frameworks, context engineering, MCP/A2A integration, workflow persistence, and multi-agent execution.
  • Experience designing secure multi-tenant protocols or executing AI workloads in trusted hardware and firmware environments.

About Lexar Enterprise

Building upon the foundation and the credibility of the well-established Lexar brand, Lexar Enterprise designs and manufactures memory and storage solutions that include embedded storage, mobile memory, solid-state drives, and memory modules in commercial, industrial, and automotive grades. Our cutting-edge products are deployed globally to drive higher performance, increase efficiencies, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities

Lexar Enterprise. 1737 N First Street, Suite 680, San Jose, CA USA

Unsolicited resumes sent to Lexar from third-party recruiters and placement agencies do not:

(a) constitute any type of relationship with Lexar or

(b) obligate Lexar to pay fees should we elect to hire from those resumes. Please do not contact or present candidates directly to any Lexar personnel.act or present candidates directly to any Lexar personnel.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Principal AI Systems Architect, MoE Runtime & Memory Hierarchy in San Jose, CA vacancy
  •  ...Lexar Enterprise seeks a senior systems architect to own the end-to-end technical architecture for sparse MoE inference and memory management. You will convert strategic direction...  ...policies, and performance models across runtimes and hardware. You will bridge model architecture... 
    Suggested

    Jobleads-US

    San Jose, CA
    4 hours ago
  • $272k - $431.25k

     ...is to innovate how we architect and develop our GPU for the changing AI and accelerated workloads...  .... We are looking for a Principal System Architect with a wealth...  ...SoC subsystems such as memory architecture, test...  ...fundamentals, including memory hierarchy, coherency, clocking,... 
    Principal
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $152k - $241.5k

    NVIDIA is seeking a motivated architect to work with a team in...  ...silicon processes. This GPU memory architecture team creates new...  ..., autonomous vehicles, AI, gaming, mobile systems. What you will be doing: Developing...  ...interconnects, QoS. Memory hierarchy, Memory model & ordering,... 
    Suggested
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    14 hours ago
  • Position: Principal Engineer/Architect - Generative AI & Edge Systems, Automotive Location: San Jose CA (4 days onsite) Employment...  ...(e.g., TensorRT-LLM, ONNX Runtime, GGUF/llama.cpp). Proven experience...  ...analysis across latency, memory bandwidth, and CPU vs. NPU vs. GPU... 
    Principal
    Full time
    Relocation package

    PER International

    San Jose, CA
    3 days ago
  •  ...innovations in Flash and advanced memory technologies, our solutions...  ...Job Description A System Architect focusing on advanced memory technology and AI Inference solution infrastructure...  ...integration of memory subsystems, cache hierarchies, and system interconnects... 
    Principal
    Full time
    Temporary work
    Remote work
    Flexible hours
    Shift work

    Sandisk

    Milpitas, CA
    26 days ago
  • $219k - $351k

     ...partners, and communities.Job Title: Principal Engineer, AI System Architect (Hardware)The Architecture Research...  ...in modern AI, particularly in memory capacity/bandwidth and system-scale...  ...architectures, including compute, memory hierarchies, and high-performance interconnects... 
    Principal
    Work at office
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    1 day ago
  • $180k - $240k

    Overview AI Memory Solution Architect, Principal Engineer role at SK hynix America Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite...  ...analysis to ensure high efficiency. Collaborate with system software and hardware design teams to ensure seamless integration... 
    Principal
    Full time
    Work at office
    Local area

    SK hynix America Inc.

    San Jose, CA
    4 days ago
  • $200k - $300k

    Job Title: Principal ASIC Architect - Memory Systems & AI InterconnectsJob Location: Santa Clara, CA or Boston, MACompensation: $200K - $300K plus bonus...  ...coherent fabrics, memory pooling, and tiered memory hierarchies.Lead specification and micro-architecture of... 
    Principal

    CyberCoders

    Santa Clara, CA
    14 hours ago
  • $152k - $241.5k

     ...candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture...  ...6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...serving generative AI and reasoning...  ...feel like a single system at datacenter scale...  ...rapidly outgrow the memory and compute budget...  ....We are seeking a Principal Systems Engineer to...  ...LLM inference.Architect and implement deep...  ...understanding of memory hierarchies (GPU HBM, host... 
    Principal
    Full time
    Local area
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  • $200k

     ...Role Overview We are looking for a Principal AI SoC Runtime Software Architect to own the software architecture...  ...candidate combines deep runtime and systems-software expertise with a strong...  ...driver interfaces, DMA and shared-memory architectures, and performance-sensitive... 
    Principal
    Contract work
    Flexible hours

    Velaura

    Santa Clara, CA
    4 days ago
  •  ...differentiated SoC architectures for AI/ML accelerators, hyperscale...  .... Job Description We are seeking a Principal / Lead DSP & Systems Architect to own the architectural direction of...  ...on micro-architecture, latency, and memory trade-offs. Silicon Bring-up & Validation... 
    Principal

    Omnidesigntech

    Milpitas, CA
    2 days ago
  • Micron Technology is seeking a System Architect to lead performance modeling and analysis of next‑generation memory architectures for AI, server, data center, and mobile platforms. You will build and validate cycle‑accurate models and translate results into architecture... 

    Micron Technology

    San Jose, CA
    3 days ago
  • Micron Technology in San Jose, California, seeks a System Architect to lead performance modeling and analysis for next‑generation memory architectures across AI, server, data center, and mobile platforms. You will build and validate timing-approximate and cycle-accurate... 

    Micron Memory Malaysia Sdn Bhd

    San Jose, CA
    2 days ago
  • Micron Technology is seeking a System Architect to lead performance modeling and analysis of next‑generation memory architectures for AI, server, data center, and mobile platforms. You will build timing‑approximate and cycle‑accurate models and translate results into architecture... 

    1000 Micron Technology, Inc.

    San Jose, CA
    3 days ago
  • $148k - $308k

     ...world leader in innovating memory and storage solutions that accelerate...  ...where memory, storage, and system architecture come together to shape the future of AI and datacenter platforms. We...  ...work here matters!As a Sr. Principal Solutions Architect for Network Interconnect... 
    Principal
    Full time
    Work at office
    Local area
    Immediate start

    Micron

    San Jose, CA
    2 days ago
  •  ...experiences—from AI and data centers,...  ...gaming and embedded systems. Grounded in a culture...  ...are seeking a Principal GenAI Inference Optimization...  ...GPU architecture, memory systems, and...  ...—from kernels and runtimes to frameworks and...  ...(compute, memory hierarchy, interconnects... 
    Principal

    AMD

    San Jose, CA
    14 hours ago
  • Dell Technologies in Santa Clara, CA, seeks a Senior Principal Systems Development Engineer to architect and deliver end-to-end AI rack solutions. You will own the system-level design across compute, mechanical, thermal, software, diagnostics and factory enablement, working... 
    Principal

    Dell Technologies

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...GPUs excel.We are seeking a GPU System Architect who will architect and design...  ...datacenter platforms for AI and HPC. The architect in this...  ...GPU compute, high-bandwidth memory, in-package interconnects and...  ...software interaction, drivers and runtimes, and performance tuning for... 
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    4 days ago
  • Samsung Semiconductor in San Jose is hiring a Staff Engineer, Memory Systems Architecture to advance fault management for data center memory...  ...failure modes, ECC schemes, and RAS algorithms to reduce downtime in AI/ML workloads. You will collaborate with customers, analyze... 

    Samsung Semiconductor

    San Jose, CA
    3 days ago
  • $272k - $431.25k

     ...NVIDIA’s Networking Systems & Software Architecture...  ...is solving some of AI’s hardest infrastructure...  ...interconnects.This Principal Architect role leads the research...  ...in communication runtimes for large-scale AI workloads...  ...architecture, memory hierarchies, DMA engines, and OS-... 
    Principal
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    3 days ago
  •  ...Role: Principal/Lead Design Engineer Memory Controller- (CHI-G / DFI) Location: San Jose, CA Let's create our future together at The AES Group...  ...their business operations with the power of cloud, data, AI, engineering, and other emerging technologies. Role... 
    Principal

    The AES Group

    San Jose, CA
    1 day ago
  • $146k - $286k

     ...a world leader in innovating memory and storage solutions that accelerate...  ...ever.Micron is seeking a System Architect to lead performance modeling...  ...memory architectures for AI, server, data center, and mobile...  ...of DRAM architecture, memory hierarchy design, and cache-coherent... 
    Full time
    Local area
    Immediate start

    Micron

    San Jose, CA
    14 hours ago
  • $141k - $319k

     ...Technology is a world leader in innovating memory and storage solutions that accelerate the...  ...and advance faster than ever.The Principal Product Marketing Manager for the Cloud Memory...  ...the value of Micron's memory portfolio for AI, cloud, and data center markets. This... 
    Principal
    Full time
    Local area
    Immediate start

    Micron

    San Jose, CA
    1 day ago
  • $180k - $260k

    Senior Staff Engineer/Principal Memory Design Engineer Senior Staff Engineer/Principal Memory Design...  ...Be among the first 25 applicants Get AI-powered advice on this job and more exclusive...  ...memory design engineer to define and architect memory circuits for high performance... 
    Principal
    Full time
    Temporary work
    Worldwide

    MediaTek

    San Jose, CA
    2 days ago
  • Micron Technology, Inc is seeking a Technical Solutions Sales Architect to lead early technical engagement with strategic customers and drive design wins for memory and storage solutions across AI, cloud, and data‑center platforms. You will translate workloads into differentiated... 

    Micron Technology

    San Jose, CA
    1 day ago
  • Micron Technology, Inc. in San Jose, CA is seeking a System Architect to drive exploration, modeling, and analysis of memory and storage subsystem architectures for AI, server, and mobile computing systems. You will collaborate across design, software, and engineering... 

    Micron Technology, Inc

    San Jose, CA
    3 days ago
  •  ...performance across compiler, runtime, and hardware layers to...  ..., high-performance AI applications....  ...quantization, scheduling, and memory optimization. Develop...  ...single-node and distributed systems. Requirements...  ...Knowledge of memory hierarchy optimization, caching,... 
    Principal
    Full time
    Temporary work
    Flexible hours

    SambaNova Systems

    San Jose, CA
    12 days ago
  •  ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture...  ...seeking a Robotics AI Architect to define and scale...  ...roadmap for our AI SDKs, runtime, and reference architectures...  ...models Memory and dataflow efficiency... 
    Principal

    Advanced Micro Devices

    San Jose, CA
    4 days ago
  •  ...are seeking an outstanding and visionary Lead Architect to drive the definition and exploration of next-generation system architectures and technologies that craft the...  ...innovative solutions—from 3D integration and advance memory systems to emerging system-level paradigms.... 
    Principal
    Work at office
    Local area
    Remote work

    ARM

    San Jose, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal AI Systems Architect, MoE Runtime & Memory Hierarchy. Be the first to apply!