Principal AI Systems Architect, MoE Runtime & Memory Hierarchy
Jobleads-US
A leader in memory solutions we deliver best-in-class products. For consumers and businesses alike, we have committed ourselves to quality, performance, and customer service. It is our mission to design a broad range of memory solutions focused on high performance, exceptional reliability, and outstanding customer service. We provide our partners with top-notch services, from instant technical support and professional services to streamlined and efficient delivery processes that leverage our many years of experience. Our cutting-edge products drive higher performance, increase efficiency, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities.
About the Role
Own the end-to-end technical architecture for sparse MoE inference and model-aware memory and storage management. This person turns the strategic direction into concrete object models, APIs, data paths, cache policies, runtime integrations, performance models, and prototype specifications.
The role should bridge model architecture, runtime software, operating-system memory management, accelerator execution, storage firmware, and performance engineering. It is the central technical owner for the AI-HLC architecture.
Primary responsibilities
- Architect the model-aware residency system. Define experts, expert bundles, shared experts, KV blocks, recurrent state, activations, routing metadata, and agent state as first-class objects with explicit lifecycle, identity, placement, quality, deadline, and security attributes.
- Design HLC policies and mechanisms. Own hot, warm, cold, compressed, resident, prefetched, and evicted states; admission and eviction policies; routing-frequency telemetry; reuse prediction; prefetching; compression decisions; and recovery behavior.
- Develop quantitative workload models. Model bytes per token, expert reuse distance, cache hit rate, cold-miss bursts, decode throughput, UFS or NVMe bandwidth, memory-write bandwidth, interconnect latency, thermal limits, and cost per token.
- Own both deployment tracks. For mobile and edge, design compressed expert storage, SoC-side decode, pinned residency, heterogeneous CPU/GPU/NPU execution, and UFS-aware scheduling. For enterprise and near-data systems, evaluate host-only execution, CXL memory, computational SSDs, in-drive acceleration, multi-drive sharding, and activation-oriented host-device protocols.
- Integrate with inference runtimes. Work with low-level and production runtimes such as llama.cpp, vLLM, SGLang, ExecuTorch, LiteRT, vendor accelerator stacks, and custom runtime components. Determine what can be implemented above the operating system and what requires kernel, firmware, compiler, or hardware changes.
- Architect agentic-state storage. Define how agent sessions, reusable prefixes, KV state, checkpoints, tool results, files, embeddings, intermediate artifacts, and shared multi-agent state move across accelerator memory, DRAM, CXL, local SSD, and remote storage.
- Design the runtime-to-system interface. Specify APIs for expert requests, prefetch deadlines, routing hints, cache state, quality constraints, memory pressure, compression eligibility, placement, cancellation, and telemetry.
- Lead performance and correctness reviews. Ensure that optimizations improve p95 and p99 latency, energy, and total system cost without silently changing routing behavior, model quality, security isolation, or data correctness.
- Guide implementation. Review C++, Python, kernel, firmware, and simulator designs; mentor the senior engineer; and translate prototype findings into production architecture.
- 15+ years in systems software, storage software, high-performance computing, operating systems, ML infrastructure, embedded systems, or related engineering.
- BS/MS degree in Computer Science or equivalent experience
- At least 5 years in AI/ML infrastructure, transformer inference, model serving, or accelerator software.
- Demonstrated experience architecting complex systems across multiple software or hardware layers.
- Strong understanding of transformer execution, MoE routing, expert FFNs, KV caches, quantization, model formats, batching, scheduling, and inference performance.
- Strong C++ and Python skills, with hands-on experience debugging concurrency, memory, I/O, and performance problems.
- Deep Linux systems knowledge, including virtual memory, mmap, page cache, faults, NUMA, DMA, asynchronous I/O, process telemetry, and memory-pressure behavior.
- Experience with at least one major storage interface or stack such as NVMe, UFS, PCIe, block I/O, SSD firmware, object storage, or distributed caching.
- Ability to translate architecture into testable requirements, interface specifications, and implementation milestones.
Preferred qualifications
- Android or embedded Linux experience, including Perfetto, cgroups, DMA buffers, vendor delegates, or mobile thermal and power analysis.
- Experience with NPU, DSP, GPU, FPGA, or custom accelerator programming.
- Familiarity with entropy coding, model compression, INT4 or mixed-precision inference, tensor packing, and accelerator-native weight layouts.
- Experience with CXL, computational storage, SmartSSD-class platforms, remote memory, RDMA, GPU-direct I/O, or disaggregated inference.
- Familiarity with agent frameworks, context engineering, MCP/A2A integration, workflow persistence, and multi-agent execution.
- Experience designing secure multi-tenant protocols or executing AI workloads in trusted hardware and firmware environments.
About Lexar Enterprise
Building upon the foundation and the credibility of the well-established Lexar brand, Lexar Enterprise designs and manufactures memory and storage solutions that include embedded storage, mobile memory, solid-state drives, and memory modules in commercial, industrial, and automotive grades. Our cutting-edge products are deployed globally to drive higher performance, increase efficiencies, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities
Lexar Enterprise. 1737 N First Street, Suite 680, San Jose, CA USA
Unsolicited resumes sent to Lexar from third-party recruiters and placement agencies do not:
(a) constitute any type of relationship with Lexar or
(b) obligate Lexar to pay fees should we elect to hire from those resumes. Please do not contact or present candidates directly to any Lexar personnel.act or present candidates directly to any Lexar personnel.
#J-18808-Ljbffr Jobleads-US- ...Lexar Enterprise seeks a senior systems architect to own the end-to-end technical architecture for sparse MoE inference and memory management. You will convert strategic direction... ...policies, and performance models across runtimes and hardware. You will bridge model architecture...Suggested
- ...Axiado is an AI-enhanced security processor company redefining... ...management of every digital system. The company was founded in... ...Summary As a Lead System Architect at Axiado, you will define... ...chip interconnects (AXI/AHB), memory hierarchies, and standard I/O interfaces...PrincipalFull time
$152k - $241.5k
NVIDIA is seeking a motivated architect to work with a team in... ...silicon processes. This GPU memory architecture team creates new... ..., autonomous vehicles, AI, gaming, mobile systems. What you will be doing: Developing... ...interconnects, QoS. Memory hierarchy, Memory model & ordering,...SuggestedFull timeWork experience placement- ...innovations in Flash and advanced memory technologies, our solutions... ...moving forward. Job Description A System Architect focusing on advanced memory technology and AI Inference solution... ...integration of memory subsystems, cache hierarchies, and system interconnects Advanced...PrincipalTemporary workRemote workFlexible hoursShift work
- ...Principal AI Systems Architect – Agentic AI & Engineering Automation Location: San Jose, CA (remote is okay) A leading global technology company... ...environments. Define architectures for context, memory, retrieval, knowledge access, provenance, and human-in-the...PrincipalRemote work
- ...Summary Celestica is seeking a Principal Hardware Systems Architect for Compute Platforms to provide senior... ...architecture and development of next-generation AI, accelerated computing, and high-... ...CPUs, GPUs and AI accelerators, memory, storage, PCIe/CXL, scale-up...PrincipalLocal area
$200k - $300k
...Job Title: Principal ASIC Architect - Memory Systems & AI InterconnectsJob Location: Santa Clara, CA or Boston, MACompensation: $200K - $300K plus bonus... ...coherent fabrics, memory pooling, and tiered memory hierarchies.Lead specification and micro-architecture of...Principal$220k - $280k
...Credo is looking for a Principal AI System Architect to join our team in San Jose, CA, reporting to AVP, XPU system AI interface. This role needs... ...PCIe, Ethernet, RDMA). Responsibilities Define the memory and interconnection architecture linking XPUs to memory and...PrincipalWork experience placement- Sandisk is seeking a System Architect to drive system-level solutions for advanced memory technologies and AI inference infrastructure in Milpitas, CA. You will optimize compute-memory bottlenecks, latency, and power while mapping workloads to hardware across chiplet,...
$220k - $275k
...Job Description Job Description Principal System Architect - AI, SOCs At the forefront of the latest semiconductor technology & security... ...environment. ~ Deep SoC knowledge: interconnects, memory hierarchies, and standard interfaces (PCIe, Ethernet, DRAM, USB)....PrincipalOngoing contractLocal areaWork visa- Principal / Lead DSP & Systems Architect Omni Design Technologies provides high-performance, ultra-low power IP solutions... ...SoC architectures for AI/ML accelerators, hyperscale datacenter... ...on micro-architecture, latency, and memory trade-offs. Silicon Bring-up & Validation...Principal
$272k - $431.25k
...serving generative AI and reasoning... ...feel like a single system at datacenter scale... ...rapidly outgrow the memory and compute budget... ....We are seeking a Principal Systems Engineer to... ...LLM inference.Architect and implement deep... ...understanding of memory hierarchies (GPU HBM, host...PrincipalFull timeLocal areaRemote work- ...Micron Technology, Inc in the United States is seeking a Lead Engineer, Applied AI and Agentic Systems to own the architecture and drive delivery of AI systems used by semiconductor engineers daily. You will lead designs, review architectures, and guide software engineers...Principal
$272k - $431.25k
...NVIDIA’s Networking Systems & Software Architecture... ...is solving some of AI’s hardest infrastructure... .... This Principal Architect role leads the research... ...bottlenecks in communication runtimes for large-scale AI... ...computer architecture, memory hierarchies, DMA engines, and OS-...Principal- ...experiences—from AI and data centers,... ...gaming and embedded systems. Grounded in a culture... ...are seeking a Principal GenAI Inference Optimization... ...GPU architecture, memory systems, and... ...—from kernels and runtimes to frameworks and... ...(compute, memory hierarchy, interconnects...Principal
$152k - $241.5k
...candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture... ...6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to...Full time- ...You Are You are PCIe/CXL System Architecture Specialist who can work directly with customer architects to close real system-level design... ...partner for next-generation AI, HPC, storage, networking, and... ...coherency, virtualization, RAS, memory expansion, or CXL-based...Principal
$141k - $319k
...Technology is a world leader in innovating memory and storage solutions that accelerate the... ...and advance faster than ever.The Principal Product Marketing Manager for the Cloud Memory... ...the value of Micron's memory portfolio for AI, cloud, and data center markets. This...PrincipalFull timeLocal areaImmediate start- ...Superintelligence Cloud, is a leader in AI cloud infrastructure serving... ...workloads.Develop compute system standards and design... ...experience (7+ years) architecting large-scale 10k-100k+ GPU HPC... ...knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies....Work at officeLocal areaWork from homeFlexible hours
$220.92k - $361.48k
...serves as the authoritative systems architect owning the end-to-end hierarchical... ...co-optimization of advanced AI models with physical hardware... ...power, thermal, latency, and memory-bandwidth constraints.Key... ...Define quantization, sparsity/MoE, context lengths, and reference...Full timeLocal areaImmediate startRemote workShift work- ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...seeking a Robotics AI Architect to define and scale... ...roadmap for our AI SDKs, runtime, and reference architectures... ...models Memory and dataflow efficiency...Principal
$323k
...are seeking an outstanding and visionary Lead Architect to drive the definition and exploration of next-generation system architectures and technologies that craft the... ...innovative solutions—from 3D integration and advance memory systems to emerging system-level paradigms....PrincipalWork at officeLocal areaRemote work- ...Micron Technology is seeking a Senior / Principal Layout Engineer to drive pathfinding and architecture exploration for next-generation memory modules and form factors. This role sits in the Module Architecture Group, using layout-driven insights to influence design decisions...Principal
- ...powering the future of physical AI. Founded in 2017 and now... ...and infrastructure, operating systems, and autonomy. Eighteen of the... ...on production-grade embedded runtime environments. You'll work across... ...quantization, and support deployment on memory constrained platforms...Full timeFor contractorsFor subcontractorCasual workWork at officeRemote workDay shift
$216k - $345k
We are seeking a Staff/Principal Solutions Architect to define and lead the technical strategy for semiconductor... ..., and debug-abilityDrive adoption of AI/ML techniques for yield learning,... ...0, Teradyne UltraFLEX/ETS, or similar systems.Strong background in DFT architectures...PrincipalFull time$272k - $431.25k
...and hands-on delivery across system software, drivers, and CUDA to... ...in C/C++, including IPC/shared memory, and bounded CPU/memory budgets... ...integrate with existing ML/AI workflows (e.g., PyTorch/XLA)... ...and GPU architecture, including runtime/driver APIs, CUDA streams/graphs...PrincipalFull time$210k - $250k
...business leaders to join us.Job Summary:The Principal Solution Architect leads a virtual team of product engineers, design engineers, system engineers and FAEs to develop best in... ...integrated systems that scale into large scale AI/HPC clusters, Scale-out Storage clusters,...PrincipalWork experience placementWorldwide- ...accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and... ...and IP architecture with a focus on SerDes, fabrics, memory subsystems, I/O, clocking, or securityExperience analyzing...
$190k - $245k
...flight. The company works across air taxis, UAS, AI and powertrain development — building real-world systems that combine software intelligence with... ...supports and celebrates all of our team members.The Principal Solution Architect – Supply Chain Systems will lead the end-to-...PrincipalLocal areaVisa sponsorshipNight shift$184k - $287.5k
...GPUs excel.We are seeking a GPU System Architect who will architect and design... ...datacenter platforms for AI and HPC. The architect in this... ...GPU compute, high-bandwidth memory, in-package interconnects and... ...software interaction, drivers and runtimes, and performance tuning for...Full timeWorldwide
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI Systems Architect, MoE Runtime & Memory Hierarchy. Be the first to apply!
- embedded systems architect San Jose, CA
- system architect San Jose, CA
- pega system architect San Jose, CA
- technical architect San Jose, CA
- senior principal cloud computing engineer San Jose, CA
- principal cloud computing engineer San Jose, CA
- senior principal scientist San Jose, CA
- principal architect San Jose, CA
- principal San Jose, CA
- embedded systems architect


