Principal AI Systems Architect, MoE Runtime & Memory Hierarchy
Jobleads-US
A leader in memory solutions we deliver best-in-class products. For consumers and businesses alike, we have committed ourselves to quality, performance, and customer service. It is our mission to design a broad range of memory solutions focused on high performance, exceptional reliability, and outstanding customer service. We provide our partners with top-notch services, from instant technical support and professional services to streamlined and efficient delivery processes that leverage our many years of experience. Our cutting-edge products drive higher performance, increase efficiency, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities.
About the Role
Own the end-to-end technical architecture for sparse MoE inference and model-aware memory and storage management. This person turns the strategic direction into concrete object models, APIs, data paths, cache policies, runtime integrations, performance models, and prototype specifications.
The role should bridge model architecture, runtime software, operating-system memory management, accelerator execution, storage firmware, and performance engineering. It is the central technical owner for the AI-HLC architecture.
Primary responsibilities
- Architect the model-aware residency system. Define experts, expert bundles, shared experts, KV blocks, recurrent state, activations, routing metadata, and agent state as first-class objects with explicit lifecycle, identity, placement, quality, deadline, and security attributes.
- Design HLC policies and mechanisms. Own hot, warm, cold, compressed, resident, prefetched, and evicted states; admission and eviction policies; routing-frequency telemetry; reuse prediction; prefetching; compression decisions; and recovery behavior.
- Develop quantitative workload models. Model bytes per token, expert reuse distance, cache hit rate, cold-miss bursts, decode throughput, UFS or NVMe bandwidth, memory-write bandwidth, interconnect latency, thermal limits, and cost per token.
- Own both deployment tracks. For mobile and edge, design compressed expert storage, SoC-side decode, pinned residency, heterogeneous CPU/GPU/NPU execution, and UFS-aware scheduling. For enterprise and near-data systems, evaluate host-only execution, CXL memory, computational SSDs, in-drive acceleration, multi-drive sharding, and activation-oriented host-device protocols.
- Integrate with inference runtimes. Work with low-level and production runtimes such as llama.cpp, vLLM, SGLang, ExecuTorch, LiteRT, vendor accelerator stacks, and custom runtime components. Determine what can be implemented above the operating system and what requires kernel, firmware, compiler, or hardware changes.
- Architect agentic-state storage. Define how agent sessions, reusable prefixes, KV state, checkpoints, tool results, files, embeddings, intermediate artifacts, and shared multi-agent state move across accelerator memory, DRAM, CXL, local SSD, and remote storage.
- Design the runtime-to-system interface. Specify APIs for expert requests, prefetch deadlines, routing hints, cache state, quality constraints, memory pressure, compression eligibility, placement, cancellation, and telemetry.
- Lead performance and correctness reviews. Ensure that optimizations improve p95 and p99 latency, energy, and total system cost without silently changing routing behavior, model quality, security isolation, or data correctness.
- Guide implementation. Review C++, Python, kernel, firmware, and simulator designs; mentor the senior engineer; and translate prototype findings into production architecture.
- 15+ years in systems software, storage software, high-performance computing, operating systems, ML infrastructure, embedded systems, or related engineering.
- BS/MS degree in Computer Science or equivalent experience
- At least 5 years in AI/ML infrastructure, transformer inference, model serving, or accelerator software.
- Demonstrated experience architecting complex systems across multiple software or hardware layers.
- Strong understanding of transformer execution, MoE routing, expert FFNs, KV caches, quantization, model formats, batching, scheduling, and inference performance.
- Strong C++ and Python skills, with hands-on experience debugging concurrency, memory, I/O, and performance problems.
- Deep Linux systems knowledge, including virtual memory, mmap, page cache, faults, NUMA, DMA, asynchronous I/O, process telemetry, and memory-pressure behavior.
- Experience with at least one major storage interface or stack such as NVMe, UFS, PCIe, block I/O, SSD firmware, object storage, or distributed caching.
- Ability to translate architecture into testable requirements, interface specifications, and implementation milestones.
Preferred qualifications
- Android or embedded Linux experience, including Perfetto, cgroups, DMA buffers, vendor delegates, or mobile thermal and power analysis.
- Experience with NPU, DSP, GPU, FPGA, or custom accelerator programming.
- Familiarity with entropy coding, model compression, INT4 or mixed-precision inference, tensor packing, and accelerator-native weight layouts.
- Experience with CXL, computational storage, SmartSSD-class platforms, remote memory, RDMA, GPU-direct I/O, or disaggregated inference.
- Familiarity with agent frameworks, context engineering, MCP/A2A integration, workflow persistence, and multi-agent execution.
- Experience designing secure multi-tenant protocols or executing AI workloads in trusted hardware and firmware environments.
About Lexar Enterprise
Building upon the foundation and the credibility of the well-established Lexar brand, Lexar Enterprise designs and manufactures memory and storage solutions that include embedded storage, mobile memory, solid-state drives, and memory modules in commercial, industrial, and automotive grades. Our cutting-edge products are deployed globally to drive higher performance, increase efficiencies, and deliver unwavering reliability. Our continued commitment to innovation ensures we are positioned to deliver superior solutions for our customers, partners, and communities
Lexar Enterprise. 1737 N First Street, Suite 680, San Jose, CA USA
Unsolicited resumes sent to Lexar from third-party recruiters and placement agencies do not:
(a) constitute any type of relationship with Lexar or
(b) obligate Lexar to pay fees should we elect to hire from those resumes. Please do not contact or present candidates directly to any Lexar personnel.act or present candidates directly to any Lexar personnel.
#J-18808-Ljbffr Jobleads-US- ...Lexar Enterprise seeks a senior systems architect to own the end-to-end technical architecture for sparse MoE inference and memory management. You will convert strategic direction... ...policies, and performance models across runtimes and hardware. You will bridge model architecture...Suggested
$272k - $431.25k
...is to innovate how we architect and develop our GPU for the changing AI and accelerated workloads... .... We are looking for a Principal System Architect with a wealth... ...SoC subsystems such as memory architecture, test... ...fundamentals, including memory hierarchy, coherency, clocking,...PrincipalFull timeRemote work$152k - $241.5k
NVIDIA is seeking a motivated architect to work with a team in... ...silicon processes. This GPU memory architecture team creates new... ..., autonomous vehicles, AI, gaming, mobile systems. What you will be doing: Developing... ...interconnects, QoS. Memory hierarchy, Memory model & ordering,...SuggestedFull timeWork experience placement- Position: Principal Engineer/Architect - Generative AI & Edge Systems, Automotive Location: San Jose CA (4 days onsite) Employment... ...(e.g., TensorRT-LLM, ONNX Runtime, GGUF/llama.cpp). Proven experience... ...analysis across latency, memory bandwidth, and CPU vs. NPU vs. GPU...PrincipalFull timeRelocation package
- ...innovations in Flash and advanced memory technologies, our solutions... ...Job Description A System Architect focusing on advanced memory technology and AI Inference solution infrastructure... ...integration of memory subsystems, cache hierarchies, and system interconnects...PrincipalFull timeTemporary workRemote workFlexible hoursShift work
$219k - $351k
...partners, and communities.Job Title: Principal Engineer, AI System Architect (Hardware)The Architecture Research... ...in modern AI, particularly in memory capacity/bandwidth and system-scale... ...architectures, including compute, memory hierarchies, and high-performance interconnects...PrincipalWork at officeFlexible hours$180k - $240k
Overview AI Memory Solution Architect, Principal Engineer role at SK hynix America Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite... ...analysis to ensure high efficiency. Collaborate with system software and hardware design teams to ensure seamless integration...PrincipalFull timeWork at officeLocal area$200k - $300k
Job Title: Principal ASIC Architect - Memory Systems & AI InterconnectsJob Location: Santa Clara, CA or Boston, MACompensation: $200K - $300K plus bonus... ...coherent fabrics, memory pooling, and tiered memory hierarchies.Lead specification and micro-architecture of...Principal$152k - $241.5k
...candidates with a track record of architecture development to join our memory system architecture team! This team drives memory system architecture... ...6, 2026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to...Full time$272k - $431.25k
...serving generative AI and reasoning... ...feel like a single system at datacenter scale... ...rapidly outgrow the memory and compute budget... ....We are seeking a Principal Systems Engineer to... ...LLM inference.Architect and implement deep... ...understanding of memory hierarchies (GPU HBM, host...PrincipalFull timeLocal areaRemote work$200k
...Role Overview We are looking for a Principal AI SoC Runtime Software Architect to own the software architecture... ...candidate combines deep runtime and systems-software expertise with a strong... ...driver interfaces, DMA and shared-memory architectures, and performance-sensitive...PrincipalContract workFlexible hours- ...differentiated SoC architectures for AI/ML accelerators, hyperscale... .... Job Description We are seeking a Principal / Lead DSP & Systems Architect to own the architectural direction of... ...on micro-architecture, latency, and memory trade-offs. Silicon Bring-up & Validation...Principal
- Micron Technology is seeking a System Architect to lead performance modeling and analysis of next‑generation memory architectures for AI, server, data center, and mobile platforms. You will build and validate cycle‑accurate models and translate results into architecture...
- Micron Technology in San Jose, California, seeks a System Architect to lead performance modeling and analysis for next‑generation memory architectures across AI, server, data center, and mobile platforms. You will build and validate timing-approximate and cycle-accurate...
- Micron Technology is seeking a System Architect to lead performance modeling and analysis of next‑generation memory architectures for AI, server, data center, and mobile platforms. You will build timing‑approximate and cycle‑accurate models and translate results into architecture...
$148k - $308k
...world leader in innovating memory and storage solutions that accelerate... ...where memory, storage, and system architecture come together to shape the future of AI and datacenter platforms. We... ...work here matters!As a Sr. Principal Solutions Architect for Network Interconnect...PrincipalFull timeWork at officeLocal areaImmediate start- ...experiences—from AI and data centers,... ...gaming and embedded systems. Grounded in a culture... ...are seeking a Principal GenAI Inference Optimization... ...GPU architecture, memory systems, and... ...—from kernels and runtimes to frameworks and... ...(compute, memory hierarchy, interconnects...Principal
- Dell Technologies in Santa Clara, CA, seeks a Senior Principal Systems Development Engineer to architect and deliver end-to-end AI rack solutions. You will own the system-level design across compute, mechanical, thermal, software, diagnostics and factory enablement, working...Principal
$184k - $287.5k
...GPUs excel.We are seeking a GPU System Architect who will architect and design... ...datacenter platforms for AI and HPC. The architect in this... ...GPU compute, high-bandwidth memory, in-package interconnects and... ...software interaction, drivers and runtimes, and performance tuning for...Full timeWorldwide- Samsung Semiconductor in San Jose is hiring a Staff Engineer, Memory Systems Architecture to advance fault management for data center memory... ...failure modes, ECC schemes, and RAS algorithms to reduce downtime in AI/ML workloads. You will collaborate with customers, analyze...
$272k - $431.25k
...NVIDIA’s Networking Systems & Software Architecture... ...is solving some of AI’s hardest infrastructure... ...interconnects.This Principal Architect role leads the research... ...in communication runtimes for large-scale AI workloads... ...architecture, memory hierarchies, DMA engines, and OS-...PrincipalFull timeRemote work- ...Role: Principal/Lead Design Engineer Memory Controller- (CHI-G / DFI) Location: San Jose, CA Let's create our future together at The AES Group... ...their business operations with the power of cloud, data, AI, engineering, and other emerging technologies. Role...Principal
$146k - $286k
...a world leader in innovating memory and storage solutions that accelerate... ...ever.Micron is seeking a System Architect to lead performance modeling... ...memory architectures for AI, server, data center, and mobile... ...of DRAM architecture, memory hierarchy design, and cache-coherent...Full timeLocal areaImmediate start$141k - $319k
...Technology is a world leader in innovating memory and storage solutions that accelerate the... ...and advance faster than ever.The Principal Product Marketing Manager for the Cloud Memory... ...the value of Micron's memory portfolio for AI, cloud, and data center markets. This...PrincipalFull timeLocal areaImmediate start$180k - $260k
Senior Staff Engineer/Principal Memory Design Engineer Senior Staff Engineer/Principal Memory Design... ...Be among the first 25 applicants Get AI-powered advice on this job and more exclusive... ...memory design engineer to define and architect memory circuits for high performance...PrincipalFull timeTemporary workWorldwide- Micron Technology, Inc is seeking a Technical Solutions Sales Architect to lead early technical engagement with strategic customers and drive design wins for memory and storage solutions across AI, cloud, and data‑center platforms. You will translate workloads into differentiated...
- Micron Technology, Inc. in San Jose, CA is seeking a System Architect to drive exploration, modeling, and analysis of memory and storage subsystem architectures for AI, server, and mobile computing systems. You will collaborate across design, software, and engineering...
- ...performance across compiler, runtime, and hardware layers to... ..., high-performance AI applications.... ...quantization, scheduling, and memory optimization. Develop... ...single-node and distributed systems. Requirements... ...Knowledge of memory hierarchy optimization, caching,...PrincipalFull timeTemporary workFlexible hours
- ...computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture... ...seeking a Robotics AI Architect to define and scale... ...roadmap for our AI SDKs, runtime, and reference architectures... ...models Memory and dataflow efficiency...Principal
- ...are seeking an outstanding and visionary Lead Architect to drive the definition and exploration of next-generation system architectures and technologies that craft the... ...innovative solutions—from 3D integration and advance memory systems to emerging system-level paradigms....PrincipalWork at officeLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal AI Systems Architect, MoE Runtime & Memory Hierarchy. Be the first to apply!
- pega system architect San Jose, CA
- technical architect San Jose, CA
- system architect San Jose, CA
- embedded systems architect San Jose, CA
- principal data scientist San Jose, CA
- senior principal scientist San Jose, CA
- principal architect San Jose, CA
- senior principal cloud computing engineer San Jose, CA
- principal San Jose, CA
- principal cloud computing engineer San Jose, CA


