Full-Stack AI Compute Architect
Oxmiq Labs
Updated Role | Now hiring: Full-Stack AI Compute Architect About OXMIQ OXMIQ provides complete hardware and software GPU IP that lets our customers build their own AI silicon. Founded by Raja Koduri, we recently closed a $35M Series A (co-led by Samsung Catalyst Fund and Fundomo, with MediaTek, Intel Capital, and others), bringing total funding to $60M, with Jim Keller joining our board. Our OxCore architecture combines scalar, tensor, SIMT, and orchestration engines on one platform — delivering CUDA compatibility alongside a fully programmable architecture. On top of it runs OxPython, our software stack that runs existing CUDA and PyTorch code unmodified with day-zero support for new AI models. We have an established team and a working stack, and we're now shifting from development to production — hardening the platform and scaling it for inference at real customer scale. You'll join the OXMIQ architecture team, which sets the direction of our software, silicon, and systems. You'll help refine our software strategy and deliver the highest-performance solutions by optimizing the whole stack — from AI models and frameworks at the top down to kernel and hardware-level code generation. The Scope of This Role This role spans the AI software stack from models on down — the full vertical path a workload travels: from a model authored in frameworks such as PyTorch, JAX, or ONNX, through graph capture and the compiler, into the runtime and scheduler, down to individual kernels and the instructions that execute on OxCore. We're looking for an architect who is fluent across that whole range — someone who can reason about how a transformer or diffusion model is structured and about warp divergence and register allocation, and who can make the choices that connect those layers into one coherent, high-performance platform. You won't be the deepest specialist at every layer, but you should be able to hold the whole range in your head and drive decisions end to end. Key Responsibilities Refine the technical direction of OxPython, the OXMIQ software stack — from framework front-ends (PyTorch, JAX, ONNX) and model ingestion, through graph capture, the compiler, runtime, and driver layers, down to kernels — in close partnership with the existing software team. Own how AI models map onto the platform: understand how modern workloads (LLMs, diffusion, vision, agentic inference, and beyond) are structured and expressed at the framework level, and shape how they are captured, partitioned, and lowered so they run unmodified with day-zero support and scale efficiently for inference. Own whole-stack performance: drive optimization end to end, identify and remove bottlenecks at every layer — from graph-level and framework-level inefficiencies down through compiler, runtime, driver, and kernels — and close the gap to peak hardware utilization. Guide the graph and compiler strategy: MLIR dialect and IR design, lowering pipelines, operator representation, and LLVM backend code generation targeting Oxmiq hardware IP. Architect work orchestration across OxCore's scalar, tensor, and SIMT engines, and shape the compiler/runtime approach that runs CUDA-optimized models unmodified — preserving day-zero support for new AI models without sacrificing programmability. Advance the SIMT execution and parallel compute strategy: thread/warp scheduling, divergence handling, occupancy, and the memory hierarchy, and how parallel workloads are expressed, lowered, and mapped onto the hardware. Refine and optimize the strategy for kernel programming models (including Triton-based and custom op lowering), kernel fusion, and tiling, in collaboration with kernel engineers. Design for scale and production: ensure solutions grow cleanly with customer workloads, deployment sizes, and model complexity, and harden them for real-world delivery. Work with architecture and hardware teams to translate performance requirements into efficient code generation and runtime strategies, and feed software needs back into hardware definition. Drive the software/hardware validation strategy — evolving how the stack and silicon are co-verified as we scale from bring-up through production. Raise the technical bar through design reviews, code reviews, and architecture discussions, and mentor engineers across the frameworks, compiler, runtime, and kernel teams. Guide performance profiling, benchmarking, and root cause analysis for compiler- and runtime-generated code. Required Qualifications 12+ years building systems and/or AI-infrastructure software, with a track record of owning architecture for a major component or a full stack that spans multiple layers. Breadth across the AI software stack — you can reason about the problem from the model and framework level down to the kernel and instruction level, and understand how choices at one layer constrain the others. Hands-on experience with ML frameworks (PyTorch, JAX, ONNX) and a solid understanding of how deep learning models are structured, expressed, captured, and optimized — not just at the operator level but as whole graphs and workloads. Strong expertise in C++ and compiler infrastructure, including LLVM (IR, pass infrastructure, backend code generation) and MLIR (dialect design, conversion passes, progressive lowering). Deep expertise in SIMT and parallel compute architectures — thread/warp execution models, divergence, synchronization, occupancy constraints, and memory hierarchies. Proven experience with GPU programming models, including CUDA, and an understanding of what it takes to deliver CUDA compatibility on non-NVIDIA hardware. Solid command of code generation concepts — instruction selection, scheduling, register allocation, vectorization, and memory hierarchy optimization. Experience co-designing software alongside silicon or shaping a hardware-software interface. Strong debugging and performance analysis skills, and a bias toward shipping: you prototype, measure, and make decisions with incomplete information. Hands-on experience with Claude Code or equivalent AI-assisted development workflows. Note: we expect exceptional depth in some of these areas and working fluency across the rest. The essential requirement is the ability to operate across the whole stack, not mastery of every layer. Preferred Qualifications Experience architecting compiler backends or runtimes for custom AI accelerators, NPUs, or GPU IP. Experience with graph-level model optimization, framework integration internals (e.g., PyTorch compile/inductor, XLA), or serving/inference infrastructure for large models. Experience with Triton or similar high-level GPU kernel frameworks, and with kernel fusion, operator tiling, and auto-tuning. Exposure to quantization, mixed-precision, or sparsity-aware compilation techniques. Experience with performance modeling and roofline analysis for GPU workloads. Background in deep learning, computer vision, image processing, or video — the workloads our customers run. Experience bringing up and scaling a software stack on new or pre-silicon hardware into production. Contributions to open-source compiler, runtime, or ML-framework projects such as LLVM, MLIR, PyTorch, or Triton. Education BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field. Why OXMIQ You'll join a well-funded company — fresh off a $35M Series A backed by Samsung, MediaTek, Intel Capital, and others, with Jim Keller on our board — with an established team and a proven stack, at the point where we're scaling it into a production platform in customers' hands. You'll work directly with the people designing the hardware it runs on. This is a role for an architect who wants to refine and optimize the AI software stack end to end — from models and frameworks down to kernels and silicon — rather than just one layer of it, and measure success in shipped, high-performance solutions. OXMIQ is an equal opportunity employer. We welcome applicants of all backgrounds. #J-18808-Ljbffr Oxmiq Labs
- Oxmiq Labs in California seeks a Full-Stack AI Compute Architect to define the software stack—from model ingestion and graph capture through compiler, runtime, and kernels—driving high-performance, CUDA-compatible solutions. You will own end-to-end performance, guide MLIR...Fullstack
- ...AI Full Stack Developer/ ArchitectLocation: San Jose, CA - Onsite 5 daysDuration: 12 + MonthsInterview: VideoJob SummaryWe are looking... ...(Docker, Kubernetes).Education: Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, or a related field....Fullstack
- codvo-team is seeking an Edge AI Architect to lead design and deployment of AI models on edge devices for medical applications. The role focuses on real-time computer vision, integration with embedded systems, and ensuring SaMD regulatory alignment. The ideal candidate...SuggestedRemote work
- Job Description: Edge AI Architect - CUDA / C++ / Computer Vision Experience Level: 10+ Years Department: Edge AI & Embedded Systems About the Role We are looking for a highly motivated and technically proficient Edge AI Engineer with a strong background in Edge AI devices...SuggestedContract work
- d-Matrix is seeking a Senior Staff SI/PI Engineer to be the technical authority for electrical integrity of high-performance AI compute platforms. You will own end-to-end SI/PI modeling, simulation, and correlation from silicon die through package to system PCBA, leading...Suggested
$120k - $275k
...tailored for the world’s best AI models. Our hardware will... ...developing vertically integrated full-stack solutions from silicon to... ...largest models. MatX is seeking an Architect to join our team as we create... ...is seeking an AI Accelerator Compute Architect to help define the...FullstackFull timeWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours- d-Matrix inc. is seeking an AI SoC Architect in Santa Clara, CA for a hybrid role. You will lead the design and delivery of system-on-chip technologies for our AI compute engine, focusing on performance and efficiency in collaboration with various engineering teams. The...
$184.7k - $324.8k
Sunnyvale, California, United States Machine Learning and AI The Video Computer Vision and Video Engineering teams are centralized applied research... ..., groundbreaking experiences, innovating through the full stack, and partnering with HW and SW teams to influence SW...FullstackRelocation$120.7k - $238.6k
...Computer Scientist Opportunity At Adobe's EPGAdobe's EPG (Emerging Products... ...in computer vision, modern AI, or computational photography,... ...mobile camera. The app offers full manual controls, a more... ...and AI, and if you enjoy full-stack challenges – from research to...FullstackWorldwide$120k - $275k
...tailored for the world’s best AI models. Our hardware will... ...developing vertically integrated full-stack solutions from silicon to... ...largest models. MatX is seeking an Architect to join our team as we create... ...is seeking an AI Accelerator Compute Architect to help define the...FullstackDaily paidFull timeWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours$120.7k - $238.6k
...the Nextcam team, is seeking Computer Scientists with graduate training... ...in computer vision, modern AI, or computational photography,... ...mobile camera. The app offers full manual controls, a more natural... ...and AI, and if you enjoy full-stack challenges - from research to...FullstackTemporary workLocal areaWorldwide- MatX Inc. in Mountain View, CA is seeking an AI Accelerator Compute Architect to define the compute architecture for next-generation GenAI accelerators. You will translate workloads into hardware architecture and evaluate ISA, memory, and interconnect tradeoffs. Join a...
- ...partnered with a well-funded startup redefining compute for state-of-the-art ML models, seeking a Computer Architect in Mountain View, CA. You will define architectures... ...next-generation compute engines, translating ML/AI workloads into hardware specs from concept through...
$272k - $431.25k
...ODC) module — a Vera Rubin–class compute platform engineered for low-... ...generation orbital roadmap to speed up AI adoption. We are looking for a strong technical architect to own end-to-end system... ...platforms. You will architect the full stack — application to libraries,...FullstackFull timeWork experience placementRemote work$164.47k - $269.1k
...networking silicon. Our team architects next-generation networking solutions... ..., cloud infrastructure, and AI workloads to achieve... ...backbone of modern distributed computing systems.We are seeking a Senior... ...-level thinking: Understands full-stack implications (hardware,...FullstackFull timeLocal areaImmediate startShift work$174k - $253k
...network, or service operations and quality.Design and implement computer vision systems, leverage ML infrastructure, and evaluate... ...qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.The Google Pixel team...Fullstack$154.38k - $193.13k
Job DescriptionAI Architect, AI & AutomationAbout the Role The applicant should have deep, hands... ...QualificationsBachelor’s degree in computer science, Engineering, Data Science, or a... ...193,125Along with competitive pay, as a full-time Infosys employee you are also eligible...Full timeTemporary workWork experience placement$272k - $431.25k
...group is solving some of AI’s hardest infrastructure... ...directly into production AI stacks. The team's charter... ...domains including quantum computing interconnects.This Principal Architect role leads the research agenda... ...Remote; US, CO, Remote; US, OR, RemoteType: Full time...Full timeRemote work$184k - $287.5k
NVIDIA is hiring an AI Hardware Architect to analyze and architect the next generation of Artificial... ...infrastructure market. If you have knowledge of computer architecture, microarchitecture, and... ...by law.SummaryLocation: US, CA, Santa Clara; US, TX, AustinType: Full timeFull time$224k - $356.5k
NVIDIA has been transforming computer graphics, PC gaming, and accelerated... ...the unlimited potential of AI to define the next era of... ...world.We’re looking for a Senior Architect to help shape the next generation... ...: US, CA, Santa Clara; US, CA, RemoteType: Full timeFull time$184k - $287.5k
NVIDIA is seeking a Compute Kernel Performance Architect who can develop, profile, and analyze... ...GPU performance and power stack. We collaborate closely... ...existing vacancy. NVIDIA uses AI tools in its recruiting... ...by law.SummaryLocation: US, CA, Santa ClaraType: Full timeFull time- ...products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and... ...ROLE: We are seeking a Robotics AI Architect to define and scale next-generation... ...experience with AMD GPU and NPU AI SW stacks and tools will be a plusACADEMIC CREDENTIALS...
$206.4k - $379.1k
...impressive content. The AI Foundations team... ...looking for a Principal Architect to build and implement... ...evolve the complete AI stack for Adobe Express — covering... ...equivalent experience in Computer Science, Data Science,... ...SummaryLocation: San Jose; San FranciscoType: Full time...Full timeTemporary workLocal areaWorldwideFlexible hours$184k - $287.5k
...now looking for a Senior Deep Learning Computer Architect! NVIDIA is seeking architects like you to... ...next-generation GPUs advance the state of AI.This position requires you to keep up... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, RedmondType: Full timeFull timeNight shift$200k - $300k
Job Title: Principal ASIC Architect - Memory Systems & AI InterconnectsJob Location: Santa Clara, CA or Boston, MACompensation: $200K - $300K plus... ...on solving the memory wall for AI. We develop memory-to-compute networking solutions for AI infrastructure, with a focus...$200k - $250k
Job Title: Performance Modeling Architect - AI Memory SystemsJob Location: Santa Clara, CA, or Boston, MACompensation: $200K - $250K base... ...performance bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workload settings....Night shift$184k - $287.5k
We are now looking for a Senior AI Training Performance... ...the largest and most powerful compute systems in the world. If you are... ...layers of the hardware/software stack - from GPU architecture to the... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full timeFull time$184k - $287.5k
...are now looking for a Senior Accelerated Computing Architect!NVIDIA is developing software and system... ...scientific computing, machine learning, AI, datacenter, and automotive computing.... ...other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full timeFull time$208k - $327.75k
...the forefront of accelerated computing, AI, and autonomous machines. From... ...are looking for a Senior AI Architect to help define the next generation... ...future autonomous vehicle stack, including Vision-Language-... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full timeFull timeWorldwide$184k - $287.5k
...modern Artificial Intelligence, Parallel Compute Infrastructure, and Agentic AI - the biggest technology... ...and we’re seeking a visionary Product Architect with strong expertise in systems architecture... ...by law.SummaryLocation: US, CA, Santa Clara; US, RemoteType: Full timeFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Full-Stack AI Compute Architect. Be the first to apply!

