Principal LLM Inference Engineer (Santa Clara)
d-Matrix
At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together, we can help shape the endless possibilities of AI. D-Matrix Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack: from pathfinding emerging use cases and novel deployment patterns to deep optimization of inference kernels, to building proof-of-concept systems that showcase D-Matrix’s unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly with product and hardware teams to shape the roadmap.We build the tools, runtimes, and frameworks that let frontier AI models run efficiently and cost-effectively across heterogeneous deployments — combining D-Matrix silicon with CPUs, GPUs, and custom accelerators. Our work powers everything from benchmarking and evaluation pipelines to production-grade inference serving.This RoleWe are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack — from kernel-level optimization to distributed orchestration to high-level serving APIs.This role could be a great match for you if you:• Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time.• Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed.• Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas.• Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff.• Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used.• Value clear communication and thrive in a small, high-ownership team environment.Responsibilities• Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.• Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.• Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.• Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.• Build and maintain inference runtimes, serving frameworks, and evaluation tooling.• Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.• Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.• Partner with product and business development to translate POCs into customer-facing demonstrations.• Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.Required Qualifications• Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience. • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.• Strong proficiency in Python and C/C++.• Hands-on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).• Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.• Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.Preferred Qualifications• Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).• Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments).• Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.• Contributions to open-source inference or ML systems projects.• Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).• Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.• Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.Why D-Matrix Frontier Group• Work on genuinely novel hardware — D-Matrix in-memory compute architecture opens up inference optimization problems that don’t exist anywhere else.• End-to-end ownership from idea to deployed system, with a short feedback loop between your work and real hardware.• Small, senior team with high autonomy and direct influence on product direction.• Competitive compensation, equity, and benefits in Santa Clara, CA.Equal Opportunity Employment Policyd-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation. Compensation Range: $195K - $285KLocationSanta ClaraEmployment TypeFull timeLocation TypeHybridDepartmentArchitectureCompensationL6$195K – $285K • Offers Equity • Offers BonusThe pay range below is for all roles at this level across all US locations and functions. Individual pay rates depend on a number of factors—including the role’s function and location, as well as the individual’s knowledge, skills, experience, education, and training. We also offer incentive opportunities that reward employees based on individual and company performance. This is in addition to our diverse package of benefits centered around the wellbeing of our employees and their loved ones. In addition to the usual Medical/Dental/Vision/401k, our inclusive rewards plan empowers our people to care for their whole selves. An investment in your future is an investment in ours.
$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post... ...will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs,... ...characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...PrincipalFull timePart time$184k - $287.5k
...looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior... ...and multimodal model inference as part of NVIDIA Inference Microservices... ...production code to TRT-LLM, NVIDIA’s open-source inference... ...law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full...SuggestedFull timePart time$195.2k - $361.2k
...people actually own. You optimize inference engines (llama.cpp, vLLM) for... ...systems-level codeExperience with LLM inference. (attention, KV... ...Primary Location: US, California, Santa ClaraAdditional Locations:US,... ...: US, California, Santa Clara; US, Oregon, Hillsboro; US, California...SuggestedFull timePart timeInternshipLocal areaImmediate startShift work- ...Location:Hybrid, working onsite at our Santa Clara, Ca headquarters 3-5 days per week... ...architecture of AI compute. As a Principal Hardware Design Engineer, you will be a cornerstone of our... ...solve the industry's most massive LLM inference challenges.Required Qualifications...PrincipalPart time3 days per week
$272k - $431.25k
...a high-throughput, low-latency inference framework for serving generative... ...deployment of cutting-edge LLM workloads.We are seeking a Principal Systems Engineer to define the vision and roadmap... ...by law.SummaryLocation: US, CA, Santa Clara; US, WA, Remote; US, MA, RemoteType...PrincipalFull timePart timeLocal areaRemote work$182.5k - $260.5k
...employees spread across offices in Santa Clara, St. Louis, Bangalore, London... ...Scientist, you own the inference and optimization layer that makes... ...with the systems and backend engineers to ship capabilities end-to-... ...(vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or...PrincipalPart timeWork at office- ...advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...head-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why...Part time
$184k - $287.5k
...highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale... ...crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang... ...protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart time$184k - $287.5k
...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis... ...status, disability status or any other characteristic protected by law.SummaryLocation: US, CA, Santa ClaraType: Full time...Full timePart timeWork experience placement$210.16k - $271.98k
...Senior Principal Systems Development EngineerHelp architect and deliver Dell's L11 rack... ...growth. As a Principal Systems Development Engineer, you'll own the system-level... ...Systems Development Engineering team in Santa Clara, California.Essential Requirements12+ years...PrincipalPart time- ...your career. THE ROLE:As a Principal Engineer, you will spearhead the next... ...gains in both training and inference pipelines through innovative... ...for multi-trillion parameter LLM training/inference including... ...experienceLOCATION: Austin, Tx or Santa Clara, Ca strongly preferred;...PrincipalPart timeRemote work
$136.88k - $205k
...support high volume production.What You Can ExpectAs a Principal Test Development Engineer in the Operations business group, you will test features... ...subject to an export license review process prior to employment.#LI-AO1SummaryLocation: Santa Clara, CAType: Full time...PrincipalPermanent employmentFull timeContract workPart timeInternshipWork from home$272k - $431.25k
...scale and complexity of these LLM systems continues to increase, we are seeking outstanding engineers to join our team and help shape the future of LLM inference.Our team is dedicated to pushing... ...by law.SummaryLocation: US, CA, Santa Clara; US, CA, RemoteType: Full time...PrincipalFull timePart time$272k - $431.25k
...silicon Aarchitects and software engineers to influence hardware... ...InfiniBand verbs is required.Inference & Serving: Advanced knowledge... ...schedulers, specifically TensorRT-LLM, vLLM, SGLang, and NVIDIA... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, RemoteType...PrincipalFull timePart time$272k - $431.25k
...computing interconnects.This Principal Architect role leads... ...SGLang, and TensorRT-LLM.Publishing findings,... ...and mentoring senior engineers across the... ...distributed training and inference patterns.Proficiency in... ...SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, TX...PrincipalFull timePart timeRemote work$147k - $299.5k
...spread across offices in Santa Clara, St. Louis, Bangalore, London... ...are available at both Principal and Distinguished Engineer levels. Candidates are... ...Architect Production-Grade Inference Systems: Design, optimize... ...systems, leveraging the latest LLM serving technologies such...PrincipalPart timeWork at office$184k - $287.5k
...millions globally. We seek a Senior Engineer to lead technical efforts in... ...By combining powerful local inference (Nemotron models) with strong... ....Proven understanding of LLM inference pipelines (Ollama,... ...law.SummaryLocation: US, CA, Santa Clara; US, WA, RedmondType: Full time...Full timePart timeLocal areaShift work$184k - $287.5k
...aspects related to tasks like large scale LLM training and inference.Conducting regular technical... ...see:BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or other... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, WA, SeattleType: Full time...Full timePart time$150.68k - $225.7k
...Looking For• Bachelor’s degree in computer science, Electrical Engineering or related fields and 8+ years of related professional... ..., all applicants may be subject to an export license review process prior to employment.#LI-SA1SummaryLocation: Santa Clara, CAType:...PrincipalPermanent employmentPart timeInternshipWork from home$220.92k - $311.89k
...seeking a highly experienced and motivated Principal Analog Design Engineer to lead the design and validation of... ...)Primary Location: US, California, Santa ClaraAdditional Locations:US, Arizona... ...: US, California, Santa Clara; US, Arizona, Phoenix; US, Oregon, HillsboroType...PrincipalFull timePart timeLocal areaImmediate startShift work$272k - $431.25k
...We're looking for a Principal Software Engineer to join our CSP Engagements team as the technical focal point for rack-scale system SW/FW, working... ...characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Remote; US, CA, RemoteType: Full...PrincipalFull timePart timeRemote workShift work$272k - $431.25k
...R&D group requires a senior software engineer. In this exciting role, you will profile... ...used for distributed Deep Learning LLM training and inference. Your primary focus will be... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Remote; US, CO, Remote; US,...PrincipalFull timePart timeRemote work$272k - $431.25k
...production-ready software.Collaborate with GPU architects, driver engineers, SDK teams, and graphics developers to understand future needs... ...characteristic protected by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR, Hillsboro; US, NC, Durham; US, WA, SeattleType...PrincipalFull timePart time$200k - $250k
...Job Title: Principal Design Verification Engineer - PCIe / High-Speed SoCLocation: Santa Clara, CACompensation: $200K - $250K base DOE plus bonus and RSUs ($300K - $350K+ total comp.)Requirements: Design Verification, High-Speed SoC, PCIe / UALink / SerDes, SystemVerilog...PrincipalPart time- ...advance your career. THE ROLE:As a Principal AI Infrastructure Solution Engineer, you will partner with AMD’s... ...to enable large‑scale LLM training and inference on AMD Instinct GPUs. You will... ...related field requiredLOCATION:Santa Clara, Ca or open to discuss other locations...PrincipalPart time
$224k - $356.5k
...growing standard for assessing LLM serving performance across various inference frameworks. Hyperscalers, cloud... ...Lead Manager, you will lead the engineering team within NVIDIA’s Dynamo organization... ...by law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, AL, Remote;...Full timePart timeLocal areaRemote workWorldwide$136.88k - $205k
...ImpactMarvell is seeking a highly motivated and talented Principal Test Engineer to join its dynamic New Product Introduction (NPI) test engineering... ...applicants may be subject to an export license review process prior to employment.#LI-TT1SummaryLocation: Santa Clara, CAType:...PrincipalPermanent employmentPart timeInternshipWork from home$272k - $431.25k
...make a lasting impact on the world.At NVIDIA, as a Principal Rack Scale Systems Infrastructure Engineer, you will build and guide the development of software... ...protected by law.SummaryLocation: US, CA, Santa Clara; US, NC, Remote; US, TX, Remote; US, Remote; US, MA...PrincipalFull timePart timeRemote workShift work$158.6k - $237.6k
...art multi-core SoCs.• Transforming the requirements from the engineering teams into software tools that are both easy to use and scalable... ..., all applicants may be subject to an export license review process prior to employment.#LI-SA1SummaryLocation: Santa Clara, CAType:...PrincipalPermanent employmentPart timeInternshipWork from home$272k - $431.25k
...We're looking for a Principal Engineer to join our CSP Engagements team as the... ...environmentsUnderstanding of inference workload performance dynamics (vLLM, TensorRT-LLM, SGLang, continuous batching)... ...law.SummaryLocation: US, CA, Santa Clara; US, TX, Austin; US, OR,...PrincipalFull timePart timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal LLM Inference Engineer (Santa Clara). Be the first to apply!
- senior director engineering Santa Clara, CA
- chief engineer Santa Clara, CA
- senior principal engineer Santa Clara, CA
- engineering director Santa Clara, CA
- senior chief engineer Santa Clara, CA
- principal infrastructure engineer Santa Clara, CA
- data center chief engineer Santa Clara, CA
- senior civil engineer project manager Santa Clara, CA
- general engineer Santa Clara, CA
- principal engineer Santa Clara, CA








