Senior Staff LLM Inference Engineer
d-Matrix
At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together, we can help shape the endless possibilities of AI. D-Matrix Frontier Group sits at the leading edge of what’s possible with LLM inference on heterogeneous hardware. Our charter spans the full stack: from pathfinding emerging use cases and novel deployment patterns to deep optimization of inference kernels, to building proof-of-concept systems that showcase D-Matrix’s unique computational fabric. We are an applied research and engineering team that moves fast, ships real systems, and works directly with product and hardware teams to shape the roadmap.We build the tools, runtimes, and frameworks that let frontier AI models run efficiently and cost-effectively across heterogeneous deployments — combining D-Matrix silicon with CPUs, GPUs, and custom accelerators. Our work powers everything from benchmarking and evaluation pipelines to production-grade inference serving.This RoleWe are hiring end-to-end inference engineers who are comfortable going from a novel research idea to a deployed, optimized system. You will work at every layer of the inference stack — from kernel-level optimization to distributed orchestration to high-level serving APIs.This role could be a great match for you if you:• Have deep intuition for modern generative AI architectures and how to squeeze performance out of them at inference time.• Are familiar with the internals of open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and can extend or replace them when needed.• Enjoy pathfinding new use cases — exploring heterogeneous deployment topologies and building early-stage POCs that prove out new ideas.• Are results-oriented with a strong bias toward action; you own problems end-to-end from prototype to optimization to handoff.• Are energized by working at the intersection of novel hardware and frontier models, and want your work to directly influence how next-generation AI silicon is used.• Value clear communication and thrive in a small, high-ownership team environment.Responsibilities• Identify and prototype emerging LLM inference use cases suited to heterogeneous hardware deployments.• Build compelling proof-of-concept systems that demonstrate D-Matrix capabilities to customers, partners, and internal stakeholders.• Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency.• Drive quantization, sparsity, and batching strategies tailored to D-Matrix computational model.• Build and maintain inference runtimes, serving frameworks, and evaluation tooling.• Contribute to distributed inference systems: tensor/pipeline parallelism, disaggregated prefill/decode, KV-cache management.• Work closely with hardware architects to provide firmware and compiler teams with actionable inference workload insights.• Partner with product and business development to translate POCs into customer-facing demonstrations.• Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility.Required Qualifications• Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience. • Master’s or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.• Strong proficiency in Python and C/C++.• Hands-on experience optimizing LLM inference — attention kernels, KV cache, batching strategies, quantization (INT8/FP8/INT4).• Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level.• Familiarity with GPU kernel programming (CUDA/Triton) and performance profiling tools.Preferred Qualifications• Experience with heterogeneous compute deployments — scheduling inference workloads across dissimilar hardware (accelerators, CPUs, GPUs).• Familiarity with custom silicon or ASIC-based inference (beyond GPU-only environments).• Experience with distributed inference: tensor parallelism, pipeline parallelism, disaggregated serving.• Contributions to open-source inference or ML systems projects.• Experience with production inference serving at scale (latency SLOs, continuous batching, multi-model serving).• Familiarity with speculative decoding, mixture-of-experts routing, or long-context serving techniques.• Working familiarity with the material in the JAX Scaling Book or equivalent systems-level understanding of modern LLM training and inference.Why D-Matrix Frontier Group• Work on genuinely novel hardware — D-Matrix in-memory compute architecture opens up inference optimization problems that don’t exist anywhere else.• End-to-end ownership from idea to deployed system, with a short feedback loop between your work and real hardware.• Small, senior team with high autonomy and direct influence on product direction.• Competitive compensation, equity, and benefits in Santa Clara, CA.Equal Opportunity Employment Policyd-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation. Compensation Range: $195K - $285KLocationSanta ClaraEmployment TypeFull timeLocation TypeHybridDepartmentArchitectureCompensationL6$195K – $285K • Offers Equity • Offers BonusThe pay range below is for all roles at this level across all US locations and functions. Individual pay rates depend on a number of factors—including the role’s function and location, as well as the individual’s knowledge, skills, experience, education, and training. We also offer incentive opportunities that reward employees based on individual and company performance. This is in addition to our diverse package of benefits centered around the wellbeing of our employees and their loved ones. In addition to the usual Medical/Dental/Vision/401k, our inclusive rewards plan empowers our people to care for their whole selves. An investment in your future is an investment in ours.
$184k - $287.5k
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful... ...language and multimodal model inference as part of NVIDIA Inference Microservices... ...bugs and deliver production code to TRT-LLM, NVIDIA’s open-source inference...SeniorFull time$184k - $287.5k
We're now looking for a Sr. Inference Engineer, for GPU Kernel Optimization! What does it take to push every LLM inference operation to its performance ceiling? Our LLM Inference Performance Analysis and Optimization team builds the answer from the ground up. We develop...SeniorFull time- ..., we advance your career. THE ROLE:We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated... ...-to-head benchmarks (AMD vs. NVIDIA) on standardized LLM workloads. Produce clear, data-backed explanations of why...Senior
- ...deliver industry-leading training and inference speeds; over 10 times faster than GPU-... ...inference.About The RoleWe are hiring a Senior Performance Engineer to join our Product team. You are an... ...inference stacks (vLLM, SGLang, TensorRT-LLM), GPU kernel-level optimization...SeniorContract workShift work
$184k - $287.5k
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency... ...out from the crowdExperience building and optimizing LLM inference engines (e.g., vLLM, SGLang).Hands-on work...SeniorFull time$184k - $287.5k
...We are now looking for a Senior High-Performance LLM Training Engineer!NVIDIA is seeking experienced engineers specializing in performance analysis and optimization to improve the efficiency of LLM training workloads, which are shaping the world's most advanced computing...SeniorFull timeWork experience placement$195.2k - $361.2k
...models fast on the hardware people actually own. You optimize inference engines (llama.cpp, vLLM) for constrained local and edge environments... ...Python; comfortable reading systems-level codeExperience with LLM inference. (attention, KV cache, decoding)Experience profiling...SeniorFull timeInternshipLocal areaImmediate startShift work- ...headquarters 3 days per week.The role: Sr. Staff, ML Researcher - LLM Algorithmic OptimizationWhat You Will... ...to optimize large language model inference on DNN accelerators we develop. You... ...mathematicians, ML researchers, and ML engineers who create and apply advanced...Senior3 days per week
$152k - $241.5k
...edge AI technology for safety-critical applications? Join NVIDIA's TensorRT team as a Senior Software Engineer, and be at the forefront of technology, enabling high-performance AI inference solutions for automotive safety and other specialized platforms. Your expertise...SeniorFull time- ...of our Polaris constellation, resulting in a system with over 99.9% accuracy. About the Role We're seeking an experienced LLM Inference Engineer to optimize our large language model (LLM) serving infrastructure. The ideal candidate has: Extensive hands‑on experience with...
$147.83k - $221.4k
...innovation, above and beyond fleeting trends, Marvell is a place to thrive, learn, and lead. Your Team, Your ImpactMarvell's Central Engineering organization provides the most advanced and key analog IPs to all businesses within Marvell. You’ll be part of a key analog team...SeniorPermanent employmentFull timeInternshipWork from home$200k - $230k
...Power Electronics FirmwareWhat You Will Be DoingChargePoint is looking for an experienced Power Electronics Controls and Firmware Engineer with more than seven years of hands-on expertise in developing embedded firmware for power control applications. The preferred candidate...Senior$138k - $206k
...closely with both hardware and software engineers to identify and address the unique... ...for modern and emerging LLM workloads.We are seeking a Senior LLM Systems Performance Engineer to... ...long-context reasoning, disaggregated inference, and Mixture-of-Experts models. Your...SeniorWork experience placementWork at officeFlexible hours$152k - $241.5k
...are looking for a motivated Deep Learning engineer to bring advanced communication... ...technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be... ...from training on scales up to 100K GPUs to inference down at microsecond latency. Communication...SeniorFull timeRemote work$230k - $250k
...Cerebras Systems in Sunnyvale, CA, seeks a Sr. Member of Technical Staff to develop resilient software for their AI chip. Responsibilities include designing robust software features, maintaining deployment workflows using AWS, and debugging software issues. Candidates...SeniorRemote work$184k - $287.5k
...millions globally. We seek a Senior Engineer to lead technical efforts in... ...By combining powerful local inference (Nemotron models) with strong... ...experience, with at least 3+ years in Staff, or Lead Architect role.BS,... ....Proven understanding of LLM inference pipelines (Ollama,...SeniorFull timeLocal areaShift work$152k - $241.5k
...impact on the world.As a Developer Technology Engineer, you will be at the forefront of... ...targeting optimal runtime performance.Improve LLM & GenAI user experience by working on... ...crowd:Experience with GPU-accelerated AI inference driven by NVIDIA APIs and SDKs, specifically...SeniorFull timeLocal area$182.5k - $260.5k
...One platform, its Zero Trust Engine, and the powerful NewEdge network... ....Positions are available at Senior Staff and above. Candidates are... ...Learning Scientist, you own the inference and optimization layer that... ...runtimes (vLLM/SGLang, TensorRT-LLM, ONNX Runtime, llama.cpp, or...Senior$168k - $270.25k
...including NeMo microservices and NVIDIA Inference Microservices (NIM), enabling... .... We are looking for a senior, technically strong test development engineer to drive quality, automation, and... ...and improve testabilityValidate LLM and AI inference workflows, including...SeniorFull time$224k - $356.5k
...sitting in their PCs.We are looking for a Senior Software Engineer to build and optimize the local... .... By combining high-performance local inference (Nemotron models) with robust privacy... ...Infrastructure: Hands-on experience with LLM inference pipelines (Ollama, llama.cpp...SeniorFull timeLocal areaWorldwide$152k - $241.5k
...”.NVIDIA is seeking top-tier AI Compiler Engineers to drive innovation within our world-class... ...problems for AI workloads (both inference and training) and successfully transition... ...and/or custom AI accelerator architectures.LLM Knowledge: Deep understanding of Large Language...Full time$244.14k - $413.16k
...respond quickly, and ensure successful implementation.Basic QualificationsBachelor's degree or higher in Computer Science, Software Engineering, Artificial Intelligence, or related fields.5-8+ years of experience in large-scale data processing or data platform development....SeniorFull timeOverseas$272k - $431.25k
As a Senior Engineering Manager for Agentic Systems & Platform Architecture, you will lead the strategy... ..., and GPU-optimized training and inference workflows.Lead integration of the AI Data... ...-source libraries; deep expertise in LLM/agent architectures—leading POCs and integrating...SeniorFull time$160k - $198k
...of our team members.What You’ll DoAs a Senior AI Systems Engineer, you will architect, deploy, and... ...for large-scale AI model training and inference. You will ensure our machine learning... ...compute scheduling alongside advanced LLM serving engines.Cross-Functional Collaboration...SeniorLocal area$272k - $431.25k
...NVIDIA is seeking a Principal Engineer to drive the performance of large-scale AI training and post-training workloads across NVIDIA’s full... ...engineering. You will analyze and optimize frontier-scale LLM workloads running on thousands of GPUs, drive improvements across...Full time$168k - $264.5k
...Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA... ..., including tracking model drift and inference latency.What we need to see:MS or BS... ...deploying and scaling Generative AI/LLM applications, integrating vector databases...SeniorFull time$248k - $396.75k
...NVIDIA's networking portfolio and aligning product, sales, engineering, architecture, marketing, and partner teams to secure... ...Ethernet, RoCE/RDMA, GPU accelerated computing, generative AI, LLM training and inference, Network SerDes, in-network compute, PCIe, and CXL....SeniorFull time$201.3k - $352.3k
Company DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis... ...contributions in agentic systems or retrieval. Exposure to LLM fine-tuning or inference optimization in production.Why join us Intelligence is...SeniorWork experience placementWork at officeImmediate startRemote workFlexible hours$184k - $287.5k
At NVIDIA, we are seeking exceptional engineers to join our autonomous driving team to design, implement, and deploy cutting-edge end-to-end... ...of our driving systems.Build, pre-train, and fine-tune LLM/VLM/VLA systems for deployment in real-world autonomous driving...SeniorFull time$136k - $218.5k
...behind next-generation silicon design? We’re seeking a Senior VLSI Library Methodology Engineer to join our team and drive the development, specification... ...efficiency in large-scale automation environmentsApplied AI/ML/LLM experience to improve EDA workflows, automation systems,...SeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Staff LLM Inference Engineer. Be the first to apply!
- engineering aide Santa Clara, CA
- technology administrator Santa Clara, CA
- senior staff engineer Santa Clara, CA
- staff engineer Santa Clara, CA
- senior staff systems engineer Santa Clara, CA
- assistant engineer Santa Clara, CA
- senior manager tax Santa Clara, CA
- senior devops Santa Clara, CA
- senior director digital marketing Santa Clara, CA
- senior international accountant Santa Clara, CA

