Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal System Software Engineer, AI Inference Execution

d-Matrix

At d-Matrix, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.We value humility and believe in direct communication. Our team is inclusive, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground? Together, we can help shape the endless possibilities of AI. Location:Hybrid, working onsite at our Santa Clara, CA, headquarters 3 days per week.The role: Principal System Software Engineer, AI Inference ExecutionWhat you will do:The role requires you to be part of the team that helps productize the SW stack for our AI compute engine. As part of the software team, you will be responsible for the development, enhancement, and maintenance of the next-generation AI deployment software. You have had past experience working across all aspects of the full-stack toolchain and understand the nuances of what it takes to optimize and trade-off various aspects of hardware-software co-design. You are able to build and scale software deliverables in a tight development window. You will work with a team of system software experts to build out the deployment infrastructure, working closely with other software (ML and compilers) and hardware experts in the company.What you will bring:Minimum:BS in Computer Science, Engineering, Math, Physics, or related degree with 12+ years of industry software development experience and MS in Computer Science, Engineering, Math, Physics, or related degree preferred with 6+ yearsStrong grasp of system software, data structures, computer architecture, and machine learning fundamentalsProficient in C/C++/Python development in Linux environment and using standard development toolsExperience with distributed, high-performance software design and implementationSelf-motivated team player with a strong sense of ownership and leadership.Preferred:MS or PhD in Computer Science, Electrical Engineering, or related fieldsExperience with inference servers/model serving frameworks (such as TensorRT-LLM, vLLM, SGLang, etc.)Experience with deep learning frameworks (such as PyTorch and TensorFlow)Experience with deep learning runtimes (such as ONNX Runtime, TensorRT, etc.).Experience with distributed systems collectives such as NCCL and OpenMPIExperience with software testing fundamentalsExperience deploying ML workloads (LLMs, VLMs, NLP, etc.) on distributed systems.Experience with Kubernetes, Ray, or other MLOps tools and techniques used from definition to deploymentPrior startup, small team, or incubation experienceWork experience at a cloud provider or AI compute/subsystem companyEqual Opportunity Employment Policyd-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We’re committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation. Compensation Range: $195K - $285KLocationSanta ClaraEmployment TypeFull timeLocation TypeHybridDepartmentSoftware EngineeringCompensationIC6 Principal$195K – $285K • Offers Equity • Offers BonusThe pay range below is for all roles at this level across all US locations and functions. Individual pay rates depend on a number of factors—including the role’s function and location, as well as the individual’s knowledge, skills, experience, education, and training. We also offer incentive opportunities that reward employees based on individual and company performance. This is in addition to our diverse package of benefits centered around the wellbeing of our employees and their loved ones. In addition to the usual Medical/Dental/Vision/401k, our inclusive rewards plan empowers our people to care for their whole selves. An investment in your future is an investment in ours.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Principal System Software Engineer, AI Inference Execution in Santa Clara, CA vacancy
  • $272k - $431.25k

     ...We are now looking for a Principal Software Engineer for LPX System Software! NVIDIA’s LPX System...  ...we engineer. We treat AI coding agents as a...  ...aggregation pipelines that execute workloads on novel silicon...  ...telemetry patterns, MPI. Inference systems and token serving... 
    Suggested
    Full time
    Shift work

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $224k - $356.5k

     ...are now looking for a Senior System Software Engineer to work on Dynamo. NVIDIA...  ...GPUs to power a revolution in AI, enabling breakthroughs in...  ...building Generative AI inference platform to make design and...  ..., and stateful, multi-turn execution.Innovate in inference-state... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a highly motivated Principal System Software Engineer to drive next-generation innovations...  ...with hardware, architecture, kernel, AI, middleware, and platform teams to deliver...  ...excellence, innovation, and execution.What We Need to See:Bachelor’s, Master... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    9 hours ago
  •  ...experiences—from AI and data centers,...  ...gaming and embedded systems. Grounded in a culture...  ...—striving for execution excellence, while...  ...LLM and Multimodal inference at scale across multi...  ...internal GPU software teams and engage with...  ...PERSON:   Skilled engineer with strong... 
    Suggested

    AMD

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

     ...now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software...  ...the team: CI, build systems, training infrastructure,...  ...tested code; fluent with AI-assisted toolingExperience...  ...certain ML layers affect execution timeFamiliarity with PyTorch... 
    Suggested
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $272k - $431.25k

     ...NVIDIA is seeking a Sr. Principal Systems Software Engineer for the Apache Spark Acceleration group. GPU accelerated data processing has moved from proof...  ...026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to... 
    Full time
    Work experience placement

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $2,000 per month

     ...Etched Etched is building AI chips that are hard-...  ...Summary Etched’s Inference SW team enables optimal...  ...skilled and motivated engineer to join our team as we...  ...architectures on Sohu systems. You’ll build SW enabling...  ...inference, intra-node execution, state management, and... 
    Full time
    Work at office
    Relocation package

    Etched

    San Jose, CA
    13 hours ago
  •  ...experiences—from AI and data...  ...gaming and embedded systems. Grounded in a...  ...challenges—striving for execution excellence,...  ...compiler and software infrastructure...  ...looking for a Principal Software Development Engineer to lead technical...  ...for ML inference workloads• Define... 

    AMD

    San Jose, CA
    1 day ago
  • $272k - $431.25k

     ...re looking for a Principal Engineer to join our CSP Engagements...  ...and drive systemic improvements in...  ...configuration, software, or workload differences...  ...audiences and executive...  ...environmentsUnderstanding of inference workload performance...  ...vacancy. NVIDIA uses AI tools in its... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $193.3k - $261.5k

     ...builds AWS Neuron, the software development kit...  ...unparalleled ML inference and training performance...  ...boundary, our engineers build systematic...  ...what's possible in AI acceleration.As...  ...across the stack from system level...  ...improving the model execution.- Software development... 
    Work experience placement
    Internship
    Local area
    Flexible hours

    Amazon

    Cupertino, CA
    2 days ago
  • $180k - $300k

     ...potential of generative AI to power the...  ...forefront of software and hardware...  ...and are driven by execution. Ready to come find...  ...looking for a Principal Software Engineer in QA to join our...  ..., product, and systems teams to design,...  ...characteristics, and the inference stack —... 

    d-Matrix

    Santa Clara, CA
    1 day ago
  • $151.8k - $332.2k

     ...you can expect We are looking for an AI Inference Engineer with a solid background in speech recognition...  ...-the-art automatic speech recognition system and ship it to various Zoom products....  ...leveraging CUDA Graphs for efficient execution scheduling, and minimizing kernel... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    2 days ago
  • $200k - $220k

     ...passionate, and committed engineers, technologists, and business...  ...is seeking an experienced AI Network Software Solution Architect to lead...  ...configuration, and monitoring.System Performance &...  ..., and work with high-level executives.Strong hands-on experience... 
    Worldwide

    Supermicro

    San Jose, CA
    2 days ago
  • $224k - $356.5k

     ...NVIDIA’s workstation-class AI computer—built on GB300 Blackwell...  ...like NemoClaw, LLM inference via NIM, Hermes agents, and...  ...looking for a deeply technical systems software engineer who will own AI stack...  ...improve kernel fusion, graph execution, operator scheduling, and memory... 
    Full time
    Local area

    Nvidia

    Santa Clara, CA
    9 hours ago
  • $152k - $241.5k

     ...about redefining how software is built in the age of Generative AI? Join NVIDIA’s TensorRT...  ...for out-of-framework inference globally. We are...  ...unprecedented scale.If you are a systems-thinking C++ engineer who wants to help...  ...solutions.Pragmatic execution: Demonstrated ability... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $272k - $431.25k

     ...Principal Systems Software Engineer at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency...  ...026.This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to... 
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $140k - $240k

    Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs...  ...training and inference speeds; over 10 times...  ...security-first based engineering. Cerebras cluster involves...  ...cluster management software stack - all the way...  ..., maintain and execute roadmap of the... 

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  •  ...computing experiences—from AI and data centers, to...  ...gaming and embedded systems. Grounded in a...  ...challenges—striving for execution excellence, while being...  ...: We are seeking a Principal Software Engineer to serve as the senior...  ...— LLM training and inference (PyTorch, vLLM, Triton... 
    Contract work
    Shift work

    AMD

    San Jose, CA
    2 days ago
  • Cerebras Systems builds the world's largest AI chip, 56 times larger than...  ...leading training and inference speeds; over 10...  ...RoleWe're hiring a Principal Engineer for our Inference...  ...tradeoffs, and drive execution without a clear...  ...of experience in software engineering, with... 

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs...  ...training and inference speeds; over 10 times...  ...'re hiring a Staff Engineer to own major areas...  ...and customer tiers.Execution on Critical Paths....  ...years of experience in software engineering, with substantial... 

    Cerebras Systems

    Sunnyvale, CA
    2 days ago
  • $117.7k - $221.4k

     ...productivity, building systems that make it easier to...  ...for embodied AI systems. We believe the...  ...model reflects how Cola engineers think: build durable intermediate...  ..., featurization, and inference foundations that power...  ...through strong execution, thoughtful tradeoff analysis... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  • $224k - $356.5k

     ...generation personal AI supercomputer—a...  ...developers, and AI engineers. As NVIDIA brings...  ...firmware, BMC, and AI software teams, collaborate...  ...both operating systems.What you’ll be doing...  ...fine-tuning, and inference—and work with the...  ...ideas with a strong execution bias. Expect to be... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $147k - $237.5k

     ...Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of...  ...building intelligent systems that fundamentally...  ...incidents at scale. As a Principal Software Engineer, you will own the...  ...performance and reduce inference costSolid skills in... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    2 days ago
  • $224k - $356.5k

     ...upon which every new AI-powered application...  ...a deeply technical software manager to lead production AI inference for NVIDIA...  ...optimized inference engines, model profiles/recipes...  ...can run production execution without managing from...  ...‑scale distributed systems, and security hardening... 

    Socket.dev

    Santa Clara, CA
    3 days ago
  • $152k - $241.5k

    We are now looking for a Senior Software Engineer for Deep Learning Inference! Would you like to make a big impact in...  ...the crowd:Experience developing System Software.Proficiency in Python as well...  ...an existing vacancy. NVIDIA uses AI tools in its recruiting processes.NVIDIA... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $152k - $241.5k

    NVIDIA is the platform upon which every new AI‑powered application is built. We are seeking a Senior Software Engineer - AI Inference to advance open‑source LLM serving by...  ...they run best‑in‑class on NVIDIA GPUs and systems-and by improving the underlying stack that... 
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    4 days ago
  • $272k - $431.25k

     ...the unlimited potential of AI to define the next era of computing...  ...MODS organization seeks a Principal Engineer to architect and scale next-...  ...L10 and L11 diagnostic systems for Cloud Service Providers...  ...distributed systems and hardware / software interfaces is essential for... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  •  ...library. He/she will participate in the core system design and development. Our target...  ...proven records on infrastructure level software development experience ~2+ years Clojure...  ...Regards Parinita Bhintade HR Executive Integrated Resources , Inc. IT Life... 
    Full time

    Integrated Resources Inc.

    Santa Clara, CA
    13 hours ago
  • $272k - $431.25k

     ...and hands-on delivery across system software, drivers, and CUDA to make...  ...integrate with existing ML/AI workflows (e.g., PyTorch/XLA...  ...technical direction for an engineering team; mentor engineers, drive...  ...experience tuning ML training/inference loops based on deep profiling... 
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $170k - $277k

     ...Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of...  ...TeamEngineering - Our engineering team is at the core of...  ....Job SummaryAs a Sr. Principal Software Engineer in the Engineering...  ...and developing systems to solve complex problems... 
    Full time
    Work at office

    Palo Alto Networks

    Santa Clara, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal System Software Engineer, AI Inference Execution. Be the first to apply!