System Software Engineer, Node & Cluster Management
$160k - $275kMatX
What MatX Is BuildingMatX's mission is to make the world’s best AI models run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability. MatX is seeking System Software Engineer to join our team as we create best-in-class silicon for high-performance and sustainable GenAI. Successful candidates for these roles will be responsible for delivering performant and functionally accurate silicon for MatX products across compute, memory management. High-speed connectivity and other key technologies.The MatX host system software team owns everything that makes our AI silicon and systems usable: from Linux kernel drivers up through node and cluster management. The team also co-owns the BMC/OpenBMC firmware stack, with dedicated firmware engineers, so host software and out-of-band management are designed together rather than bolted together. We're looking for self-driven engineers who can take a hardware spec and a register map and just start building — prototype drivers, low-level utilities that talk directly to the chip, daemons, and tooling — with minimal hand-holding. Each engineer on this team has a primary focus area, but ownership of overlapping components is shared, and you should expect (and want) to venture across the stack.What You'll Do HereDesign and build the node-level management plane for MatX's AI systems: expose node health, inventory, telemetry, and control operations through endpoints (e.g., Redfish-style or custom APIs)Design and implement cluster management solutions and failover algorithms to minimize downtimeBuild the management CLI utilities that operators and internal engineers use daily — interacting with the on-node management and telemetry daemons to query state, run diagnostics, update firmware, and recover devicesPartner with our BMC firmware engineers to present unified management and observability across in-band and out-of-band paths — so operators see one coherent node, whether data comes from the host daemons or the BMC (e.g., unified Redfish-style views, firmware update orchestration across host and BMC, and recovery flows that work even when the host is down)Extend node-level capabilities to cluster level: fleet-wide health aggregation, device inventory, alerting hooks, and integration points for our customers' own fleet-management systemsGet hands-on with the low-level stack: you'll regularly need to drop below the API layer — into the telemetry daemon, driver interfaces, or raw device access utilities — to prototype, debug, or unblock yourselfBuild tooling and automation for managing lab systems during bring-up: provisioning, test orchestration, regression monitoringDefine the software contracts between the on-node daemons, the BMC stack, and the management layer — shared-ownership boundaries you'll co-designDebug production-grade issues spanning management APIs, daemons, kernel drivers, BMC firmware, and hardwareHelp shape what "manageable at scale" means for a new hardware platform, from single node to full rack to clusterWho You AreBS or higher in Computer Science, Electrical Engineering, or equivalent practical experience, with 8+ years in systems software — this is not a pure web-services role; deep low-level systems experience is requiredStrong hands-on Linux systems development experience, including low-level userspace software; comfortable reading and debugging kernel driver and daemon codeStrong programming skills in C plus a systems language suited to services and tooling (Go, Rust, C++, and/or Python)Experience designing and building APIs and CLI tools for hardware or infrastructure managementSolid understanding of how the pieces underneath your APIs actually work — device drivers, telemetry paths, PCIe device behavior, BMC-managed subsystems — and the instinct to go look when something misbehavesExperienced debugging across API, daemon, kernel, firmware, and hardware boundariesComfortable working with firmware engineers to align host-side and BMC-side management capabilities behind common interfacesSelf-driven and pragmatic: able to stand up a working management endpoint against brand-new hardware with minimal specificationBonus Points If You HaveExperience with Redfish, OpenBMC, gNMI, IPMI, or other datacenter hardware management standardsCluster/fleet management experience for GPU or accelerator infrastructureExperience with hardware bring-up, lab automation, or manufacturing/qualification test infrastructureFamiliarity with firmware update orchestration, secure boot, or attestation flowsCompensationThe US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job-related skills, and relevant education and training. Career length is only a guideline for compensation.Early Career - $160,000 - $275,000 + equityMid Career - $175,000 - $400,000 + equitySenior Career - $250,000 - $600,000 + equityWhat We OfferTime off: 4 weeks PTO (accrued) + 12 company Holidays + up to 3 weeks remote workHealth: Company-subsidized Medical (Kaiser or Anthem) for employees & dependents, Guardian Dental and Vision insurances for employee & dependents, and life insurance (employee only), plus HSA and FSA offerings via Lively. Financial Wellbeing: Choose from Roth IRA/ 401K (or both) retirement plans with up to 5% company contribution to 401K (even if you don't contribute). Also, 100% company-paid life insurance (up to $300K) and long-term disability insurances.Professional Development: $1500 Professional Development Budget (per year)Team Meals: MatX provides onsite team lunch & dinner Monday - Friday, with your choice of ordering via WeBox, Specialty’s or via our reimbursement systemCommute on Us: Commute on our company Uber account, or reimburse your train rides. Either way, we pay 100% for your daily commute.MatX E[x]tras: $50/mo to use on the perk you value mostCell & Internet Reimbursement: $35/mo for cellular and $40/mo for wifiMental Wellbeing: 100% paid mental health benefit via SpringHealth and Guardian EAP.Support to Parents: Up to 12 weeks paid parental leave regardless of path to parenthood, 10 weeks pregnancy disability leave, flexible return-to-work hours, and Benepass reproductive health & parental benefit.AI Resources: Up to $20K/month plus a dedicated internal AI Tooling Team to support your productivityAs part of our dedication to the diversity of our team and our focus on creating an inviting and inclusive work experience, MatX is committed to a policy of Equal Employment Opportunity and will not discriminate against an applicant or employee on the basis of race, color, religion, creed, national origin or ancestry, sex, gender, gender identity, gender expression, sexual orientation, age, physical or mental disability, medical condition, marital/domestic partner status, military and veteran status, genetic information or any other legally recognized protected basis under federal, state or local laws, regulations or ordinances.All candidates must be authorized to work in the United States and work from our offices in Mountain View Tuesdays-Thursdays.This position requires access to information that is subject to U.S. export controls. This offer of employment is contingent upon the applicant's capacity to perform job functions in compliance with U.S. export control laws without obtaining a license from U.S. export control authorities.MatX does not accept unsolicited resumes from individual recruiters or third-party recruiting agencies in response to job postings. No fee will be paid to third parties who submit unsolicited candidates directly to our hiring managers or People team and any resumes submitted are deemed to be the property of MatX.LocationMountain View (HQ)Employment TypeFull timeLocation TypeHybridDepartmentHardware
$140k - $240k
Cerebras Systems builds the world's largest AI chip, 56 times... ...czar for the Cerebras’s AI cluster product. Such AI clusters have... ..., security-first based engineering. Cerebras cluster involves... ...vertically integrated cluster management software stack - all the way from a...Suggested$184k - $287.5k
...fresh hardware and software innovations to... ...Our team of skilled engineers is committed to addressing... ...for a Senior Systems Software Engineer... ...in Kubernetes node engineering, OS image... ...packaging, and nodepool management. They must have... ...to maintain cluster reliability at frontier...SuggestedFull timeWorldwide$175k - $275k
Cerebras Systems builds the world's largest AI chip, 56 times larger... ...RoleAs part of the Embedded Software team, you will help build the... ...the Cerebras Wafer Scale Engine (WSE)—the world’s largest AI... ...the WSE’s system software to cluster-level orchestration, collaborate...Suggested$224k - $356.5k
...workloads globally, our diagnostic systems need to evolve across diverse... ...technical leader to engineer and propel innovation in diagnostics... ..., which involve hardware and software tools to develop the worst... ...across rack-level or cluster-level deployments.Background...SuggestedFull time$152k - $221k
...(OKRs), and resourcing while utilizing data-driven insights to manage risks and facilitate executive-level stakeholder decision-making... ...requirements between UX researchers, product managers, and engineers to resolve constraints.Evaluate engineering feasibility, milestones...Suggested$108k - $162k
...ResponsibilitiesWe are seeking a highly skilled Sr. Systems & Infrastructure Engineer to join a dynamic, security-first IT... ..., hybrid cloud, and modern cloud-managed environments. This role spans... ...VMware vSphere/ESXi architecture, cluster operations, lifecycle management,...Permanent employmentFull time$175k - $225k
...Senior IT Systems EngineerSunnyvale, CAThe future of defense will... ...IT Infrastructure & Security Engineer to design, build, and operate... ...data segmentation to device management and IT operations.Your primary... ...devices; maintain OS patches, software updates, and hardware refresh...Full time$120.5k - $243k
...System Software Engineer This role has been designed as ‘’Onsite’ with an expectation that you will primarily work from an HPE office. Who... ...backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together,...Work experience placementWork at officeLocal areaImmediate start$266.1k - $342.9k
...technology is hyper-relevant in the age of AI. As Senior Manager of Solutions Engineering for Amazon, you are more than a leader of leaders; you are... ...massive-scale data center interconnect (DCI), and AI/ML cluster networking. Knowledge of the entire networking device stack...Full timeTemporary workLocal areaFlexible hours$127k - $182k
...Gen AI products and Gemini APIs.Collaborate closely with engineering and product management to shape future product roadmaps, translating partner... ...specialists that support Google's GenAI and open source software products.We are committed to building an ever more fair...$160k - $275k
...quality and availability. MatX is seeking System Software Engineer to join our team as we create best-... ...products across compute, memory management. High-speed connectivity and other... ...from Linux kernel drivers up through node and cluster management. The team also co-owns the...Daily paidFull timeWork experience placementLocal areaMonday to FridayFlexible hours$157k - $210k
...Infrastructure Capacity Program Manager, NPI to join our... ...Product Introduction (NPI) and Engineering Operations teams — tracking data... ...decisions, ensuring a clear system of record as hardware generations... ..., project/portfolio tracking software) is highly desirable. Bachelor...Permanent employmentFull timeTemporary workCasual workWork at officeRemote workFlexible hours- ...Work with Program & Product Management, technical leads, and product... ....), Windows Server operating systems, Windows Client operating systems... ...for approval. Implement software solutions for multiple test programs... ...test execution to test engineers at various global locations....Local areaRemote work
$146.3k - $306.4k
Defines architecture for large-scale systems software, firmware integration, and fleet automation that... ...performance, and cost at hyperscale. Champions engineering excellence: coding standards, threat modeling, resource management, concurrency, and fault isolation. Leads...Temporary workFlexible hoursShift work$184k - $287.5k
...platformsImplement power and thermal management software features in Linux Kernel and user spaceCollaborate... ...architects, hardware and software engineers on platform power estimation and... ...and power consumption, and tuning system-level performance to deliver reliable...Full time- ...Experienced and strategic Cloud Deployment Infrastructure Program Manager to lead our cloud infrastructure projects and programs. This... ...Direct and coordinate cross-functional teams of cloud architects, engineers, developers, and other stakeholders to ensure effective project...
- ...Principal Systems Software Engineer Oracle Cloud Infrastructure (OCI) is seeking a Principal Systems Software Engineer to help build and evolve the low-level systems software and platform-management capabilities that power next-generation GPU infrastructure. This...
$266.1k - $342.9k
...Technical Leader And CTO-Equivalent You will lead a team of elite Systems Engineers and serve as the senior technical leader and CTO-equivalent... ...Familiarity with AI/ML fabric design and large-scale GPU cluster networking Understanding of disaggregated/open networking...Full timeTemporary workLocal areaFlexible hours$160k - $225k
...transforming the diagnosis and management of patients with serious... ...neurological conditions. The Ceribell System is a novel, point-of-care... ...-on Senior or Staff Firmware Engineer / Embedded Engineer who is... ...system requirements into reliable software designs.What you’ll need to...Local area- ...technologies—like the da Vinci surgical system and Ion—have transformed how... ...worldwide.We’re a team of engineers, clinicians, and innovators... .... Digital Learning Strategy Manager will design, create, and... ...digital infrastructure, including software upgrades, product...Local areaRemote workWorldwideFlexible hours
$207k - $300k
...and analyze computer systems and their interactions... ...kernel, hardware, cluster software, and compute infrastructure... ...in a people management or team leadership role... ...influencing large groups of engineers and driving large-... ...internationally.The Node Performance team bridges...$184k - $287.5k
We are seeking a Sr System Software Engineer to help us build out our scientific computing platform... ...architecture, message passing (MPI, NCCL), Cluster scalability and performance.Hands on... ..., Scheduling, IPC, Memory management, File system and I/O structure.Strong...Full time$193.3k - $261.5k
...designs custom SoCs (System on Chips) that power the... ...and inference clusters. Our organization builds... ...SoCs and the low-level software stack that brings these... ...for a Systems Software Engineer who wants to work at the... ...reviews, source control management, build processes,...Local areaFlexible hours$152k - $241.5k
...motivated Performance engineer to influence the roadmap... ...NVLink, PCIe) within a node and with high-speed networking... ...-GPU and multi-node clusters.Study the interaction... ...of computer system architecture, HW-SW interactions... ...principles (aka systems software fundamentals)Implement...Full time$184k - $287.5k
...cutting‑edge hardware and software innovation to deliver... ...of forward‑thinking engineers tackling some of the globe... ...searching for a Senior Systems Software Engineer with... ...Operator, Network Operator, node-feature-discovery,... ...operation at hyperscale cluster sizes, doing in the...Full timeRemote work$184k - $287.5k
...looking for an outstanding software engineer to apply their skills in the... ...other libraries and database systems. In this role, you will research... ...ML algorithms for clustering and visualization. This is a... ...designAbility to work independently and manage your own development efforts...Full time$175k - $215k
...you will report to IT Network Operations Manager You will: Leadership : Manage a high-performing, inclusive team of engineers, technical professionals, and vendor agents... ...-functional teams to design automation systems, reduce operational toil, and onboard new...Full timeRemote work- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.What You’ll... ...maintain bare-metal Kubernetes clusters, scaling up to thousands of...Work at officeLocal areaWork from homeFlexible hours
$127k - $182k
...program charters across multiple cross-functional teams, rigorously managing program scope, timelines, OKRs, resourcing, dependencies, and... ...User Experience Research(UXR), Program Managers (PMs), and engineering partners by performing technical due diligence to evaluate...$100 - $122 per hour
...We are seeking a Technical Program Manager (TPM) to lead cross-functional... ...skills, technical depth in network engineering, and a solid understanding of IP capacity... ...with network planning tools, GIS systems, and project management software (e.g., Jira, Asana, MS Project)....Hourly payContract work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to System Software Engineer, Node & Cluster Management. Be the first to apply!
- full stack node.js developer Mountain View, CA
- react node js developer (remote) Mountain View, CA
- react node js developer Mountain View, CA
- embedded software Mountain View, CA
- software applications developer Mountain View, CA
- entry level software sales Mountain View, CA
- software technology Mountain View, CA
- software implementation project manager Mountain View, CA
- software support Mountain View, CA
- government software Mountain View, CA

