GPU System Reliability Engineer Lead
$150k - $225kCowboy Space Corp.
About Cowboy Space Corp:
Mission
Cowboy Space Corporation is solving the global energy crisis by building the infrastructure for abundant, resilient space-based solar energy. We are tackling one of humanity’s most complex engineering challenges with a world-class team dedicated to delivering a revolutionary power platform. Cowboy Space Corporation is transforming how civilization powers, computes and connects - from orbit to Earth.
Background
Cowboy Space Corporation is building the infrastructure to power and connect the orbital economy. Our modular, scalable satellites collect sunlight in Low Earth Orbit to enable multiple integrated applications: transmitting energy via infrared lasers (space-to-earth and space-to-space), powering on-orbit high-performance computing clusters (GPU/TPU), and providing secure, high bandwidth optical data transport.
Current energy and data systems rely on complex logistics and outdated infrastructure. Cowboy Space Corporation overcomes these challenges by enabling direct, on-demand, secure, and scalable energy distribution and data processing from space. This will revolutionize how we operate in orbit and on Earth, supporting the rapidly expanding space industrial base, ISAM (In-Space Servicing, Assembly, and Manufacturing) operations, remote regions, and military bases.
Baiju Bhatt founded Cowboy Space Corporation in 2024. Inspired by his father’s work with NASA Langley Research Center, Baiju earned his B.S. in Physics and M.S. in Mathematics at Stanford before co-founding Robinhood, now a public company that has helped over 20 million Americans access the financial system. Cowboy Space Corporation has raised ~$365 million from Index Ventures, Interlagos, Construct Ventures, Breakthrough Energy Ventures, Andreessen Horowitz, NEA, and others.
This is an ambitious mission that demands extraordinary talent. Cowboy Space Corporation’s team has worked at places like SpaceX, Blue Origin, Stoke Space, Astranis and NASA, and is based in San Carlos, CA. If you're ready to solve complex technical challenges and help build the most important energy company in the world, we want to hear from you.
The Role
Deploying high-performance GPU compute in Low Earth Orbit introduces a fundamentally different fault landscape than ground-based datacenter operation. This role sits at the frontier of that problem. When a fault occurs 500km above Earth, the system must detect it, classify it, contain it, and recover from it autonomously. You will own the end-to-end RAS validation strategy for GPU server systems, working directly with GPU and HBM silicon partners to analyze failures, characterize fault propagation paths, and ensure detection and recovery mechanisms function correctly. The right candidate combines deep knowledge of processor and memory architecture with hands-on system-level validation experience and the ability to drive partner engagements to resolution. This role is located in San Carlos or Seattle.
Key Responsibilities:
Lead RAS validation strategy and execution for GPU server platforms, including fault injection, detection coverage, and recovery verification.
Partner directly with GPU system designers to analyze hardware failures, review silicon errata, and align on fault handling requirements for DDR, HBM, CPU, and GPU subsystems.
Characterize fault propagation paths from hardware detection through firmware and OS layers, and validate that error signals are correctly classified, logged, and acted upon.
Validate BMC and out-of-band management visibility into hardware health events via IPMI, Redfish, and MCTP/PLDM protocols.
Debug complex failure modes spanning GPU and CPU architecture, memory subsystems, PCIe/NVLink fabric, and system management firmware.
Drive root-cause analysis for RAS failures discovered during validation and work with partners to provide input on platform design decisions that affect fault detection and serviceability.
Define RAS coverage metrics and maintain traceability from hardware fault models to test coverage.
Collaborate with firmware, software, and platform teams to validate OS-level error handling, ACPI error interfaces (EINJ, BERT,HEST), and runtime error recovery flows.
Basic Qualifications:
Bachelors degree in Electrical Engineering or a related discipline.
5+ years of experience in hardware validation, platform reliability engineering, or silicon validation on server-class compute systems.
Deep understanding of CPU and GPU architecture, including memory subsystems (DDR, HBM), cache hierarchies, and interconnect fabrics (PCIe, NVLink, XGMI).
Strong knowledge of RAS concepts: error detection and correction (ECC), fault containment, error propagation, machine check architecture (MCA/MCI), and recovery mechanisms.
Hands-on experience with fault injection methodologies at hardware, firmware, and software levels.
Familiarity with system management interfaces including BMC, IPMI, Redfish, and MCTP/PLDM.
Experience working directly with silicon vendors or ODM partners on hardware failure analysis and RAS gap closure.
Strong scripting skills in Python or equivalent for test automation and log analysis.
Compensation and Benefits:
The salary range for this position is $150,000 – $225,000 annually . The actual base salary offered will depend on factors such as job-related skills, experience, qualifications, and internal equity.
Equity in Cowboy Space Corp.
Employees and their eligible dependents may enroll in medical, dental, and vision insurance
401(k) retirement savings plan
Paid time off
10 paid holidays per calendar year
Paid parental leave
Relocation assistance if applicable
Daily lunch in the office and a fully stocked kitchen with beverages and snacks
ITAR Requirements
Export Control Requirement: To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR), applicants must be a U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State. Learn more about ITAR here .
Disclaimer
This job description is a summary of the primary duties and responsibilities of the job and position. It is not intended to be a comprehensive or all-inclusive listing of duties and responsibilities. Contents are subject to change at Cowboy Space Corp.’s discretion.
Cowboy Space Corp. is an equal employment opportunity employer. We consider individuals for employment or promotion according to their skills, abilities and experience. Cowboy Space Corp. is committed to complying with all applicable laws prohibiting discrimination based on race, color, religious creed, age, national origin, ancestry, physical, mental or developmental disability, sex (which includes pregnancy, childbirth, breastfeeding and medical conditions relating to pregnancy, childbirth or breastfeeding), veteran status, military status, marital or registered domestic partnership status, medical condition (including cancer or genetic characteristics), genetic information, gender, gender identity, gender expression, sexual orientation, as well as any other category protected by federal, state or local laws.
- ...add new capacity and deliver reliable, affordable, and sustainable... ...development and field operations for leading Fortune 500 companies, data... ...target market. Reliability engineers will use tools such as FMEA,... ...design engineers based on system and sub-system specs and possess...SuggestedLocal areaFlexible hours
$195k - $269k
..., CASupply Chain Quality and Reliability - Quality and Reliability /Full... ...with a team of world-class engineers with diverse backgrounds, such... ...Hardware and Vehicle Systems Engineering EE teams.Drive the... ...conceptsPersonable with the ability to lead and coach a high-caliber...SuggestedFull timeTemporary workRelocation package- Kforce has a client that is seeking a Site Reliability Engineer - Data & Caching Systems in Foster City, CA. Overview: We are seeking a Site Reliability Engineer to support large-scale AI infrastructure and next-generation data storage initiatives. This team manages petabyte...SuggestedTemporary work
$161.2k - $221.7k
...dispensers and charge handles to mobile all-in-one charging systems — hardware that must work reliably in airports, vertiports, and remote sites around... .... The Role We're looking for a Staff Mechanical Engineer who can lead the development of complex GSE systems end-to-end...SuggestedFull timeContract workTemporary workRemote work- ...Team: Infra Reliability • SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on... ...role for a first-principles Linux engineer. You'll be the final escalation... ...sessions to redesign systems for higher efficiency and scale...SuggestedWork experience placementRemote work
$250k - $350k
...ll work alongside some of the world's leading ML systems engineers, including leaders behind Megatron-... ...to maximize performance, scalability, reliability, and productivity for both engineers... ...execution environments Optimize memory, GPU kernels and communication for maximum...Visa sponsorship- ...add new capacity and deliver reliable, affordable, and sustainable... ...development and field operations for leading Fortune 500 companies, data... ...power and control hardware systems to achieve its best-in-class... ...for an electrical engineer to be part of the team that is...Local areaFlexible hours
$254k - $350k
...model to drive the next generation of autonomous system intelligence.As a Machine Learning and System Optimization Engineer, you will orchestrate and allocate overall... ...:Allocate and distribute system resources (CPU/GPU/interconnect) to various models and inference engines...Full timeTemporary workRelocation package- ...Reliability Engineer – RemoteBright Vision Technologies is a technology consulting and software... ...excellence of large-scale distributed systems in production. As an SRE you will live... ...failure semantics.Demonstrated experience leading incident response and conducting...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship
$195.9k - $293.9k
Principal Electrical Systems EngineerAt PacBio we create the world... ...Principal Electrical Systems Engineer, you will oversee the electrical... ...service to deliver robust, reliable, scalable, and production-... ...integrated instrument designs.Lead system-level troubleshooting...Full timeWork from homeMonday to Friday$204k - $280k
...CAVehicle Development - Vehicle Integration, Testing & Validation /Full-time /On-siteZoox is seeking a Staff Electrical Architecture System Engineer to join our Electrical Architecture team. Led by the Principal Architect, The Zoox Electrical Architecture Team is responsible...Full timeTemporary workRelocation package$105k - $130k
...CX) operations. Fresh vision. Real impact. Come build it with us. Job Description Freshworks is looking for a Lead - Procurement Systems to join the Procurement Operations team. You will be responsible for administration, configuration and maintenance of the...Flexible hours$204k - $280k
...E2E) Feature and Homologation team brings systems engineering focus and discipline to Zoox mobility-as-a... ...and debug support.In This Role, You Will Lead the design of Fail-Operational behaviors while balancing safety, reliability, and operational efficiency. Partner with...Full timeTemporary workWork experience placementRelocation package$130k - $280k
...We AreVerkada is the leading cloud-managed physical... ...with unmatched speed, reliability, and security. By continually... ...RoleAs an Embedded Engineer (Video Streaming), you... ...utilization (CPU/GPU/ISP).What You BringBS... ...experience in embedded systems development, with a focus...Full timeWorldwideWork visaFlexible hoursShift work$397.46k
...delivering scalable machine learning systems that power personalization, pricing,... ...Avatar.We’re looking for a Distinguished Engineer/Technical Director to lead the strategy and technical direction... ...ML pipelines; from data ingestion to GPU training to live inference.Complex...Full timeWork experience placementH1bWork at officeLocal areaVisa sponsorshipMonday to Friday- ...deploy million-qubit, fault-tolerant quantum systems. Quantum computers harness the laws of... ..., and industry teams work directly with leading Fortune 500 companies—including Lockheed... ...Summary:The Senior Mechanical Design Engineer is responsible for designing highly complex...Full timeFor subcontractorRemote workShift work
$135k - $225k
...world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four... ...for an experienced mechanical engineer to own the design of entire systems... ...design reviews and postmortems; lead incident root-cause...Contract workLocal areaWorldwide$160k - $275k
...inference. We are looking for a hands-on Mechanical Engineer to design, develop, and validate mechanical systems for our rack-scale AI platform from compute trays... ...with JDM/ODM partnersKnowledge of liquid cooling reliability standards and leak testing...Daily paidFull timeWork experience placementLocal areaRemote workMonday to FridayFlexible hours$135k - $225k
...world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four... ...the field.As a Mechanical Design Engineer for RF Systems, you will own the... ...parts and production processes.Lead investigations of RF field failures...Contract workLocal areaFlexible hours$185k - $245k
...operate the world’s largest autonomous logistics system, delivering critical supplies quickly and reliably. Today, Zipline operates on four continents,... ...RoleZipline is seeking a senior-level electrical engineer to lead the design and development of high-reliability embedded...Local area$131.6k - $233.7k
...Progress starts with you.Job DescriptionWe are seeking a Staff Systems Engineer, Microsoft Teams, to join Visa’s enterprise-wide subject... ...with a product mentality.The Work itselfServe as one of Visa’s lead Microsoft Teams experts, owning platform strategy, roadmap, architecture...Full timeContract workWork experience placementWork at officeLocal area$194k - $306k
...and Mission Assurance - Platform Safety Engineering and Analysis /Full-time /HybridZoox is on... ...safely deploy such a robotaxi solution. The System Design and Mission Assurance (SDMA) team... ..., and processes. In this role, you will lead the refinement of how Zoox categorizes, estimates...Full timeTemporary workRelocation package$133.4k - $183.5k
...seeking a passionate, motivated and hands-on Senior Test Engineer on the Electro-Mechanical Test Team to aid in characterization, performance, environmental and reliability testing of Environmental Control System comprising of fan, blower, compressor, heat exchanger, refrigeration...Full timeTemporary workWork experience placement$208k - $287k
...autonomous mobility solution for cities. We are seeking an Engineering Manager to lead a verification and validation team that’s evaluating the... ...driving software teams, software tool teams, test operators, systems engineers, proving ground operators, and more to develop,...Full timeTemporary workRelocation package$184k - $250k
...solution for cities and safely deploy such a robotaxi solution. The System Design and Mission Assurance (SDMA) team plays a foundational... ...-functional teams, including software developers, hardware engineers, systems engineers, simulation developers, and safety experts,...Full timeTemporary workRelocation package- ...add new capacity and deliver reliable, affordable, and sustainable... ...development and field operations for leading Fortune 500 companies, data... ...re looking for a reliability engineering leader to build and run the... ...— as scalable, repeatable systems rather than one-off engineering...Local areaRemote workFlexible hours
$170k - $267k
...City, CASystem Design and Mission Assurance - Platform Safety Engineering and Analysis /Full-time /HybridZoox is on an ambitious... ...for cities and deploy such a robotaxi solution safely and reliably. The System Design and Mission Assurance (SDMA) team plays a foundational...Full timeTemporary workRelocation package- ...Systems Engineer – M&A IntegrationBKF is a multi-service infrastructure consulting firm providing civil engineering, construction management... ...and related cloud technologies.Reporting to the M&A Integration Lead and receiving technical direction from the Senior M&A...
- ...Zoox is seeking a Senior Systems Engineer to drive new robot features across vehicle platforms in Foster City, CA. You will apply MBSE-driven... ...implementation outcomes, create verifiable requirements, and lead cross-functional triage to ensure robust, scalable solutions....
$150k - $230k
...Senior Systems Engineer - AI InfrastructureOn Site, Palo Alto, CaliforniaAbout the RoleWe're building... ...-tolerant, high-performance distributed GPU training. You'll work at the... ...that make large-scale GPU training more reliable and efficientDebug complex distributed/concurrent...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to GPU System Reliability Engineer Lead. Be the first to apply!



