Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

GPU System Reliability Engineer Lead

$150k - $225k

Cowboy Space Corporation

Cowboy Space Corp. is building the infrastructure to power and connect the orbital economy. Our satellites operate in Low Earth Orbit to collect sunlight and enable a new class of capabilities—from powering on-orbit compute, to transmitting energy via infrared lasers (space-to-earth and space-to-space), powering on-orbit compute to delivering secure, high-bandwidth optical data. By rethinking how energy and data are generated and distributed in space, we’re unlocking entirely new ways to operate both in orbit and on Earth. Cowboy Space Corp. is building the infrastructure to power and connect the orbital economy. Our satellites operate in Low Earth Orbit to collect sunlight and enable a new class of capabilities—from powering on-orbit compute, to transmitting energy via infrared lasers (space-to-earth and space-to-space), powering on-orbit compute to delivering secure, high-bandwidth optical data. By rethinking how energy and data are generated and distributed in space, we’re unlocking entirely new ways to operate both in orbit and on Earth. Founded in 2024 by Baiju Bhatt (co-founder of Robinhood), Cowboy Space Corp. is backed by leading investors and built by a team from top aerospace and defense organizations. We’re moving quickly to solve complex technical challenges and build a new category of space infrastructure. The Role Deploying high-performance GPU compute in Low Earth Orbit introduces a fundamentally different fault landscape than ground-based datacenter operation. This role sits at the frontier of that problem. When a fault occurs 500km above Earth, the system must detect it, classify it, contain it, and recover from it autonomously. You will own the end-to-end RAS validation strategy for GPU server systems, working directly with GPU and HBM silicon partners to analyze failures, characterize fault propagation paths, and ensure detection and recovery mechanisms function correctly. The right candidate combines deep knowledge of processor and memory architecture with hands-on system-level validation experience and the ability to drive partner engagements to resolution. This role is located in San Carlos or Seattle. Key Responsibilities Lead RAS validation strategy and execution for GPU server platforms, including fault injection, detection coverage, and recovery verification. Partner directly with GPU system designers to analyze hardware failures, review silicon errata, and align on fault handling requirements for DDR, HBM, CPU, and GPU subsystems. Characterize fault propagation paths from hardware detection through firmware and OS layers, and validate that error signals are correctly classified, logged, and acted upon. Validate BMC and out-of-band management visibility into hardware health events via IPMI, Redfish, and MCTP/PLDM protocols. Debug complex failure modes spanning GPU and CPU architecture, memory subsystems, PCIe/NVLink fabric, and system management firmware. Drive root-cause analysis for RAS failures discovered during validation and work with partners to provide input on platform design decisions that affect fault detection and serviceability. Define RAS coverage metrics and maintain traceability from hardware fault models to test coverage. Collaborate with firmware, software, and platform teams to validate OS-level error handling, ACPI error interfaces (EINJ, BERT,HEST), and runtime error recovery flows. Basic Qualifications 5+ years of experience in hardware validation, platform reliability engineering, or silicon validation on server-class compute systems. Deep understanding of CPU and GPU architecture, including memory subsystems (DDR, HBM), cache hierarchies, and interconnect fabrics (PCIe, NVLink, XGMI). Strong knowledge of RAS concepts: error detection and correction (ECC), fault containment, error propagation, machine check architecture (MCA/MCI), and recovery mechanisms. Hands-on experience with fault injection methodologies at hardware, firmware, and software levels. Familiarity with system management interfaces including BMC, IPMI, Redfish, and MCTP/PLDM. Experience working directly with silicon vendors or ODM partners on hardware failure analysis and RAS gap closure. Strong scripting skills in Python or equivalent for test automation and log analysis. Compensation And Benefits The salary range for this position is $150,000 – $225,000 annually . The actual base salary offered will depend on factors such as job-related skills, experience, qualifications, and internal equity. Equity in Cowboy Space Corp. Employees and their eligible dependents may enroll in medical, dental, and vision insurance 401(k) retirement savings plan Paid time off 10 paid holidays per calendar year Paid parental leave Relocation assistance if applicable Daily lunch in the office and a fully stocked kitchen with beverages and snacks ITAR Requirements Export Control Requirement: To conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR), applicants must be a U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State. Learn more about ITAR here. Disclaimer This job description is a summary of the primary duties and responsibilities of the job and position. It is not intended to be a comprehensive or all-inclusive listing of duties and responsibilities. Contents are subject to change at Cowboy Space Corp.’s discretion. Cowboy Space Corp. is an equal employment opportunity employer. We consider individuals for employment or promotion according to their skills, abilities and experience. Cowboy Space Corp. is committed to complying with all applicable laws prohibiting discrimination based on race, color, religious creed, age, national origin, ancestry, physical, mental or developmental disability, sex (which includes pregnancy, childbirth, breastfeeding and medical conditions relating to pregnancy, childbirth or breastfeeding), veteran status, military status, marital or registered domestic partnership status, medical condition (including cancer or genetic characteristics), genetic information, gender, gender identity, gender expression, sexual orientation, as well as any other category protected by federal, state or local laws. #J-18808-Ljbffr Cowboy Space Corporation

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the GPU System Reliability Engineer Lead in Seattle, WA vacancy
  • $135.2k - $306.4k

     ...hardware platform development engineering is seeking a highly driven GPU/CPU Platform System Engineer at the Principal Engineer...  ...team of talented engineers who lead the development and day-to-day...  ...scale, efficiency, reliability and velocity. What This Role Looks... 
    Suggested
    Temporary work
    Work experience placement
    Remote work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    2 days ago
  •  ...ready for the next chapter in your Uncarrier journey? The System Reliability Engineer (SRE) improves and protects the software and systems behind...  ...Security and Access Management (ISAM) organization is leading a major transformation in logical access compliance. What... 
    Suggested
    Full time
    Temporary work
    Part time
    Work experience placement
    Local area
    Flexible hours

    T-Mobile

    Bellevue, WA
    2 days ago
  • Blue Origin in Seattle, WA is seeking a Mechanical Engineer to join the New Glenn program. You will design, develop, and qualify mechanical hardware for cryogenic and high‑pressure systems, deliver drawings and analyses, and support testing and supplier coordination. The... 
    Suggested

    Blue Origin LLC

    Seattle, WA
    13 hours ago
  • Jacobs is seeking a Mechanical Engineer in Bellevue, Washington, to lead mechanical and plumbing system designs for Federal projects. You will manage project timelines, budgets, and collaborate with various disciplines while ensuring compliance with applicable codes and... 
    Suggested

    Jacobs

    Bellevue, WA
    13 hours ago
  • Blue Origin's Lunar Permanence team seeks a senior Mechanical Engineer to lead design, analysis, testing, and review of safety-critical hardware for lunar transportation systems. You will guide a cross-functional mechanical team and own major design decisions from concept... 
    Suggested

    Blue Origin

    Seattle, WA
    13 hours ago
  • A leading design firm in Seattle is seeking a Senior Mechanical Engineer to lead the design of mechanical systems for various sectors including healthcare. Responsibilities include managing project teams, producing engineering drawings, and collaborating closely with architects... 

    DLR Group

    Seattle, WA
    4 days ago
  • Google Seattle, WA, USA is seeking an Software Engineering Manager, GPU Reliability to lead a team of engineers responsible for the reliability and performance of Google's GPU infrastructure and AI platforms. You will own technical roadmaps, mentor engineers, and influence... 

    Google Inc.

    Seattle, WA
    13 hours ago
  • NVIDIA in Seattle, WA seeks outstanding AI systems engineers to advance the inference software stack for AI workloads. You will develop libraries, code generators, and GPU kernel technologies for NVIDIA hardware, including new abstractions and runtimes for large language... 

    NVIDIA

    Seattle, WA
    4 days ago
  • CoreWeave is leading the AI cloud, delivering secure, high‑performance GPU and runtime infrastructure for multi‑tenant AI workloads in Bellevue, WA. The role focuses on building sandboxed environments, integrating container runtimes with virtualization, and optimizing latency... 

    CoreWeave

    Bellevue, WA
    2 days ago
  • $165k - $242k

    CoreWeave in Bellevue, WA is seeking an experienced engineer to lead design reviews and optimize distributed systems. The role requires 5-8 years in cloud services, strong skills in Python or Go, and hands-on experience with Kubernetes. Key responsibilities include defining... 
    Flexible hours

    CoreWeave

    Bellevue, WA
    1 day ago
  • $147k - $202.4k

     ...seeking a highly skilled Senior Site Reliability Engineer to join our team. We are a SaaS company...  ...specializing in securing large-scale systems. This role is a blend of software engineering...  ...responder for critical incidents, leading root cause analysis and implementing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    Bellevue, WA
    1 day ago
  • Agile Space Industries, Inc. is seeking an experienced Systems Engineer to lead systems engineering across rocket engine and thruster programs. You will translate mission needs into requirements, drive architecture, and guide verification from concept to delivery. The... 

    Agile Space Industries, Inc.

    Seattle, WA
    1 day ago
  • $151.2k - $204.6k

    Amazon Prime Air is seeking a Sr. Systems Development Engineer to serve as our internal DO-178C Designated Engineering Representative (DER) and certification...  ...processes that satisfy SOI-1 through executing and leading SOI-2, SOI-3, SOI-4, and Conformity Audits.Your impact... 
    Flexible hours

    Amazon

    Seattle, WA
    5 days ago
  • Blue Origin in Seattle is looking for a Systems Engineer to join their Blue National Security team. You'll work on crucial spaceflight challenges, ensuring safe human spaceflight with a collaborative atmosphere. Ideal candidates will have a B.S. in engineering and 5+ years... 

    Blue Origin

    Seattle, WA
    13 hours ago
  • $130.2k - $208.2k

     ...operations in Sumner, WA. This role helps power REI’s day-to-day work by ensuring employees have reliable, secure, and easy-to-use technology wherever they are. The Systems Engineer Lead designs and supports the tools and platforms behind our Windows and macOS devices—making... 

    Recreational Equipment, Inc.

    Seattle, WA
    13 hours ago
  • Lambda in Bellevue, WA seeks a software engineer to develop storage systems, requiring a Bachelor's or Master's in CS and 5+ years of experience....  ...including Docker and Kubernetes. Join a team that values reliability and scalable architectures, offering health, dental,... 

    Lambda Corporation

    Bellevue, WA
    4 days ago
  • A leading social media platform based in Seattle is seeking a Site Reliability Engineer for its U.S. Data Security division. The role involves developing automation procedures for system efficiency, collaborating with software engineering teams, and ensuring system scalability... 

    TikTok

    Seattle, WA
    1 day ago
  •  ...become the standard.Role ScopeOwn reliability for named customer workloads:...  ...recurring customer pain into engineering fixes with the production...  ...level.You debug distributed systems methodically across layers you...  ...up until they do.Bonus: GPU training workloads. InfiniBand... 

    Fluidstack

    Seattle, WA
    4 days ago
  • $191k - $253k

    Anduril Industries is looking for an experienced Systems Engineer to head the integration engineering practice for acquisition processes. You will be responsible for the entire lifecycle of integrating newly acquired businesses, including assessments and implementation... 

    Slope

    Seattle, WA
    4 days ago
  • Gravitics Inc in Seattle, WA is seeking a Senior Systems Engineer to support the commercial spacecraft program. You will work with teams across engineering and integration to ensure spacecraft functions as an integrated system. The ideal candidate has 5+ years of aerospace... 

    Gravitics Inc

    Seattle, WA
    13 hours ago
  •  ...technology company based in Seattle is seeking a Senior Firmware Engineer to design and implement firmware for its Generation 2...  ...This role involves developing low-level drivers and ensuring system reliability and performance under real-world conditions. The ideal candidate... 

    AIM

    Seattle, WA
    1 day ago
  • $126.2k - $264.1k

     ...a Senior Principal Network Reliability Engineer, you will:Provide technical...  ...infrastructure scalability.Lead complex engineering initiatives...  ...automation solutions.Drive systemic improvements that eliminate...  ...supporting AI infrastructure, GPU networking, or large-scale data... 
    Temporary work
    Flexible hours

    Oracle Corporation

    Seattle, WA
    3 days ago
  •  ...robotics ecosystem, we build risk-aware, reliable, field-ready AI systems that solve the hardest problems in...  ...architectures, combining rigorous engineering with learning systems proven in...  ...performance across distributed CPU and GPU environments, improving throughput,... 
    Local area

    FieldAI

    Seattle, WA
    8 days ago
  • We. is seeking a Senior HRIS Analyst - Workday to join our Global HR Systems team. You will configure, optimize, and support the Workday platform, partnering with HR, Payroll, IT, and business stakeholders to ensure data accuracy, system efficiency, and scalable HR processes... 

    We. Communications

    Seattle, WA
    1 day ago
  •  ...working to develop reusable, safe, and low-cost space vehicles and systems within a culture of safety, collaboration, and inclusion. Join...  ...of the vehicle in coordination with Thermal Systems Engineers, Thermal design scope owners, and vehicle subsystem counterpartsBuild... 
    Permanent employment
    Temporary work
    Local area

    Blue Origin LLC

    Seattle, WA
    4 days ago
  • Blue Origin is seeking a senior engineering leader to drive program delivery for propulsion systems in the Engines business unit. You will manage a dedicated team, deliver product elements, and ensure quality while controlling cost and schedule. You will collaborate with... 

    Blue Origin

    Seattle, WA
    1 day ago
  •  ...LLC in the Seattle area seeks an experienced Hardware-in-the-Loop Lead to own the development, build, and checkout of high-fidelity HIL environments for in-space systems. You will lead a team of engineers for test campaigns, ensuring fidelity across real-time simulators... 

    Blue Origin LLC

    Seattle, WA
    13 hours ago
  • Blue Origin LLC is seeking a Senior Thermal Analysis Engineer to lead the system-level thermal analysis for the Thermal Subsystem of the Lunar Permanence program. You will own the delivery of thermal analyses for the spacecraft, integrating submodels from multiple analysts... 

    Blue Origin LLC

    Seattle, WA
    4 days ago
  •  ...in Seattle, Washington. The role involves owning the development of end-to-end AI systems across various platforms, integrating advanced algorithms into production systems, and leading technical projects. Candidates should have over 8 years of experience in building production... 

    Axon

    Seattle, WA
    2 days ago
  • SupportFinity™ in Seattle is seeking a Supervisor for Test System Production to oversee the team responsible for building and maintaining test systems for Starlink satellites. This role involves managing production schedules, team performance, and ensuring product quality... 

    SupportFinity™

    Seattle, WA
    13 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to GPU System Reliability Engineer Lead. Be the first to apply!