Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (High Performance Computing)

SpaceX

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (HIGH PERFORMANCE COMPUTING)SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process.We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications – the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts — you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just whether nodes ping. You’ll work alongside HPC systems engineers who design and commission clusters to help make them more reliable and provide world class services for world class engineers.Aerospace experience is not required. We value engineers who treat teammates with fairness and respect, who are self-critical, and who will take ownership of hard production problems.RESPONSIBILITIES:Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortemsManage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stackBuild observability for both HPC administrators and end users — cluster, node, and storage health for operators, and job/workflow-level signal for the people running work on the platformReduce toil with automation; split time between operating production systems and writing the software that makes that work smallerSustainably manage resources, including compute and storageLead capacity planning with users across the company: understand what they will need next, and turn that into a concrete picture of tomorrow's compute and storageCollaborate with HPC systems engineers and with engineers across all disciplines across the company on operable, maintainable infrastructureBASIC QUALIFICATIONS:Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree2+ years of experience with Linux operating systems in production2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you ownPREFERRED SKILLS AND EXPERIENCE:2+ years of professional experience in SRE, DevOps, or production infrastructure engineeringExperience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar)Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar)Experience writing scripts/code (eg. Python or similar languages) to automate common tasksExperience with containers (Docker, Podman, Singularity/Apptainer)Experience with Kubernetes administration for on-premise deploymentExperience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle managementFamiliarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute — not required; we will teach thisFamiliarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloadsGood understanding of version control, testing, continuous integration, build, deployment and monitoringAbility to communicate clearly with users, peers, and vendors in both incident and design settingsComfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilitiesEligibility for access to classified material up to TS/SCI with polygraphADDITIONAL REQUIREMENTS:Position is based in Hawthorne, CA and is primarily on-siteMust be able to participate in an on-call rotationMust be willing to work extended hours and weekends as needed for incidents, cluster bring-up, and time-critical failuresCOMPENSATION AND BENEFITS: Pay Range: Level 1: $125,000.00 - $160,000.00Level 2: $145,000.00 - $195,000.00Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience.Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law.ITAR REQUIREMENTS:To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to View email address on click.appcast.io.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (High Performance Computing) in Hawthorne, CA vacancy
  • $125k - $145k

     ...ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER, HIGH PERFORMANCE COMPUTING Starshield leverages SpaceX’s Starlink technology and...  .... The Starshield software team is building highly reliable in-space mesh networks, designing secure systems to guarantee... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Immediate start
    Weekend work

    Spacex

    Hawthorne, CA
    more than 2 months ago
  • $165k - $265k

     ...developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER SpaceX is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering... 
    Performance
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    1 day ago
  • $125k - $150k

     ...with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer...  ...range of systems engineering problems - including High Performance Computing, System performance, networking, manufacturing infrastructure... 
    Performance
    Permanent employment
    Temporary work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $125k - $145k

     ...goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity...  ...trajectory design and optimization, high-fidelity vehicle simulation,...  ...Monte Carlo simulations on our high-performance computing (HPC) cluster, automated data analysis... 
    Performance
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    9 hours ago
  • $165k - $265k

     ...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (...  ...petabyte scale bare metal compute clusters Closely...  ...across all programs to create highly operable, scalable, and maintainable...  ...refinement Focus on performance bottlenecks and performance... 
    Performance
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $165k - $270k

     ...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield...  ...software team is building highly reliable in-space mesh networks...  ...to deploy and manage compute resources both on-premises...  ...Bazel and MakefilesFocus on performance bottlenecks and performance... 
    Performance
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    9 hours ago
  • $165k - $265k

     ...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re...  ...petabyte scale bare metal compute clusters Closely...  ...across all programs to create highly operable, scalable, and maintainable...  ...refinement Focus on performance bottlenecks and performance... 
    Performance
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $145k - $195k

     ...enabling human life on Mars.SITE RELIABILITY ENGINEER (TOP SECRET CLEARANCE)As a...  ...automation to deploy and manage compute resources both on-premises...  ...engineers to create highly scalable, operable and maintainable...  ...of serversKnowledge of performance bottlenecks and performance... 
    Performance
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $153k - $185k

     ...pharmaceutical processing to reliable and economical...  ...looking for bold engineers to help us get there. As a Senior Site Reliability...  ...engineers to deliver highly operable, reliable...  ...risks; perform performance tuning...  ...QualificationsBachelor’s degree in computer science,... 
    Performance
    Permanent employment
    Full time
    Internship
    Immediate start
    Relocation package
    Flexible hours
    Weekend work

    Varda Space Industries

    El Segundo, CA
    4 days ago
  • $125k - $150k

     ...the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (RAPTOR) SpaceX is looking for a Site Reliability Engineer...  ...range of systems engineering problems - including High Performance Computing, System performance, networking, manufacturing... 
    Performance
    Temporary work

    SpaceX

    Hawthorne, CA
    21 hours ago
  • $125k - $160k

     ...life on Mars. SOFTWARE ENGINEER, SITE RELIABILITY ENGINEER (APPLICATION SOFTWARE...  ...to design and build highly operable, maintainable, and...  ...Identify and eliminate performance bottlenecks using measurement...  ...~ Bachelor's degree in computer science, information systems... 
    Performance
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    3 days ago
  • $165k - $270k

     ...human life on Mars. SR. SITE RELIABILITY ENGINEER (STARSHIELD) Starshield leverages...  ...software team is building highly reliable in-space mesh...  ...to deploy and manage compute resources both on-premises...  ...and Makefiles ~ Focus on performance bottlenecks and performance... 
    Performance
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    9 hours ago
  •  ...and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing...  ...and incident response to ensure high system availability and performance. This role will collaborate with...  ...(4-year) relevant field such as Computer Science, Information Technology,... 
    Performance
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS Group

    Los Angeles, CA
    3 days ago
  •  ...We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems...  ...performance issues such as high latency, slow login,...  ...Messaging systems * Kubernetes compute and storage * Provide evidence... 
    Performance

    Sierra Business Solution LLC

    Los Angeles, CA
    9 hours ago
  •  ...companies, despite delivering high-quality care. While...  ...Role We’re hiring a Staff Site Reliability Engineer to define and strengthen how...  ...deployment systems, networking, compute, storage, and other shared...  ..., disaster recovery, and performance engineering; Hands-on and technically... 
    Performance
    Remote work
    Flexible hours

    Pivotal Health

    Santa Monica, CA
    4 days ago
  • $181k - $265k

     ...and we’re excited to hire a high performing Staff SRE to join our growing...  ...collaborate closely with engineering leadership, product managers...  ...importance of performant and reliable systems ~ Education - Ideally...  ...in relevant fields such as Computer Science or Software... 
    Performance
    Work at office
    Immediate start
    3 days per week

    Altruist

    Los Angeles, CA
    3 days ago
  • $160k - $225k

     ...enabling human life on Mars. SR. SOFTWARE ENGINEER, COMPUTER VISION We are looking for exceptional...  ..., and Starlink programs. We deliver high-impact solutions through advanced AI...  ...necessary Ability to travel to other SpaceX sites as needed (up to 20%) COMPENSATION AND... 
    Permanent employment
    Full time
    Temporary work
    Weekend work

    Spacex

    Hawthorne, CA
    more than 2 months ago
  • $130k - $250k

     ...in the world to identify high-impact opportunities to...  ...ensure they can be executed reliably in our factories. The work spans computational geometry, CAD/CAM integrations, high-performance systems, and full-stack...  ...computing, or software engineering with strong mathematical... 
    Performance
    Permanent employment
    Full time
    Local area
    Relocation package
    Flexible hours

    Hadrian Automation

    Los Angeles, CA
    more than 2 months ago
  • $130k - $150k

     ...life on Mars. SECURITY SOFTWARE ENGINEER, APPLIED COMPUTING (STARSHIELD) Starshield leverages...  ...payloads. The Starshield team is building highly reliable in-space mesh networks, designing...  ...and systems Provide guidance and perform testing of AI integrations into new... 
    Permanent employment
    Full time
    Temporary work
    Immediate start
    Flexible hours
    Weekend work

    Spacex

    Hawthorne, CA
    more than 2 months ago
  •  ...on behalf of a global PC hardware and AI computing products manufacturer in City of...  ...Technical Product Manager will oversee highly technical AI hardware systems from concept to mass production, partnering with engineering teams and OEM/ODM partners. You will need... 
    Performance

    HR on Demand

    Los Angeles, CA
    6 days ago
  • $155k - $195k

     ...planet. About the RoleWe are seeking a Site Reliability Engineer to join our Ground Software team. As...  ...equally comfortable architecting a highly available Kubernetes platform, and defining...  ...based on merit, qualifications, and performance. We will never discriminate on the... 
    Performance
    Full time
    Work at office

    Apex Technology

    Los Angeles, CA
    9 hours ago
  • $99k - $225k

    Computer Vision AI EngineerThe Opportunity: Booz Allen is seeking an...  ..., and machine learning (ML) engineering to train, test, deploy, and maintain...  ...or RAPIDs, to optimize the performance of computer vision...  ...Resource page on our Careers site and reviewing Our Employee Benefits... 
    Performance
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    El Segundo, CA
    4 days ago
  • $164k - $270k

     ...Role What You’ll DoOwn the reliability of our robotics systems, from...  ...controls, robotics, and platform engineering teams to bake reliability in...  ...and can deliver with high velocity.CompensationFor this...  ...skills, geographic location, performance, and business or organizational... 
    Performance
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    2 days ago
  • $160k - $225k

     ...enabling human life on Mars. LEAD SOFTWARE ENGINEER, FLIGHT SYSTEMS - TOP SECRET CLEARANCE...  ...direction and delivery of scalable, high-performance applications used to control and test...  ...: ~ Bachelor's degree in computer science, mathematics, computer engineering... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Immediate start
    Weekend work

    Spacex

    Hawthorne, CA
    more than 2 months ago
  •  ...About the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data &...  ...the reliability, scalability, performance, and security of Crunchyroll's consumer...  ...leadership. You are passionate about building highly reliable platforms, eliminating... 
    Performance

    Engg

    Los Angeles, CA
    3 days ago
  • $175k - $285k

     ...hardening on both macOS and Windows.What Will Set You ApartStrong computer science fundamentals. You can reason about systems from the...  ..., certifications, experience, skills, geographic location, performance, and business or organizational needs.Benefits for Full-time EmployeesMedical... 
    Performance
    Permanent employment
    Full time
    Local area
    Remote work
    Flexible hours

    Hadrian

    Los Angeles, CA
    3 days ago
  •  ...models, and integrated customer decision engines will create the most precise weather...  ...~ Degree in Atmospheric Science, Computer Science, or related engineering / physical sciences field ~5+ years writing high-performance Python/C++/Rust for GPU workloads ~ Real... 
    Performance
    Full time
    Remote work

    Mantari

    Torrance, CA
    more than 2 months ago
  • $140k - $250k

     ...SENIOR SOFTWARE ENGINEER (EMBEDDED) Freeform builds AI-native manufacturing systems...  ...factories to operate fully autonomously at high speeds with precision. This entails...  ...speed data acquisition, and custom high performance compute systems. Your software will effectively... 
    Performance
    Full time
    Casual work
    Relocation package
    Flexible hours

    Freeform

    Hawthorne, CA
    more than 2 months ago
  • $120k - $145k

     ...human life on Mars. SOFTWARE ENGINEER, COLLISION AVOIDANCE (...  ...RESPONSIBILITIES: Develop highly reliable and available software...  ...~ Bachelor’s degree in computer science, engineering, math,...  ...profiling and improving application performance Experience with relational... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Weekend work

    Spacex

    Hawthorne, CA
    more than 2 months ago
  •  ...addition to innovative solutions for high-performance milling and the project-focused engineering of special cutting tool...  ...experience and profit from the reliability and quality of our cutting tools...  ...plus Education BA / BS in Computer Science, Electrical Engineering... 
    Performance
    Full time
    Work at office
    Flexible hours

    Scale Search Group

    Torrance, CA
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (High Performance Computing). Be the first to apply!