Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (High Performance Computing)

$125k - $160k

SpaceX

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal ofenableing human life on Mars. SITE RELIABILITY ENGINEER (HIGH PERFORMANCE COMPUTING) SpaceX HPC is a shared compute platform used across the company — vehicle and structures simulation, machine learning, AI inference, and more. We support every program at SpaceX to design and operate the worlds most advanced rockets and satellites. This role exists to put a real Site Reliability Engineer operating model on these capabilities and accelerating the world class engineering at SpaceX: toil reduction, automation, observability, and a sustainable incident process. We are looking for a Site Reliability Engineer who wants to own everything from Linux machines and our Infrastructure as Code, storage, and user facing applications – the whole ecosystem as a product, not as a ticket queue. You do not need a prior HPC title. You do need production instincts — you have operated real infrastructure, you write code to delete toil, and you care about whether users can actually get work done, not just whether nodes ping. You’ll work alongside HPC systems engineers who design and commission clusters to help make them more reliable and provide world class services for world class engineers. Aerospace experience is not required. We value engineers who treat teammates with fairness and respect, who are self-critical, and who will take ownership of hard production problems.

RESPONSIBILITIES:

Participate in the team's on-call rotation; practice sustainable incident response and blameless postmortems Manage node lifecycle with infrastructure as code: OS images, firmware, configuration management, kernel and driver stack Build observability for both HPC administrators and end users — cluster, node, and storage health for operators, and job/workflow-level signal for the people running work on the platform Reduce toil with automation; split time between operating production systems and writing the software that makes that work smaller Sustainably manage resources, including compute and storage Lead capacity planning with users across the company: understand what they will need next; and turn that into a concrete picture of tomorrow's compute and storage Collaborate with HPC systems engineers and with engineers across all disciplines across the company on operable, maintainable infrastructure

BASIC QUALIFICATIONS:

Bachelor's degree in computer science, engineering, math, or a scientific discipline; OR 2+ years of professional experience operating production infrastructure in lieu of a degree 2+ years of experience with Linux operating systems in production 2+ years of experience operating production infrastructure (servers, services, or networks), including monitoring, debugging, and repairing what you own

PREFERRED SKILLS AND EXPERIENCE:

2+ years of professional experience in SRE, DevOps, or production infrastructure engineering Experience with monitoring and alerting (Prometheus, Grafana, Nagios, or similar) Experience deploying and maintaining configuration management or infrastructure as code (Ansible, Puppet, Terraform, or similar) Experience writing scripts/code (eg. Python or similar languages) to automate common tasks Experience with containers (Docker, Podman, Singularity/Apptainer) Experience with Kubernetes administration for on-premise deployment Experience with distributed or high-performance storage (VAST or similar), including capacity, performance, and lifecycle management Familiarity with HPC clusters, schedulers (Slurm, PBS, LSF), or GPU compute — not required; we will teach this Familiarity with scientific computing (CFD, FEA) and/or ML training workloads (PyTorch, TensorFlow, CUDA) and/or AI inference workloads Good understanding of version control, testing, continuous integration, build, deployment and monitoring Ability to communicate clearly with users, peers, and vendors in both incident and design settings Comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities Eligibility for access to classified material up to TS/SCI with polygraph

ADDITIONAL REQUIREMENTS:

Position is based in Hawthorne, CA and is primarily on-site Must be able to participate in an on-call rotation Must be willing to work extended hours and weekends as needed for incidents, cluster bring-up, and time-critical failures

COMPENSATION AND BENEFITS:

Pay Range: Level 1: $125,000.00 - $160,000.00 Level 2: $145,000.00 - $195,000.00 Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience. Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law.

ITAR REQUIREMENTS:

To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status. Applicants wishing to view a copy of SpaceX’s Aff…? View email address on click.appcast.io. #J-18808-Ljbffr Spacex

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (High Performance Computing) in Hawthorne, CA vacancy
  • $120k - $145k

     ...ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER, HIGH PERFORMANCE COMPUTING Starshield leverages SpaceX’s Starlink technology and launch...  .... The Starshield software team is building highly reliable in-space mesh networks, designing secure systems to guarantee... 
    Performance
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $165k - $265k

     ...developing the technologies to make this possible, with the ultimate goal ofenabling human life on Mars. SR. HPC SYSTEMS ENGINEER (HIGH PERFORMANCE COMPUTING) the company is looking for an HPC Systems Engineer with strong knowledge and experience in a world class... 
    Performance
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    United States Digital Space LLC

    El Segundo, CA
    4 days ago
  • $165k - $265k

     ...enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging...  ...to deploy and manage on-premise compute resources, create highly scalable and maintainable...  ...such as Bazel and MakefilesFocus on performance bottlenecks and performance improvement... 
    Performance
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    1 day ago
  • $140k - $180k

     ...Title: Senior Site Reliability Engineer Location: El Segundo, CA (On-Site)...  ...mindset focused on building highly available, scalable, and resilient...  ...also be responsible for performance optimization, reliability...  ...Bachelor's degree in Computer Science, Information Systems... 
    Performance
    Contract work

    ConsultNet

    El Segundo, CA
    1 day ago
  • $170k - $190k

     ...HiveWatch is seeking a Senior Site Reliability Engineer to join our Platform Team,...  ...You’ll ensure exceptional performance, reliability, and...  ...architectures ~ Bachelor's degree in Computer Science, Engineering, or...  ...specifically Background in high‑availability, multi‑tenant... 
    Performance
    Flexible hours

    Saasventurecapital

    El Segundo, CA
    3 days ago
  • $160k - $200k

     ...Senior Site Reliability Engineer El Segundo, CA HiveWatch is a tech-forward...  ...ll help ensure exceptional performance, reliability, and...  ...physical security, IoT, or edge computing environments Experience...  ...specifically Background in high-availability, multi-tenant... 
    Performance
    Flexible hours

    Hivewatch

    El Segundo, CA
    5 days ago
  • $153k - $185k

     ...pharmaceutical processing to reliable and economical...  ...looking for bold engineers to help us get there. As a Senior Site Reliability...  ...engineers to deliver highly operable, reliable...  ...risks; perform performance tuning...  ...Bachelor's degree in computer science, engineering... 
    Performance
    Permanent employment
    Full time
    Internship
    Immediate start
    Relocation package
    Flexible hours
    Weekend work

    Varda Space Industries

    El Segundo, CA
    1 day ago
  • $129k - $193.5k

     ...Mission IT pillar supports engineering teams across...  ...is seeking a skilled Site Reliability Engineer with deep expertise...  ..., upgrade, patching, performance tuning, capacity...  ...maintaining very high uptime ~ Ownership...  ...Bachelor’s degree in STEM, Computer Science. or other... 
    Performance
    Full time
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    The Aerospace Corporation

    El Segundo, CA
    3 days ago
  • $140k - $180k

     ...flown, unlocking performance levels previously...  ...class of spacecraft. Engineered to survive the...  ...the success of a high-growth Series D-funded...  ..., deployable, and reliable products   Reduce...  ...'s degree in computer science, Information...  ...Software Engineering, Site Reliability... 
    Performance
    Permanent employment
    Shift work

    K2 Space

    Los Angeles, CA
    28 days ago
  • $130k - $145k

     ...entertainment. The Role The Site Reliability Engineer (SRE) II is responsible for...  ...and incident response to ensure high system availability and performance. This role will collaborate with...  ...(4-year) relevant field such as Computer Science, Information Technology,... 
    Performance
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS

    Los Angeles, CA
    2 days ago
  • $110k - $145k

     ...liaise with product and engineering teams to ensure...  ...platform and product reliability. The ideal candidate...  ...to ensure the health, performance, and availability of the...  ...or Master's degree in Computer Science, Information Technology...  ...a platform engineer, site reliability engineer,... 
    Performance
    Work experience placement
    Local area

    CoSM

    El Segundo, CA
    20 hours ago
  • $181k - $265k

     ...and we’re excited to hire a high performing Staff SRE to join our growing...  ...collaborate closely with engineering leadership, product managers...  ...importance of performant and reliable systems ~ Education - Ideally...  ...in relevant fields such as Computer Science or Software... 
    Performance
    Work at office
    Immediate start
    3 days per week

    Altruist

    Los Angeles, CA
    3 days ago
  • $81.5k - $141.3k

     ...DESCRIPTION: Position Title: Site Reliability Engineer II Team: CRM DevOps...  ...Division. We are seeking a highly skilled and mission-driven...  ...reliability, scalability, performance, and operational excellence...  ...Qualifications Bachelor's in Computer Science, Software... 
    Performance
    Full time
    Remote work
    Shift work

    Abbott

    Los Angeles, CA
    8 hours ago
  •  ...on behalf of a global PC hardware and AI computing products manufacturer in City of...  ...Technical Product Manager will oversee highly technical AI hardware systems from concept to mass production, partnering with engineering teams and OEM/ODM partners. You will need... 
    Performance

    HR on Demand

    Los Angeles, CA
    6 days ago
  • $135.2k - $202.8k

     ...leaders, and innovators. Join us and take your place in space. The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will develop, implement... 
    Performance
    Full time
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    The Aerospace Corporation

    El Segundo, CA
    5 days ago
  • SpaceX is seeking a Site Reliability Engineer for its HPC platform in Hawthorne, CA. You will own the end-to-end reliability of Linux-based infrastructure, IaC pipelines, storage, and user-facing services, collaborating with cross-functional teams to deliver scalable, repeatable... 

    Spacex

    Hawthorne, CA
    4 days ago
  • $107.8k - $162k

     ...work remotely on the remaining days. On-site expectations may evolve over time to support...  ...Operates company's complex high traffic, business critical internet site...  ...including IP and VOIP). Responsible for system performance; supports/troubleshoots network issues and... 
    Performance
    Work at office
    Local area
    Remote work
    3 days per week

    Green Dot

    Los Angeles, CA
    3 days ago
  •  ...Role: Site Reliability Engineering (SRE) Location: Los Angeles, CA Remote position...  ...management skills. We are seeking a highly skilled - Site Reliability Engineer (...  ...focus on reliability, scalability, and performance. Good to have Logic Monitor and... 
    Performance
    Full time
    Remote work

    SARIAN Co

    Los Angeles, CA
    1 day ago
  •  ...About the Role We are seeking a Site Reliability Engineer to join our Ground Software team. As...  ...equally comfortable architecting a highly available Kubernetes platform, and defining...  ...based on merit, qualifications, and performance. We will never discriminate on the... 
    Performance
    Full time
    Work at office

    Apex Space, Inc.

    Los Angeles, CA
    20 hours ago
  • $180k - $200k

     ...created by our award-winning engineering teams. Ateme (PARIS:...  ...we enable clients to deliver high-quality video experiences on...  ...events and ensure proper system performance during critical operations....  ...international teams that value reliability, innovation, knowledge... 
    Performance

    ATEME

    Culver City, CA
    21 hours ago
  • $28 - $35 per hour

     ...Moog is a performance culture that empowers people to achieve...  ...: Intern, Software Engineering Reporting To:...  ...you: ~ Enrolled in a Computer Engineering, Computer...  ...activities. ~ Additional site-specific benefits may...  ...the low and high end of the Moog salary... 
    Performance
    Hourly pay
    Full time
    Internship
    Summer internship
    Relocation
    Relocation package
    Flexible hours

    Moog Inc.

    Torrance, CA
    8 hours ago
  • $32 - $35 per hour

     ...Aerospace Corporation is hiring Software Engineering Grad Interns for the Software...  ...system technologies, such as elastic compute clouds, containerization, microservices...  ...development to deliver responsive, resilient, high-performance software intensive systems to our... 
    Performance
    Hourly pay
    Full time
    Temporary work
    Internship
    Work at office
    Immediate start
    Remote work
    Relocation package
    Flexible hours

    The Aerospace Corporation

    El Segundo, CA
    8 hours ago
  •  ...actively seeking a Principal Engineer for an immediate full-time...  ...company is seeking experienced Site Reliability Engineers to take ownership...  ...environments. This is a high-impact role focused on ensuring...  ...content delivery and CDN performance Develop and execute capacity... 
    Performance
    Full time
    Contract work
    Immediate start
    Work from home
    Flexible hours

    Kesta IT

    Culver City, CA
    3 days ago
  • $210.5k - $263.1k

     ...the role We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights...  ...the reliability, scalability, performance, and security of Crunchyroll's consumer...  .... You are passionate about building highly reliable platforms, eliminating operational... 
    Performance
    Flexible hours

    Crunchyroll

    Los Angeles, CA
    1 day ago
  • $125k - $160k

     ...human life on Mars. SOFTWARE ENGINEER (FLIGHT RELIABILITY) The Flight...  ...RESPONSIBILITIES: Develop highly reliable software...  ...~ Bachelor's degree in computer science, engineering, math...  ...and improving application performance. ~ Experience designing... 
    Performance
    Permanent employment
    Temporary work
    Remote work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    8 hours ago
  •  ...Job Description Software Engineer – Flight Reliability Inglewood, CA – Relocation...  ...Functions Develop highly reliable software...  ...~ Bachelor's degree in computer science, engineering, math...  ...and improving application performance ~ Experience designing and... 
    Performance
    Relocation package

    Galaxy Technology Hires LLC

    Inglewood, CA
    21 days ago
  • $135k - $175k

     ...human life on Mars. DATABASE RELIABILITY ENGINEER SpaceX is looking for a...  ...anticipate future needs. Perform root cause analysis of...  ...QUALIFICATIONS: ~ Bachelor’s degree in computer science, information systems...  ...tuning databases to provide highly available and performant... 
    Performance
    Permanent employment
    Temporary work
    Remote work
    Flexible hours
    Weekend work

    SPACE EXPLORATION TECHNOLOGIES CORP

    Hawthorne, CA
    20 hours ago
  • $70 - $85 per hour

     ...in Dublin, Ohio (USA) with Engineering Centers in Pune, India and...  ...foster a culture built on high performance, collaboration, continuous...  ...Qualifications: ~ Degree in Computer Science, software engineering...  ...~ Development of secure, reliable Over-The-Air (OTA) update systems... 
    Performance
    Work experience placement

    Goken

    Torrance, CA
    28 days ago
  • $218.52k - $323.08k

     ...Concepts and Enterprise Engineering (ACE), supporting Blue...  ...to transform how compute infrastructure is deployed...  ...will operate with a high degree of autonomy to...  ...design a platform that is performant, operable, and meets...  ...0 - $323,081.85 Other site ranges may differ Culture... 
    Performance
    Permanent employment
    Full time
    Temporary work
    Local area
    Weekend work

    BLUE ORIGIN

    Los Angeles, CA
    2 days ago
  • $126.9k - $158.6k

     ...Title Senior Software Engineer KBR’s National...  ...Solutions team provides high-end engineering and advanced...  ...Location: On-site Travel Requirements:...  ...degree in engineering or computer science, or related technical...  ...optimizing application performance, scalability, and resiliency... 
    Performance
    Contract work
    For contractors
    Work at office
    Local area

    KBR Careers

    El Segundo, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (High Performance Computing). Be the first to apply!