Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer — HPC & Automation (Silicon Engineering)

SpaceX

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER — HPC & AUTOMATION (SILICON ENGINEERING)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most advanced broadband internet system. Starlink is the world’s largest satellite constellation and is providing fast, reliable internet to millions of users worldwide. We design, build, test, and operate all parts of the system – thousands of satellites, consumer receivers that allow users to connect within minutes of unboxing, and the software that brings it all together. We’ve only begun to scratch the surface of Starlink’s potential global impact and are looking for best-in-class engineers to help maximize Starlink’s utility for communities and businesses around the globe. We are seeking a motivated, proactive, and intellectually curious engineer who will work alongside world-class cross-disciplinary teams (systems, firmware, architecture, design, validation, product engineering, ASIC implementation). As a Site Reliability Engineer on the Silicon Engineering team you will get the opportunity to design, operate, scale, and automate the high performance computing infrastructure we use to develop the chips powering the world's largest satellite constellation and a global internet service. This position will have a meaningful impact on Starlink silicon by enabling faster design-iterations, simulations, and regression turnaround times that gate how fast our chip teams can ship. RESPONSIBILITIES:Deploy, upgrade, operate, maintain, and scale our suite of clusters and servicesCollaborate with engineers to develop automated, full turnkey solutions for silicon simulation workflows to speed up project timelinesManage our underlying infrastructure as code and use modern observability tools to provide a complete picture of cluster and infrastructure healthOperate the continuous integration pipeline, build and release systems, and version control across the environmentIdentify and eliminate performance bottlenecks using measurement and creative engineeringBASIC QUALIFICATIONS:Bachelor’s degree in computer science, information systems, or an engineering discipline; OR 2+ years of professional experience in system administration, high performance computing, or site reliability engineering1+ years of development experience with Bash, Python, and/or other programming languages1+ years of experience with Linux operating systemsPREFERRED SKILLS AND EXPERIENCE:Familiarity with containerization technologies (i.e. Docker, Kubernetes)Knowledge in computer system concepts (computer architecture, computer organization, operating systems and concurrency)Experience with databases and data modeling (e.g., MySQL, PostgreSQL, SQLite)Networking knowledge of TCP/IPExperience with high performance computing and workload managers (e.g., Slurm, LSF)Experience with Terraform, Ansible, Puppet, or similar automation frameworksExperience building monitoring and alerting as code (e.g., Grafana, Prometheus, custom exporters)Experience with CI/CD automation at scale (e.g., Jenkins, Bamboo, build systems)Experience with infrastructure as code (IaC) tools for managing fleets of serversExperience with using & building REST API clients/serversExperience with enterprise/networked storage automation (e.g., NetApp ONTAP REST API/CLI, NFS)Experience with ASIC design flows and tools (e.g., Cadence, Synopsys, Ansys, Keysight, Siemens)Strong desire to find performance bottlenecks and performance improvement techniquesExcellent communication skills with the ability to communicate with customers, peers, management, etc. in both formal and informal situationsAbility to quickly learn new tools and frameworksInterest in or experience with AI/LLM-assisted tooling (e.g., Grok, Claude Code)ADDITIONAL REQUIREMENTS:Ability to work extended hours and weekends as needed to meet critical milestonesCOMPENSATION AND BENEFITS:Pay Range:Level 1: $125,000.00 - $150,000.00Level 2: $145,000.00 - $175,000.00Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience.Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees in Washington State accrue paid sick time in compliance with state and federal law. Company shuttles are offered to employees for roundtrip travel from select Seattle locations to the SpaceX Redmond office Monday to Friday.ITAR REQUIREMENTS:To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. Learn more about the ITAR here. SpaceX is an Equal Opportunity Employer; employment with SpaceX is governed on the basis of merit, competence and qualifications and will not be influenced in any manner by race, color, religion, gender, national origin/ethnicity, veteran status, disability status, age, sexual orientation, gender identity, marital status, mental or physical disability or any other legally protected status.Applicants wishing to view a copy of SpaceX’s Affirmative Action Plan for veterans and individuals with disabilities, or applicants requiring reasonable accommodation to the application/interview process should reach out to View email address on click.appcast.io.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer — HPC & Automation (Silicon Engineering) in Redmond, WA vacancy
  •  ...actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER — HPC & AUTOMATION (SILICON ENGINEERING)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink,... 
    Suggested
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    2 days ago
  •  ...career. THE ROLE: The AMD AI Group is looking for a Senior Silicon Design Engineer to co-design the hardware-software interface for AMD Instinct...  ...building custom tooling on top of them.Publications or patents in HPC, ML systems, or GPU kernel optimization.PREFERRED ACADEMIC... 
    Suggested

    AMD

    Bellevue, WA
    3 days ago
  • $184k - $287.5k

     ...searching for a highly motivated, technical engineer to join the Tegra system-on-chip (SoC)...  ...for our next-generation SoCs. In both pre-silicon and post-silicon phases of execution....  ...programming and/or GPUs.Experience with HPC or large-scale computing environments.Your... 
    Suggested
    Full time
    Remote work

    Nvidia

    Redmond, WA
    4 days ago
  • $165k - $230k

     ...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology...  ..., and operational support. RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and... 
    Suggested
    Permanent employment
    Temporary work
    Work at office
    Immediate start
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    1 day ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 0 days / week in-office -...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...workloads. As a Site Reliability Engineer II, you will take ownership of...  ...issues, design and implement automation to reduce toil, and contribute... 
    Suggested
    Ongoing contract
    Work at office
    Local area

    Microsoft

    Redmond, WA
    3 days ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...than the Microsoft Defender engineering team. We are looking for a Site...  ...within SLA timelines. Automation & Deployment: Contribute to... 
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...world.Microsoft’s Azure Data engineering team is leading the transformation...  ...and shaping the Livesite Automation and AI Ops stack in Cosmos... 
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    2 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...seeking a Senior Site Reliability Engineer to lead a team that builds and...  ...that deliver production code, automation, and self-healing... 
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    2 days ago
  • $102.1k - $202.2k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...team, you will collaborate with engineers across disciplines to deliver...  ...incident response, and building automation to reduce operational toil.Microsoft... 
    Ongoing contract
    Work experience placement
    Local area
    Remote work
    3 days per week

    Microsoft

    Redmond, WA
    5 days ago
  • $165k - $230k

     ...with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology...  ..., and operational support. RESPONSIBILITIES: Develop automation to deploy and manage compute resources both on-premises and... 
    Permanent employment
    Temporary work
    Work at office
    Immediate start
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    1 day ago
  • $184k - $287.5k

     ...fully optimized NVIDIA AI and HPC software stack. We’re...  ..., and a passion for building reliable, debuggable, and scalable manufacturing...  ...tool design.Drive pre-silicon readiness for factory & manufacturing...  ....Mentor architects and engineering teams to grow them into future... 
    Full time
    Remote work
    Shift work

    Nvidia

    Redmond, WA
    3 days ago
  •  ...Type: Full TimeIndustry: Computer SoftwareClient: WiproContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Post Silicon Validation Engineer NO 16+ yrs profilesLocation: Redmond WAJob Description: Look for profiles with 5 to 12 years Exp only- UVM with C, C++ Coding... 

    SRI Tech

    Redmond, WA
    4 days ago
  • $165k - $230k

     ...enabling human life on Mars.SR. HARDWARE / INFRASTRUCTURE SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building...  ..., to our internal Kubernetes platforms. You will develop automation to deploy and manage on-premise compute resources, create... 
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    4 days ago
  • $119.8k - $234.7k

     ...type: Full-TimeWork site: 3 days / week in-...  ...Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team...  ...infrastructure teams to build reliable, high-performance...  ...patterns. Create and automate network stress, scale...  ...supporting AI, HPC, cloud, or large-scale... 
    Ongoing contract
    Work at office
    Local area
    Worldwide
    3 days per week

    Microsoft

    Redmond, WA
    1 day ago
  • $165k - $260k

     ...of enabling human life on Mars.SR. TECHNICAL PROJECT LEAD (SILICON ENGINEERING)At SpaceX we’re leveraging our experience in building rockets...  ...world’s largest satellite constellation and is providing fast, reliable internet to millions of users worldwide. We design, build,... 
    Permanent employment
    Temporary work
    Work at office
    Worldwide
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    6 days ago
  • $174k - $253k

     ...Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent...  ...Science or Engineering. About The Job Site Reliability Engineering (SRE) is what you get when...  ...systems sustainably through mechanisms like automation, and evolve systems by pushing for... 
    Temporary work

    Google

    Kirkland, WA
    4 days ago
  • $142.8k - $274.8k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringDiscipline: Site Reliability EngineeringCompany:...  ...Principal Site Reliability Engineering Manager to lead a team responsible...  ...production code, build automation and self-healing systems, and... 
    Ongoing contract
    Temporary work
    Fixed term contract
    Local area
    Immediate start
    3 days per week

    Microsoft

    Redmond, WA
    3 hours ago
  • $152k - $241.5k

     ...team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding...  ...full tools lifecycle, initiating from pre-silicon software development through silicon bring...  ...or SoC platforms.Familiarity with test automation frameworks and CI/CD pipelines for hardware... 
    Full time

    Nvidia

    Redmond, WA
    4 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...EngineeringCompany: MicrosoftOverviewMicrosoft Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind...  ...product goals across quality, reliability, and performance. Collaborate... 
    Ongoing contract
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    3 days per week

    Microsoft

    Redmond, WA
    3 days ago
  • $165k - $225.6k

     ...we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site...  ...enablement systems. With a strong focus on automation, testing, and operational excellence, you will... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    3 days ago
  • $45 - $48 per hour

    DescriptionKforce has a client that is seeking an Android Test Automation Engineer in Redmond, WA.Key Tasks:* Develop, maintain, and expand...  ...component-level tests for Android applications to improve reliability, regression coverage, and overall product quality* Maintain... 

    KForce

    Redmond, WA
    3 days ago
  • $184k - $287.5k

     ...looking for a Senior Software Engineer for AI Resiliency!At NVIDIA,...  ...AI systems remain robust and reliable at all times.What You’ll Be Doing...  ...and JAX/XLA.Testing & Automation: Develop and implement tests...  ...scale AI workloads in cloud and HPC environments, ensuring seamless... 
    Full time

    Nvidia

    Redmond, WA
    3 days ago
  • $320k

     ...networking, NVIDIA Grace CPUs, and a fully optimized NVIDIA AI and HPC software stack. NVIDIA NVLink Fusion will enable industry-...  ...across NVIDIA's Software, Architecture, Networking and Systems engineering teams in defining the architecture for NVLink Fusion. Ensuring... 
    Full time
    Shift work

    Nvidia

    Redmond, WA
    4 days ago
  • $152k - $241.5k

     ...future architectures.Optimize kernels for peak throughput on both silicon and software performance simulators.Collaborate with teams...  ...need to see:Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).3+ years of... 
    Full time

    Nvidia

    Redmond, WA
    2 days ago
  • $143.7k - $194.4k

     ...network. Our mission is to deliver fast, reliable internet connectivity to customers...  ...with every device we design, from custom silicon to secure software, to enable innovative...  ...reality.As a Device Software Development Engineer on the Amazon Leo for Government (ALG) team... 
    Internship
    Local area
    Relocation package
    Flexible hours

    KForce

    Redmond, WA
    6 days ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole...  ...Error Correction Software Engineer. This position offers an opportunity...  ...gathering, day to day task automation).Additional or Preferred...  ...equivalent experience. Experience with HPC, scientific programming, and/... 
    Ongoing contract
    Permanent employment
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    3 days ago
  • $152k - $241.5k

    NVIDIA is looking for outstanding software engineers to help us expand our enterprise GPU management and monitoring tools. In this role, you...  ...ecosystem. We are focused on supporting NVIDIA products across HPC, cloud, and enterprise on both bare metal and virtualized... 
    Full time

    Nvidia

    Redmond, WA
    6 days ago
  • $125k - $145k

     ...ultimate goal of enabling human life on Mars.SOFTWARE ENGINEER, HARDWARE TEST & AUTOMATION (OPTICAL PAYLOADS)As a Software Engineer on the Starlink...  ...that test flight components for maximum performance and reliability in extreme environments, with mission success being... 
    Permanent employment
    Temporary work
    Work at office
    Monday to Friday
    Weekend work

    SpaceX

    Redmond, WA
    2 days ago
  • Build the sound of next-gen consumer devicesWe are hiring a hands-on Acoustic Test Engineer to join our audio design team working on cutting-edge consumer electronics and IoT products. If you love measuring, breaking, and improving real-world audio systems - this is your... 

    CyberCoders

    Redmond, WA
    2 days ago
  • $194k - $267k

     ...be challenged and have a passion for solving large-scale automation, testing, and tuning problems, we would love to hear from...  ...self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    Bellevue, WA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer — HPC & Automation (Silicon Engineering). Be the first to apply!