Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff HPC Infrastructure Engineer

$173k - $237.95k

Guardant Health, Inc.

Company DescriptionGuardant Health is a leading precision oncology company focused on guarding wellness and giving every person more time free from cancer. Founded in 2012, Guardant is transforming patient care and accelerating new cancer therapies by providing critical insights into what drives disease through its advanced blood and tissue tests, real-world data and AI analytics. Guardant tests help improve outcomes across all stages of care, including screening to find cancer early, monitoring for recurrence in early-stage cancer, and treatment selection for patients with advanced cancer. For more information, visit guardanthealth.com and follow the company on LinkedIn, X (Twitter) and Facebook. Guardant's HPC team builds and operates the computational technology backbone of the company. This includes scalable data storage that holds petabytes of genomics data, high-performance compute clusters running a custom bioinformatics pipeline in production and R&D environments, and the software infrastructure that hosts an ecosystem of services for internal data processing and external data integration.The HPC engineering team is looking for a Staff-level engineer with broad, all-round HPC competency and specialized depth in one or more of the following: Red Hat-family OS management, networking, storage, Kubernetes, or Slurm. Experience applying these skills both on-premise and in the cloud is a plus, as we continue to evolve our HPC footprint. This role carries technical depth and cross-team representation for HPC compute, networking, and storage-integration initiatives, partnering closely with our engineering team and our managed service provider (MSP) as we scale operations. To support Guardant Health's fast growth over the next few years, we need a strong technical engineer who can help maintain and grow the HPC infrastructure through this expansion while working closely with corporate IT, SQA, and DevOps/SRE teams. Strong teamwork and the ability to manage multiple in-flight, cross-functional projects at once are essential for success in this role.About the RoleYou enjoy an agile, very fast paced and highly technical environment. You are a self-driven, accomplished technologist who strives to continually improve your skills as the HPC landscape evolves and the computational infrastructure scales. You are dedicated to engineering excellence yet pragmatic and flexible. You have the ability to maintain the day-to-day support SLA while running various key projects that move the business forward. You are comfortable being the deepest technical voice in the room on your declared specialism, even among more senior colleagues, while operating as a peer to the other engineers on the team.Essential Duties and ResponsibilitiesBroad HPC Skills (all-round)· Manage multiple HPC clusters and cluster file systems· Integrate cloud bursting as part of the HPC abstraction work· Research, develop, and implement the next generation HPC solutions· Troubleshoot the production system stack down to source code level, e.g shell scripts, Python, and others· Maintain, monitor, and support the infrastructure environment and/or facilities· Use and maintain enhanced production monitoring and addition capability· Support improvements for increased system reliability and performance· Support multiple systems or applications of medium to high complexity complexity defined by size, technology used, and system feeds and interfaces) with multiple concurrent users, ensuring control, integrity, and accessibility· Support systems at remote locations, including internationally· Mentor junior engineers on HPC best practices· Work with offsite consultants to maintain the infrastructure· Work with vendors to troubleshoot, upgrade, and repair systems as needed· Represent HPC infrastructure networking and storage-integration topics in cross-functional planning with networking, SQA, DevOps/SRE, and the MSP· Set up and ownership supporting XDMoD instances for HPC metric and monitoring· Participate in a 24/7 on-call rotationNetworking· Act as the technical peer for HPC networking and interconnect initiatives with the dedicated networking engineer· Collaborate on design, performance tuning, and troubleshooting of HPC Ethernet· Work with enterprise networking on integration of HPC systems with the bandwidth-on-demand system that connects our sites and cloud infrastructure· Work with the networking infrastructure team to manage and optimize connectivity to and from HPC systems and global locationsStorage· Act as the technical peer for the architecture and integration strategy for. HPC storage in partnership with the dedicated storage engineer and MSP· Serve as a technical point of contact for the MSP storage relationship and help define and evolve SLAs, validate delivery, and escalate technical issues· Support the transition of day-to-day storage operations to the MSP without loss of performance or reliabilityRequired QualificationsBachelor’s degree in Computer Science or a related field with 8–12 years of relevant experience; Master’s degree with 6–8 years of relevant experience; or PhD with 3–5 years of relevant experienceStrong experience in systems and/or infrastructure engineering, including Linux/Unix administration and TCP/IP networking.Hands-on experience with automation tools, such as Ansible or equivalent technologies.Experience with high-performance networking technologies, such as InfiniBand, RoCE, RDMA, or equivalent, including troubleshooting in production environments.Experience supporting large-scale data storage and high-performance computing (HPC)/compute environments.Experience working with both on-premise and cloud-based infrastructure, such as AWS, Google Cloud Platform (GCP), Azure, or similar environments.Experience developing and supporting software release, operations, and infrastructure automation processes and toolsets.Strong experience creating and maintaining system administration and technical documentation.Preferred QualificationsCisco Certified Network Professional (CCNP) certificationExperience with Arista and compatible networking, up to and including 400 Gb/s linksExperience administering IBM's General Parallel File System (GPFS)Experience administering the Slurm schedulerExperience using WarewulfLinux support and OS management. Red Hat family is a must, but Debian or Suse is nice to have.Experience with cloud bursting technologiesExperience with wide area file systemsExperience with Docker and Apptainer container technologiesExperience with KubernetesOperating infrastructure compliant with HIPAA and SOX standardsAI & Digital FluencyDemonstrate curiosity, sound judgment, and the ability to critically evaluate and responsibly leverage AI-enabled tools in accordance with company policies, ethical standards, and regulatory requirements to improve the efficiency, effectiveness, and quality of work.Hybrid Work Model:This section is applicable to onsite employees who are eligible for hybrid work location as specified by management and related policies. Guardant has defined days for in-person/onsite collaboration and work-from-home days for individual-focused time. All U.S. employees who live within 50 miles of a Guardant facility will be required to be onsite on Mondays, Tuesdays, and Thursdays. We have found aligning our scheduled in-office days allows our teams to do the best work and creates the focused thinking time our innovative work requires. At Guardant, our work model has created flexibility for better work-life balance while keeping teams connected to advance our science for our patients. The annualized base salary ranges for the primary location and any additional locations are listed below. This range does not include benefits or, if applicable, bonus, commission, or equity. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, education, job-related skills, job duties, and business need. Primary Location: Palo Alto, CA Primary Location Base Pay Range: $173,000 - $237,950 Other US Location(s) Base Pay Range: $147,100 - $202,300 If the role is performed in Colorado, the pay range for this job is: $155,700 - $214,150Employee may be required to lift routine office supplies and use office equipment. Majority of the work is performed in a desk/office environment; however, there may be exposure to high noise levels, fumes, and biohazard material in the laboratory environment. Ability to sit for extended periods of time.Guardant Health is committed to providing reasonable accommodations in our hiring processes for candidates with disabilities, long-term conditions, mental health conditions, or sincerely held religious beliefs. If you need support, please reach out to View email address on us.fitly.work background screening including criminal history is required for this role. GH will consider qualified applicants with criminal arrest or conviction histories in a manner consistent with applicable law including but not limited to the LA County Fair Chance Policies and the Fair Chance Act (Gov. Code Section 12952).Guardant Health is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or protected veteran status and will not be discriminated against on the basis of disability.All your information will be kept confidential according to EEO guidelines.To learn more about the information collected when you apply for a position at Guardant Health, Inc. and how it is used, please review our Privacy Notice for Job Applicants.Please visit our career page at: SummaryLocation: Palo Alto, CAType: Full time

Vacancy posted 7 days ago
Similar jobs that could be interesting for youBased on the Staff HPC Infrastructure Engineer in Palo Alto, CA vacancy
  •  ...software to expose next-generation hardware IO capabilities for AI/HPC application teams Govern a generic IO API with multiple...  ...Qualifications Master's/PhD in Computer Science or Electrical Engineering + 1 year industry experience, OR 3+ years industry experience.... 
    Suggested
    Full time

    Cerebras Systems

    Sunnyvale, CA
    9 hours ago
  •  ...the firmware/PHY connectivity layer. You will drive development and verification across software and firmware, enabling cutting-edge HPC networking technologies for large-scale data centers. The role requires 8+ years of technical experience with 3+ years of leadership... 
    Suggested

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $150k - $250k

     ...small, highly motivated, and focused on engineering excellence. This organization is for...  ...networks that underpin training and inference infrastructure, including high-performance /...  ...supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA-capable designs) as... 
    Suggested
    Temporary work
    Night shift

    SpaceXAI

    Palo Alto, CA
    5 days ago
  • $165.5k - $289.6k

     ...It all started when engineer Fred Luddy wrote code that automated a tedious task for...  ...is seeking a highly experienced Senior Staff Cloud FinOps Analyst to lead enterprise...  ...Microsoft Azure, Google Cloud Platform, private infrastructure, and emerging AI services. This is... 
    Suggested
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    4 days ago
  • $161k - $264.5k

     ...City, a test course for mobility; and Cloud & AI, the digital infrastructure powering our collaborative foundation. Business-critical...  ...automotive software developers. That's why the Enterprise Technology Engineering Team (EnTec) builds solutions that enhance productivity, so... 
    Suggested
    Temporary work
    For contractors
    Work at office
    Flexible hours

    Woven by Toyota

    Palo Alto, CA
    26 days ago
  • $160k - $190k

     ...Robotics is currently seeking an IT Systems Engineer to join our growing team working out of...  ...support services and collaborating on infrastructure-as-code projects with the greater IT...  ...onsite and remote guidance for Engineers and staff Provide end user guidance to Windows,... 
    Temporary work
    Work at office
    Remote work
    Visa sponsorship
    Flexible hours

    Kodiak

    Mountain View, CA
    7 days ago
  • $189.3k - $290.7k

     ...Posting summary Production Mapping is looking for a Staff Cloud and AI Solutions Engineer to build and scale future-ready mapping foundations...  ...work across data engineering, backend services, cloud infrastructure, distributed systems, developer tooling, and AI-enabled... 
    Full time
    Local area
    Work from home
    Flexible hours

    General Motors

    Sunnyvale, CA
    2 days ago
  •  ...Woven by Toyota is seeking a Head of Infrastructure Engineering in our Palo Alto office to lead a global team responsible for corporate infrastructure systems including network, cloud, productivity and AI tools. You will cultivate an internationally distributed team... 
    Work at office

    Jobleads-US

    Palo Alto, CA
    4 days ago
  • $170k - $210k

     ...and 10-50x more efficient. As a Staff Backend Developer on the ALSO backend...  ...and architectural decisions for the core infrastructure behind the customer experience on ALSO'...  ...monitoring, and deployment safety Mentor engineers and contribute to technical decision-... 
    Local area
    Remote work
    Flexible hours

    ALSO

    Palo Alto, CA
    8 days ago
  • $140k - $210k

     ...Stanford researchers and veteran systems engineers who share a vision for redefining the...  ...grow increasingly complex, traditional infrastructure struggles to meet the demands of performance...  ...to Have Experience supporting AI/ML, HPC, storage, or GPU cluster infrastructure... 

    Clockwork.io

    Palo Alto, CA
    5 days ago
  • $168k - $270.25k

     ...Enterprise Experience (NVEX) Solutions Engineering team is looking for a senior Computer or...  ...-X that link GPUs and AI compute infrastructure. Candidates must have a software development...  ...:Background with AI infrastructure and HPC networkingExperience programming switch... 
    Full time
    Weekend work

    Nvidia

    Santa Clara, CA
    28 days ago
  • Company DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She...  ...product-specific search solutions to a composable, agent-native infrastructure foundation that agents and applications build on to locate,... 
    Full time
    Work at office
    Immediate start
    Remote work
    Flexible hours
    Shift work

    Moveworks

    Mountain View, CA
    7 days ago
  • $160k - $275k

     ...system is visible without anyone having to ask.We expect every engineer to handle everyday collaboration: open a clean PR, review one...  ...a migration that cannot.Bonus Points If You Have:Developer-infrastructure work under hard deadlines in another industry (silicon, aerospace... 
    Daily paid
    Full time
    Work experience placement
    Local area
    Remote work
    Monday to Friday
    Flexible hours

    MatX

    Mountain View, CA
    5 days ago
  • $109k - $160k

     ...enterprises, CoreWeave combines superior infrastructure performance with deep technical...  ...at  . About the role A Software Engineer contributes to the design, implementation...  ...Experience testing hardware at scale. HPC Experience. Experience with AI/ML infrastructure... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Remote work
    Flexible hours

    Core Weave

    Sunnyvale, CA
    9 hours ago
  • $160k - $275k

    What MatX Is BuildingWe're a small engineering team designing a custom chip. The work is compute...  ...engineers depend on every day. The infrastructure that supports all of this — CI/CD,...  ...equivalentOperating batch compute or job schedulers — HPC, Slurm, Nomad, Kubernetes batch, or... 
    Daily paid
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    Day shift
    3 days per week

    MatX

    Mountain View, CA
    4 days ago
  • $150k - $170k

     ...cryogenic cooling and standard fiber-optic infrastructure. In 2024, PsiQuantum announced...  ...OverviewPsiQuantum's Applications Software Engineering Team (ASET) builds tools for quantum algorithm...  ...PsiQuantum, including to leverage HPC/GPU resourcesMake deployments faster to... 
    Full time
    Shift work

    PsiQuantum

    Palo Alto, CA
    a month ago
  •  ...accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with... 
    Full time
    Temporary work
    Remote work

    X.Ai

    Palo Alto, CA
    9 hours ago
  • $193.93k - $291.15k

     ...Role The ability to monitor and assist our vehicles remotely plays a key role in our business strategy. As a Senior Software Engineer, Networking you will work on our in-house Teleoperations platform. You will work with a diverse team of engineers to build the core... 
    Full time
    Remote work

    Nuro

    Mountain View, CA
    9 hours ago
  • $148.2k - $222.2k

     ...Staff Engineer, IT Infrastructure, Platforms Position SummaryThe Staff Engineer, IT Infrastructure, Platforms provides technical leadership for the engineering...  ..., enterprise storage, High Performance Computing (HPC), and modern platform services.The ideal candidate... 
    Full time
    Work from home
    Monday to Friday

    Pacific Biosciences

    Menlo Park, CA
    a month ago
  • $127k - $204k

     ...and accessible for all.We're searching for a Senior Network Engineer (P6) to join our Network Engineering team. This is a senior individual...  ...you.In this role you willDesign, deploy, and operate network infrastructure (switches, routers, wireless access points, and firewalls)... 
    Work at office
    Local area
    3 days per week

    Aurora Innovation

    Mountain View, CA
    5 days ago
  • $148.2k - $222.2k

    Staff Engineer, IT InfrastructureAbout the RolePacBio is looking for a Staff Engineer, IT Infrastructure to help design, operate, secure, and evolve the infrastructure supporting our research...  ...You will work across Enterprise Linux, HPC and scientific computing, cloud... 
    Full time
    Immediate start
    Work from home
    Monday to Friday

    Pacific Biosciences

    Menlo Park, CA
    4 days ago
  • $168k - $264.5k

     ...how the world builds and operates computing infrastructure for accelerated computing and AI. Our Network Deployment Engineering team designs and delivers the global network...  ...architectures.Design or deployment background with AI/HPC networking, including NVIDIA Spectrum-X,... 
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $118k - $231k

    We are looking for a Senior Network Engineer to take ownership of key areas of MongoDB's global network infrastructure. You'll design, deploy, and troubleshoot network and VPN solutions, including Palo Alto Prisma Access at global scale, partner with InfoSec to harden... 
    Casual work
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Palo Alto, CA
    14 days ago
  • $180k

     ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  ...understand the universe.  We are looking for exceptional ML Infrastructure Engineers with deep expertise in high-speed interconnect technologies... 
    Temporary work

    SpaceXAI

    Palo Alto, CA
    18 days ago
  •  ...Preference for candidates in/near Palo Alto. Visa Sponsorship not available Role: Infrastructure Engineer/architect Location: Palo Alto, CA Full-Time Position Hybrid What You'll Do Keep highly available systems highly available. We process mortgage... 
    Full time
    Work experience placement

    HARAMAIN SYSTEMS INC.

    Palo Alto, CA
    a month ago
  • Job Title 12+ years in platform engineering, SRE, or DevOps. Experience with HPC clusters (Slurm, PBS, Grid Engine). Cloud infrastructure expertise (GCP/AWS preferred). Proficiency with Terraform, Ansible, Prometheus, Grafana, ELK. Strong Linux administration and scripting... 

    Saxon Global

    Mountain View, CA
    6 days ago
  • Senior Network Engineer Location: Pal Alto, CA 3days a week(Hybrid) Contract Experience: 10+ Job Description : Rubrik is looking...  ...( AWS, Azure, GCP, and OCI ). You will join the Network Infrastructure Services team to build, deploy, and maintain standardized network... 
    Contract work

    Argyll Infotech, Inc.

    Palo Alto, CA
    6 days ago
  • $180k - $240k

     ...maintenance, turn-ups) Team: Technology Services - Network Engineering Employment Type: Full-Time Overview We are seeking an experienced...  ..., maintaining, and optimizing modern enterprise network infrastructure. In this position, you will design, deploy, and support... 
    Full time

    Tenth Revolution Group

    Menlo Park, CA
    3 days ago
  • $100k - $258k

     ...knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who...  .... ABOUT THE ROLE: We are seeking a talented and motivated Infrastructure Security Engineer to join our security team. In this role, you... 
    Temporary work
    Relocation

    SpaceXAI

    Palo Alto, CA
    6 days ago
  • Job Title: Network Engineer Client: (Financial Software Company) Location: 2700 Coast Avenue, Mountain View, CA 94040 (Onsite - Local...  ...security, scalability, and high availability of networking infrastructure. Required Skills & Qualifications Mandatory : Hands-on... 
    Contract work
    Local area
    Immediate start

    Kaav Inc.

    Mountain View, CA
    6 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff HPC Infrastructure Engineer. Be the first to apply!