Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

System Infrastructure / Platform Engineer, HPC Technology Department

$156.86k - $191.72k

Berkeley Lab

The National Energy Research Scientific Computing Center (NERSC) is seeking a System Infrastructure / Platform Engineer to help build and manage HPC systems and Linux-based infrastructure. NERSC operates some of the world’s largest supercomputers, supporting thousands of researchers tackling major scientific challenges. In this role, you will manage high-performance computing environments, including HPC systems, containers, virtual machines, and core infrastructure services. You'll work with cutting-edge technologies such as CPU/GPU clusters, parallel storage, high-speed networking, Slurm, and Kubernetes, balancing innovation with reliability, performance, and security at scale. Collaborating with engineers, researchers, vendors, and open-source communities, you will help develop scalable solutions that advance scientific discovery and the future of HPC. If you have Linux experience, an interest in science, and enjoy fast-paced collaborative environments, NERSC would love to hear from you. We’re here for the same mission, to bring science solutions to the world. Join our team and YOU will play a supporting role in our goal to address global challenges! Have a high level of impact and work for an organization associated with 17 Nobel Prizes! Why join Berkeley Lab? Exceptional health and retirement benefits , including pension or 401K-style plans Opportunities to grow in your career - check out our Tuition Assistance Program A culture where you’ll belong - we are invested in our teams! In addition to accruing vacation and sick time, we also have a Winter Holiday Shutdown every year. Parental bonding leave (for both mothers and fathers) Pet insurance What You Will Do if hired at a Level 3: Build and manage Linux systems and storage infrastructure Troubleshoot complex technical issues with team members Install, upgrade, and secure systems and services Develop and maintain scripts and automation tools Participate in a 24/7 on-call rotation Lead small projects, upgrades, and service rollouts Collaborate with vendors to improve technologies and user experience Support reliable operations of NERSC’s Perlmutter supercomputer and Spin Kubernetes platform Develop and integrate services across NERSC and DOE facilities, including the upcoming Doudna supercomputer Present technical work to the HPC community at conferences and industry events Responsibilities In Additional Responsibilities if hired at a Level 4: Solve complex technical problems with independent judgment Develop team strategies and project plans Provide technical leadership and mentorship Lead system improvements for performance, reliability, and security Evaluate emerging HPC technologies and capabilities Represent NERSC in HPC and DOE technical communities and advocacy groups What is Required to be hired at a Level 3: Typically, 8+ years of related experience with a Bachelor’s degree; alternatively, 6+ years with a Master’s degree; or equivalent career experience 4+ years of experience managing large-scale Linux-based system deployments in a high-performance computing, cloud computing, or hyper-scale environment Mastery of Linux concepts and operations (processes, networking, system logs, performance) Proficiency with bash and Python scripting Experience with some or all of our key technologies: containers (such as Docker or Kubernetes) virtualization (such as Proxmox or VMware) cloud-based deployment (such as AWS, Azure or GCP) identity and access management database administration, tuning, and troubleshooting storage systems technologies (such as iSCSI and NAS appliances) parallel filesystems (such as Lustre, GPFS, or VAST) high-speed networking/interconnect (such as InfiniBand, Slingshot, or RoCE) advanced performance analysis and debugging tools (such as strace, lsof, ebpf, or gdb) DevOps tools (such as Gitlab or Jira) and processes (such as issues, merge requests, and API/automation) Familiarity with automated provisioning systems (such as Chef, Foreman, or Terraform) Familiarity with configuration management systems (such as Ansible or Puppet) Working knowledge of Linux system engineering and security practices Ability to resolve complex issues in creative and effective ways and derive technical solutions in a collaborative environment to meet end user requirements or needs Demonstrated ability to work independently as well as collaboratively in large projects, and contribute to an active and respectful intellectual environment Creative, positive, and collaborative work style Excellent oral and written communication skills Requirements Additional Requirements to be hired at a Level 4 Typically, 12+ years of related experience with a Bachelor’s degree; alternatively, 8+ years with a Master’s degree; or equivalent career experience Proven ability to lead troubleshooting and resolution of high-impact incidents in complex, large-scale environments Demonstrated leadership in cross-team collaboration and mentoring Experience in software engineering, Linux systems programming, or complex scripting Experience managing one or more of the following: data center networking (TCP/IP, Ethernet, BGP, ECMP) batch workload managers (such as Slurm), including installation, configuration, routine operations, job lifecycle concepts, and troubleshooting common failure modes Cray/HPE HPC ecosystems (e.g., CSM/COS, Slingshot interconnect, and related components) Ability to lead and coordinate projects with traditional or Agile methodologies (such as Scrum or Kanban) Ability to analyze and resolve significant and unique issues requiring evaluation of multiple intangible factors Ability to exercise independent judgment in methods, techniques and evaluation criteria for obtaining results Additional Information Applications will be accepted until the job posting is removed. Appointment type: This is a full-time, career appointment, exempt (monthly paid) from overtime pay. Salary range: Level 3: The expected salary for this position is $156,864 - $191,724, which fits into the full salary of $139,440 - $235,308 depending upon the candidate’s skills, knowledge, and abilities. This includes education, certifications, and years of experience. Level 4: The expected salary for this position is $178,644 - $218,364, which fits into the full salary of $158,808 - $267,996 depending upon the candidate’s skills, knowledge, and abilities. This includes education, certifications, and years of experience. Background check: This position is subject to a background check. Any convictions will be evaluated to determine if they directly relate to the responsibilities and requirements of the position. Having a conviction history will not automatically disqualify an applicant from being considered for employment. Work modality: This position requires substantial on-site presence, but is eligible for a flexible work mode, and hybrid schedules may be considered. Hybrid work is a combination of performing work on-site at Lawrence Berkeley National Lab, 1 Cyclotron Road, Berkeley, CA and some telework. Individuals working a hybrid schedule must reside within 150 miles of Berkeley Lab. Work schedules are dependent on business needs. In rare cases, full-time telework or remote work modes may be considered. Multi-level Posting: This position will be hired at a level commensurate with the business needs and the skills, knowledge, and abilities of the successful candidate. Export Control Access: This position will involve access to hardware, commodities, and technical information subject to export control regulations including, but not limited to, the Export Administration Regulations (EAR) and/or International Traffic in Arms Regulations (ITAR). Accordingly, any hiring decision may depend in part on Berkeley Lab’s ability to obtain or rely on federal government authorizations as required, if you are not a U.S. citizen, lawful permanent resident of the U.S. (“green card holder”), asylee, refugee, or other qualifying protected individual as defined by 8 U.S.C. 1324b(a)(3). Equal Employment Opportunity Employer: The foundation of Berkeley Lab is our Stewardship Values: Team Science, Service, Trust, Innovation, and Respect; and we strive to build community with these shared values and commitments. Berkeley Lab is an Equal Opportunity Employer. We heartily welcome applications from all who could contribute to the Lab's mission of leading scientific discovery, excellence, and professionalism. In support of our rich global community, all qualified applicants will be considered for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, protected veteran status, or other protected categories under State and Federal law. Misconduct Disclosure Requirement: As a condition of employment, the finalist will be required to disclose if they are subject to any final administrative or judicial decisions within the last seven years determining that they committed any misconduct, are currently being investigated for misconduct, left a position during an investigation for alleged misconduct, or have filed an appeal with a previous employer. #J-18808-Ljbffr Berkeley Lab

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the System Infrastructure / Platform Engineer, HPC Technology Department in Berkeley, CA vacancy
  • $90 per hour

     ...and other software engineering projects to help...  ...join a group of systems and software engineers...  ...operated by the Department of Energy Office...  ...enhance their technologies to meet user needs...  ...container cloud platform based on Kubernetes...  ...and the broader HPC community at... 
    Suggested
    Contract work
    Work at office

    Bay Systems Consulting Inc.

    Berkeley, CA
    a month ago
  •  ...programs. We build the infrastructure behind some of...  ...Software / API Engineer to join a dynamic...  ...software systems that integrate scientific...  ...edge science and technology, contributing to...  ...container cloud platforms based on Kubernetes...  ...and the broader HPC community at conferences... 
    Suggested
    Temporary work

    LTD Global

    Berkeley, CA
    a month ago
  • Berkeley Lab is hiring a System Infrastructure / Platform Engineer to build and manage high-performance computing systems. This role involves working with HPC systems, Linux infrastructure, and collaborating with engineers and researchers to develop scalable solutions.... 
    Suggested

    Berkeley Lab

    Berkeley, CA
    3 days ago
  • $180k - $225k

    AI Infrastructure Engineer - Agent Sandbox PlatformAs a Software...  ...our agent sandboxing platform — the secure, high-...  ...researchers using this system as they do about the...  ...and virtualization technologies (e.g., Docker,...  ...see the United States Department of Labor's Know Your... 
    Suggested
    Full time
    Immediate start
    Remote work

    Scale AI

    San Francisco, CA
    9 hours ago
  • $148.7k - $201.2k

     ...team of scientists, engineers, and technicians, on...  ...are looking to hire an HPC Platform Engineer to develop,...  ...performance computing (HPC) infrastructure on AWS that CQC...  ...users, and securing systems.* Support computer-aided...  ...quantum computing technologies.Basic qualifications-... 
    Suggested
    Local area
    Flexible hours

    Amazon

    San Francisco, CA
    4 days ago
  •  ...impact.As a Lead Software Engineer- Linux at JPMorgan Chase within...  ...Corporate Sector Compute Infrastructure Platform (CIP) organization, you...  ...initiatives consisting of multiple technologies and applications. Job...  ...certify new Red Hat operating systems, create multiple image... 

    JP Morgan Chase

    San Francisco, CA
    4 days ago
  • $260k - $350k

     ...harnessing emerging technologies to redefine transportation...  ...software to the data platform that will power the...  ...impact for the engine of the American economy...  ...our foundational data systems, ensuring scalability...  ...pipelines, optimize cloud infrastructure, and champion best... 
    Full time
    Work at office
    Immediate start

    Baton Trucking

    San Francisco, CA
    1 day ago
  • $15k

    Voleon is a technology company that applies state-of-...  ...Cluster Site Reliability Engineer (SRE), you will help...  ...a world-class HPC platform for researchers to focus...  ...both on-prem and cloud infrastructure, and work to provide...  ...while also engineering systemic improvements and... 
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    4 days ago
  • DutchTech is seeking a talented Software Engineer to develop next-generation software for illumination systems. This foundational role involves architecting connections between hardware, firmware, and cloud infrastructure. The ideal candidate has a degree in Software Engineering... 

    DutchTech

    Emeryville, CA
    5 days ago
  •  ...California is looking for an experienced platform engineer to design, build, and operate significant components of their trading infrastructure. Candidates should have over 5 years in...  ...role offers the chance to influence systems that support a robust quantitative trading... 

    The Voleon Group

    Berkeley, CA
    4 days ago
  • $90k - $180k

     ...portfolio of life-changing technologies spans the spectrum of...  ...health platform that combines continuous...  ...healthier lives. Our systems ingest millions of sensor...  ...OPPORTUNITYThis Senior Platform Engineer position works onsite...  ...improve the cloud infrastructure, CI/CD pipelines,... 
    Worldwide

    Abbott

    Alameda, CA
    1 day ago
  • $15k

    Voleon is a technology company that applies state-of-the-art...  ...TeamAs a Senior Software Engineer on the Software Platform team, you will design and evolve the distributed systems that power research and trading...  ...that abstract away infrastructure complexity.Operating at the... 
    Work at office
    Local area

    The Voleon Group

    Berkeley, CA
    4 days ago
  • $180k - $225k

     ...reality, leading platform companies are scrambling...  ...AI Data Engine, SGP, Donovan, and...  ...Scale’s core cloud infrastructure and orchestration...  ...platforms and systems, working closely...  ...development processes and technologies.Ideally you’d...  ...United States Department of Labor's Know Your... 
    Full time
    Live in

    Scale AI

    San Francisco, CA
    9 hours ago
  • $15k

    Voleon is a technology company that applies state-of-the-art AI and machine learning techniques...  ...daily catered lunches, and more.Strategy Platform owns the infrastructure between quantitative research and live trading. Our systems orchestrate the transformation of... 
    Work at office
    Local area

    The Voleon Group

    Berkeley, CA
    4 days ago
  • $170k - $215k

    A pioneering micromanufacturing company is seeking a Staff Systems Engineer to develop innovative manufacturing hardware. The role involves collaboration with various engineering teams, prototype integration, and hands-on work with robotics. Ideal candidates will have... 

    Atomic Machines

    Emeryville, CA
    5 days ago
  • $148.7k - $297.3k

     ...portfolio of life-changing technologies spans the spectrum of...  ...health platform that combines continuous...  ...healthier lives. Our systems ingest millions of sensor...  ...Senior Manager, Platform Engineering. You are a platform...  ...of Lingo’s cloud infrastructure on Azure, ensuring it... 
    Worldwide

    Abbott

    Alameda, CA
    1 day ago
  • $133.65k - $222k

     ...transportation with technology that's powering commercial...  ...for complex embedded systems based on requirements...  ...and maintaining test infrastructure including Hardware-in...  ...of embedded platform components (including...  ...- Bachelors in an engineering discipline (MS/PhD preferred... 
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    23 days ago
  • $170k - $277k

     ...with cutting-edge technology and bold thinking...  ...cybersecurity platform that provides comprehensive...  ..., analytics engine, and user...  ...contribute to shared infrastructures, tackling complex...  ...engineers across the department, fostering a...  ...scale distributed systems.Extensive experience... 
    Full time
    Work at office

    Palo Alto Networks

    San Francisco, CA
    2 days ago
  •  ...understand your body — and the platform engineering team that underpins that...  ...expanding. As a Senior IT Systems/Platform Engineer, you'll...  ...engineering, network infrastructure, and day-to-day systems engineering...  ...a high-growth, fast-paced technology company — infrastructure,... 
    Work at office
    Local area
    Flexible hours
    3 days per week

    Oura

    San Francisco, CA
    3 days ago
  • $194k - $267k

     ...building the trusted, neutral infrastructure that enables organizations...  ...looking for a Staff level DevOps Engineer to join a team of highly...  ..., work with cutting-edge technologies, and directly share in Okta'...  ...businessEvaluate and scale existing systems to meet specialized... 
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    1 day ago
  • Staff Web Frontend Engineer (Frontend Platform)Location: San Francisco, CA (Hybrid...  ...at the intersection of technology, accessibility, and trust—...  ...horizontally—someone who builds systems other engineers rely on....  ...such as: Design system infrastructure (tokens, theming)Build tooling... 
    Local area

    Talener

    San Francisco, CA
    1 day ago
  •  ...the Team The Applied Engineering team works across...  ...design to bring OpenAI’s technology to consumers and...  ...for running the core infrastructure that supports products...  ...ChatGPT and the API. The systems we support include...  ...development and production platforms that power our... 
    Full time
    Relocation package

    OpenAI

    San Francisco, CA
    9 hours ago
  • $250k - $350k

     ...The Role As a Senior Software Engineer, Infrastructure and Platform, you'll design and build the core infrastructure...  ..., evaluation, and agentic systems — the shared platforms that let engineers...  ....js/Next.js) or similar backend technologies ~ Experience designing and... 
    Full time
    Visa sponsorship

    David Joseph & Company

    San Francisco, CA
    20 days ago
  • $180k - $400k

     ...interactive entertainment, technology that enables millions...  ...technical founders, engineers that made 100+ games...  ...AI and large-scale platform engineering. Your...  ...and owning the core infrastructure, platforms, and toolchains...  .../mobile, and cloud systems to speed up iteration... 
    Full time
    Visa sponsorship
    Relocation package

    ROAM

    San Francisco, CA
    13 hours ago
  •  ...AI Product Engineer Anything is the AI product...  ...else. You'll build the systems that supports millions...  ...that leverage platform telemetry, execution...  ...operate multi-tenant cloud infrastructure, including isolation,...  ...and make decisions on technology choices, and service... 

    Anything Corp.

    San Francisco, CA
    13 hours ago
  • $200k - $230k

     ...services and healthcare technology company based on...  ...DescriptionDirector, AI Platform EngineeringLocations:...  ...of AI Platform Engineering to lead the design, development...  ..., architect critical infrastructure, and drive the...  ...prompt engineeringStrong systems design: distributed... 
    Ongoing contract
    Full time
    Casual work
    Work at office
    Flexible hours

    SS&C Technologies

    San Francisco, CA
    3 days ago
  • $190k - $210k

     ...Senior Engineer Gridware is a San Francisco-based technology company dedicated to protecting and...  ...Active Grid Response platform uses high-precision sensors...  ...the underlying cloud infrastructure and security posture,...  ...experience with CI/CD systems, ideally GitHub Actions... 
    Local area

    Gridware

    San Francisco, CA
    4 days ago
  •  ...businesses deserve financial infrastructure tailored to how they...  ...is, at its core, a technology company and is on a...  ...to build the best engineering team in the world. We...  ...Senior Infrastructure/Platform Engineer focused on...  ...responsible for keeping our systems reliable, secure, and... 
    Full time
    Work at office

    Slash Financial

    San Francisco, CA
    4 days ago
  •  ...robust, scalable trading platform to serve high-traffic,...  ...applications. Our infrastructure leverages state-of-the-art technologies to support real-time trading...  ...of our platform and engineering culture. Job Summary...  ...Grafana) for real-time system monitoring. Incident... 
    Remote work
    Flexible hours

    OnHires

    San Francisco, CA
    more than 2 months ago
  • $160k - $200k

     ...problems, build exciting technology, and work alongside...  ..., smart and humble engineers from all around the world...  .... The result is a system that handles long...  ...Overview: Our Software Platform team builds all the...  ...sensor, data, and demo infrastructure that makes Zendar run... 
    Full time
    Work at office
    Flexible hours

    Zendar

    Berkeley, CA
    17 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to System Infrastructure / Platform Engineer, HPC Technology Department. Be the first to apply!