HPC Systems Engineer
University of California , San Francisco
The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day operations of the Institute's HPC clusters.
The HPC Systems Engineer will:
- Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with unique computing, networking, and storage systems designed to address cutting-edge research problems
- Apply their engineering and design skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity of CI systems.
- Select methods, techniques and evaluation criteria to develop new CI solutions to address complex research problems.
- Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and availability.
- Provides consultation, support, and guidance to researchers on how to address computational problems using standard tools, packages, and approaches.
- Develop enhancements of monitoring to maintain the integrity of CI systems.
- Participate in multiple technical projects simultaneously.
- Applies working knowledge of security control frameworks to maintain the integrity of the CI systems and the research being performed on them.
- Gives presentations to the associated team and other technical units.
Evaluates new technologies, including performing moderate to complex cost/benefit analyses.
This position may lead to cross-functional technical working groups and projects in support of onboarding research customers, or making systems improvements.
Department Overview
Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA-compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers' needs.
The Core HPC team of the Academic Research Service (ARS) focuses on large-scale, high-performance computational and storage services for UCSF researchers so they can address complex computational, AI, and data science problems.
%
of time
Essential Function (Yes/No )
Key Responsibilities
(To be completed by Supervisor)
15
Applies advanced systems / infrastructure concepts to define, design, implement, and operate highly complex, research cyberinfrastructure systems, services and technology solutions. Proposes and implements highly complex system or device enhancements such as software, hardware and network configuration, updates and installations for projects or services of broad scope. Sets standards for monitoring and maintaining the health and integrity of CI systems including upgrading and patching.15
Independently manages systems and services for a large facility, campuswide, medical center or Office of the President and / or institution-wide scope and makes recommendations for purchases or upgrades. Performs complex and advanced analysis to acquire, install, modify and support operating systems, databases, utilities and web-related tools. Selects methods and techniques to obtain solutions. Interacts with senior management. May perform complex network integration tasks and interoperability assessments for interconnected servers or components of clusters for communication. Support and collaborate with researchers and other key IT (e.g. network and security) and Data Center partners in a timely manner15
Specifies, writes and executes highly complex software and scripts to support systems management, log analysis, monitoring, deployment, configuration management, and other system administration duties for multiple, highly integrated systems.30
Provides consultation, training, support, and guidance to researchers enabling them to utilize HPC resources effectively.10
Maintains complex security systems. Interprets and adopts campus, medical center or Office of the President, system and regulation-based security policies to control access to networked resources. Provides recommendations and requirements on network access controls.5
Collaborates and may provide leadership with other Systems Engineers within the CI ecosystem/higher-education community. Regularly contribute best practices documentation, present at conferences, or publish in peer reviewed journals.10
Define and track performance metrics to ensure efficient current and future use of cyber infrastructure resources.100%
(To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.)REQUIRED QUALIFICATIONS
- Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems
- Expert knowledge of HPC systems infrastructure design
- Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc.
- Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements
- Demonstrated testing and test planning skills. Demonstrated ability to create automated testing.
- Knowledge of HPC job scheduler system design and operation such as SLURM or PBS,
- Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters
- Ability to elicit and communicate technical and non-technical information in a clear and concise manner.
- Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines.
- Understanding of system performance monitoring and actions that can be taken to improve or correct performance.
- Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems.
- Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration.
-
Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner
PREFERRED QUALIFICATIONS
- Knowledge of the design, development, and application of technology and systems to meet business needs.
- General knowledge of other areas of IT. Thorough understanding of and experience with systems-related issues and actions that can be taken to improve or correct performance.
- Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions.
REQUIRED QUALIFICATIONS
- Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems
- Expert knowledge of HPC systems infrastructure design
- Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc.
- Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements
- Demonstrated testing and test planning skills. Demonstrated ability to create automated testing.
- Knowledge of HPC job scheduler system design and operation such as SLURM or PBS,
- Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters
- Ability to elicit and communicate technical and non-technical information in a clear and concise manner.
- Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines.
- Understanding of system performance monitoring and actions that can be taken to improve or correct performance.
- Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems.
- Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration.
-
Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner
PREFERRED QUALIFICATIONS
- Knowledge of the design, development, and application of technology and systems to meet business needs.
- General knowledge of other areas of IT. Thorough understanding of and experience with systems-related issues and actions that can be taken to improve or correct performance.
- Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions.
$123.27k - $167.3k
...CommVault Systems Engineer (Data Protection / Backup) Employment Type: Full-Time, Experienced Department: Technology Support CGS is seeking an experienced CommVault Data Protection Engineer with extensive knowledge and experience in designing, developing, configuring...SuggestedFull timeFlexible hours- ...Systems Engineer Atlas Technica's mission is to shoulder IT management, user support, and cybersecurity for our clients, who are hedge funds and other investment firms. Founded in 2016, we have grown year over year through our uncompromising focus on service. We...SuggestedWork at office
$150k - $250k
...already has millions in recurring revenue is building next‑generation AI infrastructure for video intelligence and is looking for a Systems Engineer to join their team. What You'll Be Doing Design and engineer systems that handle compute, scheduling, and orchestration of...SuggestedFull timeImmediate startSleeping nights$174.5k - $240k
...as we power the shop local movement. If you believe in community, come join ours. About this role GTM Engineering builds and operates the intelligent systems, integrations, and automations that power Faire's GTM revenue org. We're the AI and technical backbone of...SuggestedWork at officeLocal areaRemote workFlexible hours3 days per week- ...high intensity to push the frontier forward. The Production Engineering Team Examples of key exciting problems the team is working... ...interface for the whole company, not a hundred scripts. Make the system's view of itself always match reality: integrate fleet state...SuggestedLocal area
$196k - $248k
...Waymo Systems Engineering Role Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced...Full timeRemote work- ...Infrastructure for the World's Largest Dataset You'll be our eleventh engineering hire. You'll have full ownership over major features, play a... ...better products. Your Qualifications ~4+ years of systems software development experience ~ Deep customer empathy,...Immediate start
- ...Systems Engineer - Hybrid Location: San Francisco Bay Area, CA (Hybrid) Reports to: Client Technology Manager Type: Full‑Time Hourly, Non‑Exempt Atlas Technica's mission is to shoulder IT management, user support, and cybersecurity for our clients, which are hedge funds...Hourly payFull timeWork at office
- ...System Engineer Location: Foster City, CA OR San Francisco, CA (Hybrid) Role Summary: The Systems Engineer will be responsible for developing system requirements, test cases, test automation & power train systems issue RCA support for an autonomous vehicle....Shift work
- ...Description Specify and develop procedures and mechanisms for system status monitoring and performance reporting on an LTE system.... .... Qualifications ~ Bachelor's degree in Electrical Engineering, Computer Science, or related field is required. ~5+ years...Permanent employmentFull timeRemote work
- ...Distributed Systems Engineer As a distributed systems engineer, you'll work across the stack to solve problems as they come up and help build Archil volumes. You'll have significant influence over the technical and product direction. We'll expect you to be able...Flexible hours
$200k - $300k
...Team: We've assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis, Microsoft, Splunk, Ngrok,... ...~ An insatiable desire to ship. ~7+ years of software engineering experience comprising of: ~5+ years of backend development...Work at officeShift work$190k - $280k
...A tech company in San Francisco is seeking a Senior Software Engineer to design and operate infrastructure within a distributed architecture. Responsibilities include enhancing control plane systems and improving reliability. Ideal candidates have over 5 years of experience...- ...Technical Staff (MTS) in San Francisco, California, to develop production-grade systems that support continuous optimization for AI agents. This unique position blends machine learning engineering with backend development, emphasizing customer collaboration to align...
- ...sensors for precision mapping, robotics, automotive, security systems, smart cities and various industrial solutions. We’ve transformed... ...lower price. The Role We’re seeking a proactive and skilled engineer who thrives in a dynamic, fast‑paced environment to join our San...Work experience placementLocal area
- ...and County of San Francisco is seeking a TCUP Technology Project Engineer (Electrical Assistant Engineer) to support CBTC integration for... ...to ensure alignment across disciplines and successful system delivery. The role involves design reviews, field testing, and...
$190k - $230k
...part of a high-performing team that believes in each other, come build with us at Crusoe. About This Role We’re seeking a Senior Systems Engineer to play a key role in executing Crusoe’s 2026 Enterprise AI Strategy. In this role, you will design and build agentic AI...Temporary work$165k - $200k
...Senior Systems Engineer Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power...Temporary work- ...to join a team redefining hardware technology? We are searching for dedicated engineers to join our exciting team of innovators. In this role, you will be a key member of the RM modeling system engineering team working on modem systems features and requirements....
$180k - $240k
...Senior Systems Engineer Houston; New York; San Francisco; Seattle; US About the Role We are looking for a Senior Systems Developer to lead the design, development, and operation of large-scale distributed systems powering high-performance infrastructure. This...Flexible hours$90k - $100k
...4 technical escalation support for mission-critical emergency response systems, playing a vital role in our commitment to "Solving for safer".Job Description We are seeking a Sr. Systems Engineer to serve as a high-level technical escalation point for our Next-Generation...Remote workRelocationShift workRotating shift$200k - $300k
...Base pay range: $200,000.00/yr - $300,000.00/yr Direct message the job poster from Acceler8 Talent Senior Neuro-Symbolic Systems Engineer - San Francisco, CA A company building AI systems that can interact with the physical world at scale – designing experiments, controlling...Full timeImmediate start- ...handling What We’re Looking For Strong experience building agent systems, LLM tool-calling pipelines, or orchestration frameworks Deep... ...production, not just prototypes Work on complex, high-impact engineering workflows with real constraints High ownership and influence...
- ...research and product development. We oversee large-scale systems that span data centers, GPUs, networking, and more,... ...growth. About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and...Full time
$196k - $248k
...autonomously driving over 100 million miles on public roads and tens of billions in simulation across 15+ U.S. states. Waymo's Systems Engineering team works together to blend software and hardware systems in groundbreaking new ways. We set the high performance standards...Full timeRemote work- Systems Engineer @ Dedalus Labs Mission Dedalus Labs is an AI research neolab building infrastructure for AI agents. We’re building the persistent compute layer that powers the next generation of autonomous software. Our platform spans virtualization, distributed systems...Work at officeVisa sponsorshipRelocation package
- ...company that specializes in the development of unique life‑saving Medical Technologies. They are currently seeking an experienced Systems Engineer to join their dynamic team. The successful candidate will have a robust background in the medical device industry, with...
$100k - $200k
...assembly lines. You’ll be joining a team of extremely hardcore and self‑motivated engineers, scientists, and operators who focus on winning 24/7. You will develop and own entire systems from design to deployment, playing a foundational role in deploying 5000+ robots by...Remote work- ...Lyft’s success. The health and sustainability of our applications, systems, tooling and processes are critical for daily operations and for Lyft’s ability to grow. This Technical Business Systems Engineering role is responsible for supporting and sustaining Lyft's...Work experience placement
- ...things. The Team: • We're a team of passionate, mission-driven engineers and scientists from companies like Tesla, Amazon, SpaceX, and... ...and empathetically. The Role: We're looking for a Systems Test Engineer to be our front line of robot uptime at Medra Lab...Full timeContract workLocal areaRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to HPC Systems Engineer. Be the first to apply!
- operating system engineer San Francisco, CA
- space systems engineer San Francisco, CA
- computer system validation engineer San Francisco, CA
- system performance engineer San Francisco, CA
- ground systems engineer San Francisco, CA
- system engineer contract San Francisco, CA
- senior linux systems engineer San Francisco, CA
- digital communications systems engineer San Francisco, CA
- director systems engineering San Francisco, CA
- sr systems engineer San Francisco, CA



