HPC Systems Engineer
University of California , San Francisco
Certain terms and conditions of employment for this position, including the rate of pay, benefits, etc., are currently subject to negotiation with the appropriate union The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day operations of the Institute's HPC clusters. The HPC Systems Engineer will: Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with unique computing, networking, and storage systems designed to address cutting-edge research problems Apply their engineering and design skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity of CI systems. Select methods, techniques and evaluation criteria to develop new CI solutions to address complex research problems. Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and availability. Provides consultation, support, and guidance to researchers on how to address computational problems using standard tools, packages, and approaches. Develop enhancements of monitoring to maintain the integrity of CI systems. Participate in multiple technical projects simultaneously. Applies working knowledge of security control frameworks to maintain the integrity of the CI systems and the research being performed on them. Gives presentations to the associated team and other technical units. Evaluates new technologies, including performing moderate to complex cost/benefit analyses. This position may lead to cross-functional technical working groups and projects in support of onboarding research customers, or making systems improvements. Department Overview Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA-compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers' needs. The Core HPC team of the Academic Research Service (ARS) focuses on large-scale, high-performance computational and storage services for UCSF researchers so they can address complex computational, AI, and data science problems. % of time Essential Function (Yes/No ) Key Responsibilities (To be completed by Supervisor) 15 Applies advanced systems / infrastructure concepts to define, design, implement, and operate highly complex, research cyberinfrastructure systems, services and technology solutions. Proposes and implements highly complex system or device enhancements such as software, hardware and network configuration, updates and installations for projects or services of broad scope. Sets standards for monitoring and maintaining the health and integrity of CI systems including upgrading and patching. 15 Independently manages systems and services for a large facility, campuswide, medical center or Office of the President and / or institution-wide scope and makes recommendations for purchases or upgrades. Performs complex and advanced analysis to acquire, install, modify and support operating systems, databases, utilities and web-related tools. Selects methods and techniques to obtain solutions. Interacts with senior management. May perform complex network integration tasks and interoperability assessments for interconnected servers or components of clusters for communication. Support and collaborate with researchers and other key IT (e.g. network and security) and Data Center partners in a timely manner 15 Specifies, writes and executes highly complex software and scripts to support systems management, log analysis, monitoring, deployment, configuration management, and other system administration duties for multiple, highly integrated systems. 30 Provides consultation, training, support, and guidance to researchers enabling them to utilize HPC resources effectively. 10 Maintains complex security systems. Interprets and adopts campus, medical center or Office of the President, system and regulation-based security policies to control access to networked resources. Provides recommendations and requirements on network access controls. 5 Collaborates and may provide leadership with other Systems Engineers within the CI ecosystem/higher-education community. Regularly contribute best practices documentation, present at conferences, or publish in peer reviewed journals. 10 Define and track performance metrics to ensure efficient current and future use of cyber infrastructure resources. 100% (To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.) REQUIRED QUALIFICATIONS Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems Expert knowledge of HPC systems infrastructure design Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements Demonstrated testing and test planning skills. Demonstrated ability to create automated testing. Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters Ability to elicit and communicate technical and non-technical information in a clear and concise manner. Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines. Understanding of system performance monitoring and actions that can be taken to improve or correct performance. Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems. Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration. Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner PREFERRED QUALIFICATIONS Knowledge of the design, development, and application of technology and systems to meet business needs. General knowledge of other areas of IT. Thorough understanding of and experience with systems-related issues and actions that can be taken to improve or correct performance. Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions. About UCSF The University of California, San Francisco (UCSF) is a leading university dedicated to promoting health worldwide through advanced biomedical research, graduate-level education in the life sciences and health professions, and excellence in patient care. It is the only campus in the 10-campus UC system dedicated exclusively to the health sciences. We bring together the world's leading experts in nearly every area of health. We are home to five Nobel laureates who have advanced the understanding of cancer, neurodegenerative diseases, aging and stem cells. Pride Values UCSF is a diverse community made of people with many skills and talents. We seek candidates whose work experience or community service has prepared them to contribute to our commitment to professionalism, respect, integrity, diversity and excellence - also known as our PRIDE values. In addition to our PRIDE values, UCSF is committed to equity - both in how we deliver care as well as our workforce. We are committed to building a broadly diverse community, nurturing a culture that is welcoming and supportive, and engaging diverse ideas for the provision of culturally competent education, discovery, and patient care. Additional information about UCSF is available here. Join us to find a rewarding career contributing to improving healthcare worldwide. Equal Employment Opportunity The University of California is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, protected veteran status, or other protected status under state or federal law. Salary Information The final salary and offer components are subject to additional approvals based on UC policy. Your placement within the salary range is dependent on a number of factors including your work experience and internal equity within this position classification at UCSF. For positions that are represented by a labor union, placement within the salary range will be guided by the rules in the collective bargaining agreement. To learn more about the benefits of working at UCSF, including total compensation, please visit: REQUIRED QUALIFICATIONS Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems Expert knowledge of HPC systems infrastructure design Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc. Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements Demonstrated testing and test planning skills. Demonstrated ability to create automated testing. Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters Ability to elicit and communicate technical and non-technical information in a clear and concise manner. Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines. Understanding of system performance monitoring and actions that can be taken to improve or correct performance. Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems. Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration. Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner PREFERRED QUALIFICATIONS Knowledge of the design, development, and application of technology and systems to meet business needs. General knowledge of other areas of IT. Thorough understanding of and experience with systems-related issues and actions that can be taken to improve or correct performance. Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions.
$166k - $343k
Senior Presales, Systems EngineerThis role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week... ...a wide range of customer segments.The Sr. Presales Systems Engineer plays a critical role in driving the success of HPE Networking...SuggestedFull timeWork experience placementWork at officeLocal areaImmediate start2 days per week- ...includes hundreds of modern software companies building the next generation of enterprise-ready products. About the Role As a Systems Engineer at WorkOS, you will be the technical backbone of our internal IT organization — designing the systems, automations, and...SuggestedRemote work
$100k - $200k
...performance across the robotics software stack by diagnosing system bottlenecks and failure modes. Implement and maintain software... ...Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, Robotics, or a related technical field, or equivalent...SuggestedRemote work$150k - $250k
...already has millions in recurring revenue is building next‑generation AI infrastructure for video intelligence and is looking for a Systems Engineer to join their team. What You'll Be Doing Design and engineer systems that handle compute, scheduling, and orchestration of...SuggestedFull timeImmediate startSleeping nights$175k - $300k
...high intensity to push the frontier forward. The Production Engineering Team Examples of key exciting problems the team is working... ...interface for the whole company, not a hundred scripts. Make the system's view of itself always match reality: integrate fleet state...SuggestedLocal area- ...explorers. We are motivated by the challenge of solving tough engineering problems, committed to creating sustainable change, and driven... ...defense missions — including space-based IR sensing and interceptor systems. If you join the team, you will have the opportunity to work...Flexible hours
- ...Distributed Systems EngineerAs a distributed systems engineer, you'll work across the stack to solve problems as they come up and help build Archil volumes. You'll have significant influence over the technical and product direction.We'll expect you to be able to:Be oncall...Flexible hours
- ...Job Description Job Description CommVault Systems Engineer (Data Protection / Backup) Employment Type: Full-Time, Experienced Department: Technology Support CGS is seeking an experienced CommVault Data Protection Engineer with extensive knowledge and experience...Full timeFlexible hours
$152k
...About the role We are seeking an experienced Senior Systems Engineer to own our macOS platform and drive endpoint management and device trust standards across Chime’s IT ecosystem. As a Senior Systems Engineer, you will drive multi-system initiatives across our IT domains...Full timeShift work$190k - $230k
...We are seeking a full-time Senior Systems Engineer to enhance the performance and efficiency of our deployments. You will be in charge of owning the entire end to end system level performance for our deployments. In this role, you will collaborate closely with teams across...Full timeImmediate start$175k - $308.5k
...to join a team redefining hardware technology? We are searching for dedicated engineers to join our exciting team of innovators. In this role, you will be a key member of the RM modeling system engineering team working on modem systems features and requirements....$160k - $185k
...sensors for precision mapping, robotics, automotive, security systems, smart cities and various industrial solutions. We have transformed... ...your help! The Role We're seeking a proactive and skilled engineer who thrives in a dynamic, fast‑paced environment to be an...Work experience placementLocal area- ...autonomous software. Our platform spans virtualization, distributed systems, storage, networking, scheduling, orchestration, and low-level... ...and scheduling decision matters. We are looking for systems engineers who want to understand computers all the way down and build...Full timeWork at officeVisa sponsorshipRelocation package
- ...data, run data applications, so they can spend more time putting knowledge into action. We’re looking for engineers who want to build the operating system for AI Data Applications and Workflows. About the role We're looking for experienced distributed systems...
- ...software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding team... ...on real-world tasks. About the Role We're hiring a GTM Systems Engineer to design, build, and scale the technical...
$200k - $300k
...Building an AI agent that can securely take action inside enterprise systems is hard. The moment an agent accesses customer data, executes... ...a user, authorization, governance, and trust become the real engineering challenge. Arcade is the MCP runtime that gives agents the...Work at officeShift work- ...System Engineer Location: Foster City, CA OR San Francisco, CA (Hybrid) Role Summary: The Systems Engineer will be responsible for developing system requirements, test cases, test automation & power train systems issue RCA support for an autonomous vehicle....Shift work
$175k - $308.5k
...RF Modeling Systems Engineer At Apple, new ideas have a way of becoming extraordinary products and customer experiences. Bring passion and dedication to your job, and there's no telling what you could accomplish. Dynamic environment, inquisitive people, and innovative...Relocation$190k - $230k
...high-performing team that believes in each other, come build with us at Crusoe. About This Role We’re seeking a Senior Systems Engineer to play a key role in executing Crusoe’s 2026 Enterprise AI Strategy. In this role, you will design and build agentic AI systems...Temporary work- ...Job Title4-yr Technical Degree 8+ yrs IT/Engineering 4+ yrs Windows Server/Systems Administration 3+ yrs of Palo Alto Network Security 3+ yrs of O365 Administration 3+ yrs of PowerShell and Python scripting 2+ yrs of Dell Storage, Cisco UCS (servers), Meraki, VMware, Cisco...
$200k - $300k
...Base pay range: $200,000.00/yr - $300,000.00/yr Direct message the job poster from Acceler8 Talent Senior Neuro-Symbolic Systems Engineer - San Francisco, CA A company building AI systems that can interact with the physical world at scale – designing experiments, controlling...Full timeImmediate start$135k - $145k
...Senior Systems EngineerLyra Technology Group is a private equity-backed holding company that invests in and operates industry leading... ...long term.People 1st IT is looking for an experienced Level 3 Engineer / Senior Systems Engineer to join our technical team. This is a...Work at officeRemote workRelocation$114.2k - $142.7k
...hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a variety... ...transforming large datasets into actionable metrics.As a Space Systems Engineer, you'll support daily satellite operations, analyze...Full timeTemporary workFor contractorsWork at officeLocal areaRemote workHome office3 days per week$148.5k - $223.9k
...and you are the future of Salesforce.Role OverviewAs a Software Engineer at Salesforce, you'll join a dynamic global team responsible... ...with AI: Develop AI-based automation to reduce errors, increase system availability, and accelerate operational velocity.Build orchestration...Full time$215k - $260k
...We are seeking a Hardware Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team and close critical skill gaps in debugging... ...edge GPU architectures and how to leverage them in AI/HPC environments. Expertise supporting or designing...$180k - $300k
...Job Description Job Description Systems Engineer Company: Dedalus Labs Location: San Francisco, CA (on-site 5 days per week; relocation support available) Compensation: $180,000 - $300,000 + 0.5% - 1% equity Employment Type: Full-time Visa Sponsorship...Full timeH1bImmediate startVisa sponsorshipRelocation package$120k - $250k
...Job Description Job Description Distributed Systems Engineer @ Dedalus Labs Mission Dedalus Labs builds persistent computers for AI agents. Our flagship product, Dedalus Machines, gives agents an isolated environment where they can run software, keep files and state...Work at officeVisa sponsorshipRelocation package$100 - $120 per hour
...Benchmark , General Catalyst , Peter Thiel , Adam D'Angelo , Larry Summers , and Jack Dorsey . Position: Systems & Integration Engineering Experts Type: Contract Compensation: $100–$120/hour Location: Remote Role Responsibilities...Contract workSummer workRemote work$68.25k - $93k
...Langan provides expert land development engineering and environmental consulting services for major developers, renewable energy producers, energy companies, corporations, healthcare systems, colleges/universities, and large infrastructure programs throughout the U.S....Hourly payFull timeTemporary workWork experience placementWork at officeLocal areaWorldwideFlexible hours$134.9k - $185k
As a System Electrical Engineer at eero, you will work with the rest of the Hardware Engineering team to design, implement, debug, and characterize embedded systems for wired and wireless networking applications. You may own one or more electrical sub-systems within a...Permanent employmentLocal areaWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to HPC Systems Engineer. Be the first to apply!
- distributed systems engineer San Francisco, CA
- digital communications systems engineer San Francisco, CA
- space systems engineer San Francisco, CA
- sr systems engineer San Francisco, CA
- system engineer contract San Francisco, CA
- senior linux systems engineer San Francisco, CA
- mission system engineer San Francisco, CA
- operations support system engineer San Francisco, CA
- healthcare systems engineer San Francisco, CA
- ground systems engineer San Francisco, CA



