Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff HPC Engineer

$214k - $268k

Biohub

Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere.The TeamThe HPC Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of hybrid GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared. Our infrastructure supports day-to-day AI researcher workflows. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization.The OpportunityWe seek a Staff HPC Engineer to help lead the evolution of our advanced computing infrastructure into a next-generation hybrid HPC and AI platform. This role will help shape strategy, architecture, and operations for high-performance computing resources — including cutting-edge GPUs, large-scale storage, and high-speed networks — while enabling transformative science through AI and machine learning at scale.You will design, implement, and optimize a unified HPC-AI ecosystem blending on-prem Slurm-managed clusters, cloud GPU resources, and containerized environments. This hybrid environment will power everything from traditional HPC workloads to large AI training jobs, generative model development, real-time inference, and data-intensive pipelines.The successful candidate will be a thought leader in HPC infrastructure , capable of partnering with scientists, computational biologists, and software engineers to translate complex research needs into high-impact computing solutions. You will also foster adoption of emerging AI tools, and ensure our systems can scale to meet the demands of next-generation biomedical research.What You'll DoHPC EngineeringBuild and support a hybrid HPC-AI environment with large-scale on-prem compute/storage and elastic cloud GPU clusters (Coreweave, AWS, GCP).Architect and optimize environments for large-scale AI training and tuning, and low-latency scientific workloads.Integrate MLOps and model deployment pipelines into HPC infrastructure, ensuring reproducibility and efficiency.Implement advanced resource scheduling and orchestration (Slurm, Kubernetes, SUNK) optimized for mixed HPC and AI workflows.Operational ExcellenceSupport researchers with job optimization, GPU utilization best practices, and performance tuning for AI and HPC applications.Evaluate, deploy, and maintain AI/ML software stacks (e.g., PyTorch, TensorFlow, Hugging Face, RAPIDS) and HPC toolchains.Ensure robust data ingest, analysis, and management capabilities for AI and HPC workloads, including integration with parallel file systems and object storage.Collaboration & EnablementWork with diverse science teams to translate research requirements into hardware/software solutions, from experimental design through publication.Promote best practices for AI model training, validation, and deployment in shared computing environments.Foster a culture of shared learning by running internal workshops on HPC-AI tooling (e.g., VS Code remote dev, containerization, MLOps workflows).What You'll BringEssentialBachelor’s or advanced degree in Computer Science, AI/ML, Data Science, Systems Engineering, or related field.10+ years building and managing HPC infrastructure, with significant experience integrating AI/ML workloads.Proven track record architecting environments for large-scale GPU AI training and inference in hybrid on-prem/cloud environments.Deep expertise with HPC scheduling (Slurm), container orchestration (Kubernetes), and cloud GPU services.Strong hands-on experience with AI frameworks (PyTorch, TensorFlow, JAX) and distributed training strategies (Horovod, DeepSpeed, Ray).Knowledge of MLOps best practices, including CI/CD for ML, model registry, experiment tracking, and performance monitoring.Exceptional ability to collaborate with multidisciplinary teams and communicate complex technical concepts clearly.Demonstrated leadership in guiding infrastructure teams, influencing organizational strategy, and fostering adoption of new technologies.TechnicalAdvanced Linux systems administration, HPC networking (Infiniband, Ethernet), and storage systems administration (VAST Lustre, Weka and ZFS)Cloud platform expertise (Coreweave, AWS, GCP) including GPU provisioning, storage, and networking for AI workloads.Proficiency in automation tools (Terraform, Ansible, Puppet), containerization (Docker, Singularity), and orchestration frameworks.Strong experience debugging and troubleshooting hardware across the stack (network, GPU, compute and storage systems).Strong scripting/programming skills (Python, Bash) and familiarity with version control (Git).Experience integrating AI LLMs, AI coding assistants, and custom model development into HPC workflows.CompensationThe San Francisco, CA base pay range for a new hire in this role is for a Staff HPC Engineer 214,000–$268,000 and for a Senior Staff HPC Engineer $241,000–$300,000.New hires are typically hired into the lower portion of the range, enabling employee growth in the range over time. Actual placement in range is based on job-related skills and experience, as evaluated throughout the interview process. This position may be eligible to participate in our discretionary annual performance bonus program. Bonus eligibility and targets are determined in accordance with our total rewards philosophy and may vary by role.Better TogetherAs we grow, we’re excited to strengthen in-person connections and cultivate a collaborative, team-oriented environment. This role is a hybrid position requiring you to be onsite for at least 60% of the working month, approximately 3 days a week, with specific in-office days determined by the team’s manager. The exact schedule will be at the hiring manager's discretion and communicated during the interview process.Benefits for the Whole You We’re thankful to have an incredible team behind our work. To honor their commitment, we offer a wide range of benefits to support the people who make all we do possible. Provides a generous employer match on employee 401(k) contributions to support planning for the future.Paid time off to volunteer at an organization of your choice. Funding for select family-forming benefits. Relocation support for employees who need assistance movingIf you’re interested in a role but your previous experience doesn’t perfectly align with each qualification in the job description, we still encourage you to apply as you may be the perfect fit for this or another role.#LI-Hybrid

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Staff HPC Engineer in San Francisco, CA vacancy
  • $225k - $275k

     ...at Crusoe.About this RoleCrusoe Cloud is seeking a Senior Staff Network Deployment Engineer to serve as the technical owner of how we deploy network infrastructure...  ...rapidly scale our footprint of high-performance compute (HPC) and GPU-based AI infrastructure, you will define the... 
    Suggested
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    2 days ago
  • $209k - $253k

     ...-first cloud infrastructure, and our Compute-focused Production Engineers are the backbone of that mission. This role is centered on supporting...  ..., ensuring performance, security, and scale for modern AI and HPC workloads.What You'll Be Working On:In this role, you will... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    4 hours ago
  • $215k - $260k

     ...About This Role:We are seeking a Hardware Production / Sustaining Engineer to strengthen Crusoe’s Hardware Systems Engineering team and...  ...with cutting-edge GPU architectures and how to leverage them in AI/HPC environments.Expertise supporting or designing systems across... 
    Suggested

    Crusoe

    San Francisco, CA
    3 days ago
  • $300 per month

     ...build with us at Crusoe.About This Role:We are seeking a Staff Hardware Systems Engineer to strengthen Crusoe’s Hardware Systems Engineering team and...  ...GPU or accelerated computing infrastructure for AI/ML or HPC workloads.Hands-on experience with distributed training and... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $225k - $275k

     ...at Crusoe.About this RoleCrusoe Cloud is seeking a Senior Staff Network Production Engineer to own production reliability across our global network, including...  ...RDMA/RoCE (v1 and v2) lossless fabrics for GPU and HPC workloads, including PFC, ECN, and DCQCN tuning. Required... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    4 hours ago
  • $193k - $234k

     ...RoleCrusoe Cloud is seeking a high-energy, detail-oriented Staff Network Production Engineer to lead the physical and logical implementation of our...  ...rapidly expand our footprint of high-performance compute (HPC) and GPU-based AI infrastructure, you will be the primary... 
    Temporary work
    Remote work

    Crusoe

    San Francisco, CA
    4 days ago
  • About The RoleThe Staff Regulatory Affairs Engineer is the primary regulatory strategist and architect for our AI/ML Software as a Medical Device (SaMD) portfolio. While many companies view Regulatory Affairs as a purely administrative function, at Hinge Health we leverage... 
    Local area

    Hinge Health

    San Francisco, CA
    2 days ago
  • $142k - $196k

     ...five continents, spurring a new generation of clean alternatives to car ownership. Learn more at li.me.Lime is looking for a Staff CAE Engineer to lead the structural integrity, safety, and weight optimization of our next-generation shared micro-mobility vehicles and infrastructure... 
    Contract work
    Local area
    Remote work
    2 days per week

    LimeBike

    San Francisco, CA
    2 days ago
  • $173.5k - $331.05k

     ...the modern, AI-powered video rendering pipeline behind the next generation of these products. We are looking for a senior, hands-on engineer to own and evolve the cross-platform GPU rendering platform at the heart of that mission — leading major initiatives through to... 
    Full time
    Temporary work
    Local area
    Worldwide

    Adobe Systems

    San Francisco, CA
    4 hours ago
  • $210k - $255k

     ...be part of a high-performing team that believes in each other, come build with us at Crusoe.About the Role:We are looking for a Staff Engineer to be the detection authority for Crusoe’s Command Center platform—the person who owns what “something is wrong” means. You... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 hours ago
  •  ...balance near-term delivery with long-term platform health. Advise engineering and product leadership on technical trade-offs and their...  ...software engineering experience, with demonstrated impact at a staff or equivalent scope (leading initiatives across multiple teams... 

    Stripe

    San Francisco, CA
    4 days ago
  • $180k - $200k

     ...believe that what makes us different makes us stronger. So add your voice. Make an impact. Find your fit — and your future.As a Staff Engineer on the Order Management System (OMS) team, you will be an important technical voice. You will shape how we build, operate, and... 
    Full time
    Temporary work
    Seasonal work
    Work at office
    Immediate start
    3 days per week

    Levi Strauss & Co

    San Francisco, CA
    4 hours ago
  • $223k

     ...members on a mission to help them unlock financial progress. Growth Engineering builds the product surfaces and systems that turn that mission...  ...to be both quick to experiment on and dependable at scale.As a Staff Software Engineer, you'll operate as a subject matter expert... 
    Full time
    Work at office
    Local area
    Remote work

    Chime

    San Francisco, CA
    1 day ago
  •  ...every step of the way. Join us to invest in yourself, your career, and the financial world.The roleWe are seeking a Staff Vulnerability Management Engineer to lead the most complex technical work in SoFi’s Vulnerability Management program. You will design and build... 
    Remote work

    SoFi

    San Francisco, CA
    2 days ago
  •  ...THE ROLE As a Senior/Staff Engineer, you may work on projects that require strong execution, communication, analytical judgment, and the ability to move quickly in ambiguous environments. This posting is intended for candidates who want to be considered for this... 

    Career Launch

    San Francisco, CA
    1 day ago
  •  ...companies, including GitHub, Yelp, Paramount, and JetBlue. We're building a more trustworthy Internet. Come join us. Staff Backend Software Engineer, Monetization Staff Backend Software Engineer About the Role: Staff Backend Software Engineer - Monetization... 
    Full time
    Work at office
    Local area
    Flexible hours
    Night shift

    Fastly Inc.

    San Francisco, CA
    8 hours ago
  •  ...products offered to merchants, connecting merchant intent with consumer demand across search and discovery experiences. As a Senior Staff Engineer, you will lead the technical direction for AI-first experiences, including ranking and relevance systems that sit at the core... 
    Hourly pay
    Work at office
    Local area
    Remote work
    Flexible hours

    DoorDash USA

    San Francisco, CA
    22 hours ago
  • $156k - $190k

     ...each other, come build with us at Crusoe.About the Role:As a Staff Cloud Support Engineer, you are a technical authority within Crusoe Cloud and a...  ...You Bring to the Team:8+ years experience in SRE, DevOps, HPC, or Cloud Infrastructure roles.Advanced Linux systems expertise... 
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $212k - $265k

     ...ready to see your impact and unlock incredible career growth opportunities, join us, and build real world value.THE WORK:As a Staff Security Engineer within the Secure Digital Asset Operations (SDAO) function, you will collaborate with leadership and cross-functional... 
    Full time
    Work at office
    Local area

    Ripple

    San Francisco, CA
    3 days ago
  • $252k - $315k

     ...papers emerge every week. Yet building AI systems that reliably solve real-world problems remains one of the hardest engineering challenges.As a Staff Frontier Agent Engineer (Applied AI), you'll bridge the gap between cutting-edge AI research and production deployment... 
    Full time

    Scale AI

    San Francisco, CA
    4 hours ago
  • $208k - $260k

     ...be at the forefront of revolutionizing global payments and crafting the future of financial transactions? Join Ripple as a Staff Partner Engineer and be part of our dedicated team that crafts and deploys innovative solutions for senders, receivers, exchanges, and fund... 
    Full time
    Work at office
    Local area
    Worldwide

    Ripple

    San Francisco, CA
    2 days ago
  •  ...world’s best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to tackle technical challenges, from designing next-gen UI/UX for... 
    Worldwide

    DataBricks

    San Francisco, CA
    1 day ago
  • $134k - $184.8k

     ...The Technology, Data, and Intelligence (TDI) organization is the engine that powers Okta's global workforce, providing the technology and systems that enable our employees to do their best work.The Staff Security Engineer OpportunityWe are seeking a highly skilled and... 
    Local area
    Worldwide
    Flexible hours
    Shift work

    Okta

    San Francisco, CA
    3 days ago
  • $204k - $280.5k

     ...detection, data loss prevention, email security, corporate threat detection, compliance training, and AI governance. As our Staff Security Engineer focusing on Enterprise AI, you will own the AI governance domain within Enterprise Security, serving as IT's technical... 
    Work experience placement
    Work at office
    Local area
    Remote work
    Monday to Friday
    Flexible hours
    3 days per week

    Faire

    San Francisco, CA
    3 days ago
  • $232k - $290k

     ...see your impact and unlock incredible career growth opportunities, join us, and build real world value.THE WORK:As a Senior Staff Security Engineer focused on AI Security, you will be Ripple's deepest technical expert at the intersection of artificial intelligence and... 
    Full time
    Work at office
    Local area

    Ripple

    San Francisco, CA
    1 day ago
  • $175k - $270k

     ...teamAirwallex’s Information Security team partners closely with engineering, IT, and other stakeholders to protect our systems, data, and...  ...built into how we operate, not treated as a blocker.Your roleAs a Staff Corporate Security Engineer, you will be a critical part of... 
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  • $210k - $255k

     ...with us at Crusoe.About This RoleCrusoe is building the world’s favorite AI-first cloud infrastructure. We are seeking a Staff Corporate Security Engineer to act as the principal architect for our corporate security posture.In this role, you will move beyond tactical tool... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  • $141.7k - $250.8k

    P-1011Job Location: San Francisco Bay Area, CA As a Sr. Staff Technical Solutions Engineer and tech subject matter expert, you will partner closely with our Field and Engineering teams to deliver high-touch specialized support and tailored technical solutions for Databricks... 
    Local area
    Worldwide
    Night shift

    DataBricks

    San Francisco, CA
    4 days ago
  • $243k - $284k

     ...forefront of new technology, helping founders and their companies impact and change the world.The RoleWe're hiring a Staff Incident Response Engineer to anchor a16z's detection and response work. You'll own incident triage and response across AWS and GCP, write the detections... 
    Work at office
    2 days per week

    a16z

    San Francisco, CA
    2 days ago
  • $232k - $290k

     ...see your impact and unlock incredible career growth opportunities, join us, and build real world value.THE WORK:As a Senior Staff Security Engineer, you will be one of Ripple's most senior technical security practitioners, operating at the intersection of application... 
    Full time
    Work at office
    Local area

    Ripple

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff HPC Engineer. Be the first to apply!