Staff AI Infrastructure Engineer
$241k - $331kBiohub
Biohub is the first large-scale initiative bringing frontier AI models, massive compute, and frontier experimental capabilities under one roof. We're building a general-purpose system to accelerate scientific discovery, integrating frontier AI models, biological foundation models, and lab capabilities, with the ultimate goal of curing disease. Our technology powers scientists around the world, translating AI capabilities into tools that accelerate research everywhere. The Team The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization. The Opportunity CZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping people, not optimizing ad revenue. What You'll Do Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes. Build the systems, automation, and processes that keep clusters healthy, and that enable fast, efficient recovery when things break. Debug and resolve deep infrastructure failures across storage, networking, scheduling, and GPU compute layers. Build the tooling and operational patterns that make these failures easier to detect, diagnose, and prevent. Design and execute GPU cluster scaling plans, systematically validating storage, networking, interconnect, and scheduler behavior as clusters grow to support larger training runs. Build automation and tooling to manage cluster operations at scale: capacity planning, GPU utilization monitoring workload manager policy management, and pod lifecycle automation. Drive configuration-as-code practices, ensuring cluster state is reproducible and auditable, and managed through version-controlled pipelines. Collaborate directly with AI researchers and hero run leads to understand training workload patterns and design infrastructure that meets frontier-scale requirements. Own the vendor relationship on technical issues - escalating SEV1s, coordinating across multiple partners and network backbone teams, holding them accountable to root/proximate cause analysis and SLAs. Contribute to capacity planning: projecting GPU demand, managing cluster expansion across GPU generations, and coordinating multi-cluster strategy. Improve operational resilience, reducing mean time to detect and resolve incidents, reducing toil through automation, and developing runbooks that scale the team's operational knowledge beyond any individual. What You'll Bring 8+ years of AI/ML infrastructure engineering experience, with deep expertise in at least one of: HPC/Slurm cluster operations, Kubernetes at scale, distributed systems debugging, or GPU compute infrastructure. Strong Linux systems fundamentals - networking (TCP/IP, InfiniBand, RDMA, MTU/MSS/PMTUD), storage (NFS, VAST, WEKA, POSIX semantics), kernel internals (cgroups, namespaces, eBPF, sysctls). Hands-on experience with Kubernetes and cloud-native infrastructure - pod lifecycle, CNI plugins (Cilium preferred), StatefulSets, Helm, ArgoCD, or equivalent GitOps tooling. Experience with HPC workload managers - Slurm strongly preferred (QoS, partitions, preemption, accounting, Sunk/CoreWeave patterns a plus). Debugging instinct: ability to form hypotheses quickly, design controlled experiments, and root cause complex multi-system failures under pressure. You enjoy finding the hard bugs. Proficiency in Python and Bash for automation and tooling. Go, Rust, or C/C++ a plus. Experience with observability stacks - Prometheus/VictoriaMetrics, Grafana, DCGM metrics, distributed tracing. You know how to instrument systems you don't control. Excellent communication - you can write a crisp incident summary for researchers, a technical escalation to a vendor CTO, and a system design doc for teammates, all in the same day. Bonus: experience with distributed AI training infrastructure (NCCL, PyTorch DDP, multi-node job debugging, checkpoint/restart patterns, container environments for large-scale training). Compensation The Redwood City, CA base pay range for a new hire in this role is $241,000 - $331,000 New hires are typically hired into the lower portion of the range, enabling employee growth in the range over time. Actual placement in range is based on job-related skills and experience, as evaluated throughout the interview process. Better Together As we grow, we're excited to strengthen in-person connections and cultivate a collaborative, team-oriented environment. This role is a hybrid position requiring you to be onsite for at least 60% of the working month, approximately 3 days a week, with specific in-office days determined by the team's manager. The exact schedule will be at the hiring manager's discretion and communicated during the interview process. Benefits for the Whole You We're thankful to have an incredible team behind our work. To honor their commitment, we offer a wide range of benefits to support the people who make all we do possible. Provides a generous employer match on employee 401(k) contributions to support planning for the future. Paid time off to volunteer at an organization of your choice. Funding for select family-forming benefits. Relocation support for employees who need assistance moving If you're interested in a role but your previous experience doesn't perfectly align with each qualification in the job description, we still encourage you to apply as you may be the perfect fit for this or another role. #LI-Hybrid Biohub
- Staff AI Infrastructure Engineer You'll own the reliability of Luma's 10k+ GPU fleet: the scheduling, efficiency, and resilience that research and products depend on. As a Staff AI Infrastructure Engineer, you'll be a technical authority who turns deep systems knowledge...Suggested
- The Phoenix Group is seeking an experienced AI Software Engineer to design and deliver AI-powered applications for production environments in Redwood City. You will own projects end-to-end, partnering with technical and business teams to build scalable, reliable systems...Suggested
- ...Description Job Description Palona’s AI agents operate continuously in... ...systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...SuggestedTemporary work
$190k - $260k
...has developed an artificial intelligence (AI) powered technology stack purpose-built... ...large-scale world models – depends on infrastructure that turns thousands of hours of multimodal... ...training throughput. We are looking for engineers who make model training fast: streaming...SuggestedTemporary workWork at officeVisa sponsorshipFlexible hours- ...Senior Lead Software Engineer Be an integral part of an agile team that's constantly... ...JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part... ...scalable cloud platforms optimized for AI/ML workloads. Partner with AI teams...SuggestedFor contractors
- ...Senior Lead Software Engineer Be an integral part of an agile team that's constantly... ...JPMorgan Chase within the Corporate Sector, Infrastructure Platforms team, you are an integral part... ...scalable cloud platforms optimized for AI/ML workloads. Partner with AI teams...For contractors
$151.3k - $283.8k
...technological advancements such as cloud, AI, and network security. While driving the digitalization... ...of emerging technologies within cloud infrastructure.Who We Look For1.Education: Master’s or Ph.D. degree in Computer Engineering, Electronic Engineering, Microelectronics,...Full timeRelocation package- Notable is the leading healthcare AI platform for transforming... ...growth without hiring more staff.We are on a mission to improve... ...healthcare. As a Senior AI Platform Engineer, you will design, build, and... ...and Helm Charts for infrastructure and deployment.Google Cloud Platform...Full timeTemporary workWork at officeRemote work3 days per week
$120k - $220k
...news and information powered by advanced AI, recommendation systems, and adtech.... ...team to fulfill our mission: building the infrastructure layer for content intelligence. If you... ...hiring our first dedicated Agent Platform engineer to own this layer end-to-end. You'll...Full timeLocal areaWork from home- ...JPMorgan Chase & Co. in Palo Alto is seeking a Senior Director of Software Engineering to shape data and AI platforms within the Commercial and Investment Bank Operations team. You lead multiple technical areas, manage departments, and collaborate across Data and AI...Bank staff
- ...JPMorganChase is seeking a Senior Director of Software Engineering to lead Data and AI platforms within the Commercial and Investment Bank Operations group. You will oversee multiple departments, set technical strategy, and drive enterprise-wide AI-enabled engineering...
- ...mission is simple: Improving patient outcomes. As a market leader in AI-driven, data-powered, and privacy-compliant healthcare... ...hey, we are not hiring for A/B testing. As an AI-Forward Engineer , you will be at the frontier of how AI transforms every function...Live outWork at officeFlexible hoursWeekend work
- ...Role: AI Engineer Location: Bellevue, WA. Let's create our future together at The AES Group! About The AES Group The AES Group is a premier technology and engineering consulting company that has been bringing businesses and talent together...Casual work
- ...Vehicles, Scania and MAN. Here in the US, we are blending German engineering with American ingenuity. As ADMT, we develop and realize fully... ...has the main focus areas: to know and understand the future of AI and to derive requirements on sensor technologies and compute...Local areaWorldwide
$133.4k - $222.3k
...We're hiring a Staff AI Engineer to build GenAI and voice agents for medical devices, deployed both on-device and in the cloud. You'll own the technical direction for these systems — connecting clinical use cases to the models behind them (ASR, TTS, SLMs, speech-to...$170k - $225k
...and dedication. Come join Dexterity and help make intelligent robots a reality!About the RoleWe’re looking for an Senior/Staff AI Algorithms Engineer with deep foundations in machine learning, reinforcement learning, and optimization—and a strong drive to apply those...Worldwide- Cognichip is seeking a versatile Sr. Staff AI Engineer in Redwood City, California, to integrate Artificial Intelligence into the semiconductor design lifecycle. This role involves bridging ML research and hardware engineering, facilitating productivity gains in chip design...
- ...Citizens ONLY - NO SUBS, NO Subcontracting, NO outside firms. AI Engineer – AI Agents & Generative AI Mid-to-Senior Level- Hands-on... ...and/or Kubernetes . Familiarity with Terraform or Infrastructure as Code. Understanding LLM security, governance, observability...
- Box, headquartered in Redwood City, CA, is seeking a Solutions Engineer to empower sales teams with technical solutions across Box AI and the full platform. You will own technical wins, craft customer-centric demonstrations, and work with cross-functional teams to deliver...
- ...EngineerCompany OverviewA leading artificial intelligence and advanced analytics organization is seeking a Data & Knowledge Engineer to help power next-generation AI and decision-support platforms. Based in Los Angeles, California, the company specializes in integrating complex...
$108k - $170k
About Us Observe.AI is the AI Agents platform for customer experience, designed to help organizations deliver faster, smarter, and... ...customer interaction. Why Join Us We're looking for an AI Agent Engineer to lead the charge in building and deploying enterprise-grade...Full timeWork at officeLocal areaRemote workFlexible hours- ABOUT RETELL AI Retell AI is using first-principles thinking to reimagine the call... ...Airlines, Lenovo, and Grab. We're building the infrastructure that powers millions of AI-powered phone... ...started. We're hiring an Applied AI Engineer to work directly with enterprise...Full timeH1b
$70 - $80 per hour
...MatchPoint Solutions is a fast-growing, young, energetic global IT-Engineering services company with clients across the US . We provide... ...We look forward to hearing from you! Job Title: Embedded AI Engineer Location: Menlo Park , CA Employment Type: 6...Contract workLocal area- Retell AI is hiring an Applied AI Engineer to work with enterprise customers to design, build, and deploy production AI voice applications. You'll sit at the intersection of AI, software engineering, and customer success, partnering with customer teams to integrate Retell...Full time
- Sr. Staff AI Engineer, Silicon Design Position Overview We are seeking a versatile Sr. Staff AI Engineer to drive the integration of Artificial Intelligence into the semiconductor design lifecycle. In this role, you will bridge the gap between advanced ML research and...
- ...Corp. in California is seeking an experienced Field Applications Engineer to partner with customers, sales, product and engineering. From... ...early product evaluation to production, you will help deploy edge AI and sensor solutions, lead demos, workshops, and training, and...
- ...optimization of scalable Artificial Intelligence ("AI") and Machine Learning ("ML") solutions.... ...cross-functional teams (Product, Data, Engineering) to develop and implement AI/ML... ...Design, build, and manage scalable ML infrastructure Develop, validate, deploy, and monitor ML...Full timeLocal areaRemote workFlexible hours
- ...Software Engineer – SaaS Platform How often do you get the chance to make a global impact developing the latest AI inside of the "built world"? Reconstruct's Visual Command Center... ..., and coding standards Ensure infrastructure-as-code is in place for reliable, repeatable...Work at office
- ...C3 AI (NYSE: AI), is the Enterprise AI application software company. C3 AI delivers... ...for a highly motivated Senior Software Engineer - Platform to join our Platform Engineering... ...organization. Develop and maintain infrastructure automation to provision, manage, and...
$175k - $225k
...industry's leading vector database for enterprise-grade AI. Founded by the engineers behind Milvus, the world's most popular open-source vector... ...efficient Work deep across Kubernetes, multi-cloud infrastructure, networking, storage, and database engine runtimes to deliver...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff AI Infrastructure Engineer. Be the first to apply!
- engineering aide Redwood City, CA
- technology administrator Redwood City, CA
- senior staff systems engineer Redwood City, CA
- staff engineer Redwood City, CA
- assistant engineer Redwood City, CA
- software engineer staff Redwood City, CA
- ai engineer Redwood City, CA
- ai developer Redwood City, CA
- remote infrastructure engineer Redwood City, CA
- infrastructure engineer Redwood City, CA



