Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Full-time

Together AI

Role Description

In this role, you will design and deliver multi-petabyte storage systems purpose-built for the world’s largest AI training and inference workloads. You’ll architect high-performance parallel filesystems and object stores, evaluate and integrate cutting-edge technologies such as WekaFS, Ceph, and Lustre, and drive aggressive cost optimization—routinely achieving 30-50% savings through intelligent tiering, lifecycle policies, capacity forecasting, and right-sizing.

You will also build Kubernetes-native storage operators and self-service platforms that provide automated provisioning, strict multi-tenancy, performance isolation, and quota enforcement at cluster scale. Day-to-day, you’ll optimize end-to-end data paths for 10-50 GB/s per node, design multi-tier caching architectures, implement intelligent prefetching and model-weight distribution, and tune parallel filesystems for AI workloads.

Qualifications

  • 8+ years in storage engineering with 3+ years managing distributed storage at multi-petabyte scale
  • Proven track record deploying and operating high-performance storage for GPU/HPC clusters
  • Deep Kubernetes and cloud-native storage experience in production environments
  • Strong coding skills in Go and Python with demonstrated ability to build production-grade tools
  • BS/MS in Computer Science, Engineering, or equivalent practical experience
  • History of technical leadership: designing systems that significantly improved performance (>3x), reliability (99.9%+ uptime), or cost efficiency
  • Distributed Storage Systems: Deep expertise in WekaFS, Lustre, GPFS, BeeGFS, or similar parallel filesystems at multi-petabyte scale
  • Object Storage: Production experience with S3, MinIO, Ceph, or R2 including performance optimization and cost management
  • Kubernetes Storage: CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers
  • Storage optimization for GPU workloads, RDMA/InfiniBand networking, parallel filesystem optimization (100+ GB/s aggregate cluster throughput)
  • Programming: Go and Python for automation, operators, and tooling
  • Infrastructure as Code: Terraform, Ansible, Helm, GitOps (ArgoCD)
  • Linux Storage Stack: Advanced knowledge of filesystems (ext4, xfs), LVM, NVMe optimization, RAID configurations
  • Observability: Prometheus, Grafana, Thanos architecture and operations

Requirements

  • GPU Direct Storage (GDS), NVMe-oF, storage networking (100GbE/400GbE)
  • ML/AI storage patterns (model weights, checkpointing, dataset caching)
  • Kubernetes operator development (controller-runtime, kubebuilder)
  • Storage snapshots, cloning, and thin provisioning
  • Backup and disaster recovery (Velero, Restic, cross-region replication)
  • Storage encryption (at-rest and in-transit), security and compliance
  • Storage benchmarking and profiling tools (fio, iperf3, iostat, blktrace)

Benefits

  • Competitive compensation
  • Startup equity
  • Health insurance
  • Flexibility in terms of remote work

Company Description

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Staff Engineer, Distributed Storage and HPC & AI Infrastructure in Remote vacancy
  • $189k - $301k

     ...communities.About the RoleWe are seeking a Senior Staff Engineer to build and optimize the EDA design environment and large-scale compute infrastructure that supports our semiconductor design...  ..., we use Artificial Intelligence (AI) tools in the recruitment process to enhance... 
    Suggested
    Contract work
    Work at office
    Remote work
    Flexible hours

    Samsung Semiconductor

    San Jose, CA
    2 days ago
  • $180.1k - $278.7k

     ...candidates outside of these locations.The Staff Infrastructure Reliability Engineer is responsible for the technical...  ...of Redfin’s production database and storage systems. They will work with the...  ...You will use and evangelize approved AI code generation tools to document, architect... 
    Suggested
    Minimum wage
    Full time
    Immediate start
    Remote work

    Rock central

    Seattle, WA
    3 days ago
  • $207k - $340k

     ...and approval. We’re hiring a Principal Staff Software Engineer to lead LinkedIn’s GPU-Based Retrieval Platform, a foundational AI infrastructure stack that powers candidate generation...  ...indexing and retrieval, low-latency distributed serving, GPU scheduling, memory efficiency... 
    Suggested
    For contractors
    Work at office
    Remote work
    Work from home
    Flexible hours

    Linkedin

    Mountain View, CA
    2 days ago
  • Senior/Staff Backend Engineer - Distributed System Zettabyte Inc About Us At Zettabyte ,...  ...we’re on a mission to make AI compute ubiquitous,...  ...of the team designing the infrastructure for the AI-first world. Why...  ...Bonus Qualifications GPU or HPC cluster management experience... 
    Suggested
    Hourly pay
    Work at office
    Work from home
    Visa sponsorship

    Zettabyte Inc

    Palo Alto, CA
    3 days ago
  •  ...collaboration and AI-powered workflow software...  ...real. We are a distributed team of builders...  ...Onebrief's infrastructure team owns the platform...  ...to military staffs. Our charter spans...  ...an Infrastructure Engineer who builds security...  ...operators, networking, storage, multi-cluster... 
    Suggested
    Remote work

    Onebrief, Inc

    United States
    5 days ago
  • $276.5k - $300k

     ...Staff Infrastructure EngineerTools for Humanity (TFH) designs and builds technology...  ...people in the age of AI. As bots and autonomous...  ..., AI, cryptography, mobile engineering, and global operations. Our...  ...to work effectively with a distributed team and patiently support... 
    Flexible hours

    Tools for Humanity

    San Francisco, CA
    2 days ago
  • $189.3k - $290.7k

     ...driving? Join the Embodied AI team at General Motors. Our...  ...real-world scenarios.As a Staff ML Infra Engineer, you will drive the development...  ...building large-scale distributed systems, applications, or advanced...  ...systems on modern cloud infrastructure‑performance End-to-end... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    4 days ago
  • $185k - $335.3k

     ...driven expert in ML Training Infrastructure with a demonstrated ability...  ..., and high-performance AI/ML platform infrastructure...  ...development at scale.As a Staff ML Engineer, you will operate as a technical...  ...efforts across distributed training workflows, improving... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Sunnyvale, TX
    8 hours ago
  •  ...upgrades, recovery, and decommissioning. Use AI to automate and accelerate infrastructure delivery and operations. Provision dedicated...  ...MetalLB, VLAN, VXLAN, BGP, and ECMP. Configure distributed and shared storage for high-performance workloads. Build... 
    Full time

    fal

    Remote
    14 days ago
  •  ...that powers dashboards, ML, and AI across the company. We're seeking a Staff Engineer to architect, operate, and...  ...AI-native team of seasoned data infrastructure engineers, ship continuously, and...  ...languages) with a robust systems and distributed computing mindset. ~... 
    Full time
    Remote work

    Shopify

    Remote
    17 days ago
  • $126k - $204.5k

     ...Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do...  ...SummaryWe're looking for an Infrastructure Engineer to build developer tooling that enables...  ...productivity and reliability across a distributed cloud environment. We operate in a 100... 
    Full time
    Local area
    Remote work

    Palo Alto Networks

    New York, NY
    3 days ago
  •  ...Own server deployments, infrastructure automation, software...  ...infrastructure and/or security engineering experience for the...  ...level, 6+ years for staff, or 8+ years for...  .... ~ Experience with distributed systems, secure multi-...  ...~ Familiarity with AI tools for debugging, documentation... 
    Permanent employment
    Full time
    Work at office
    Home office
    Relocation package
    Flexible hours
    2 days per week
    1 day per week

    AllSpice

    Boston, MA
    9 days ago
  •  ...announcement here! As a Sr./Principal Infrastructure Engineer, you will improve our cloud...  ...engineering experience [Staff] 6+ years of cloud...  ...on the details Fluency with AI tools — you use them day-to-...  ...Kubernetes and/or Helm experience Distributed systems design and operation... 
    Permanent employment
    Home office
    Relocation package
    Flexible hours
    2 days per week
    1 day per week

    Linuxconfig

    Boston, MA
    3 days ago
  •  ...growth at a massive scale. As a Staff Machine Learning Infrastructure Engineer , you will own the technical...  ...offers an opportunity to influence key AI-driven systems across Reddit while...  ...teams to build high-performance, distributed training systems that efficiently... 
    Full time
    Flexible hours

    Reddit

    Remote
    17 days ago
  • $148.2k - $222.2k

    Staff Engineer, Infrastructure Platforms Position SummaryThe Staff Engineer, Infrastructure Platforms...  ..., virtualization, enterprise storage, High Performance Computing (HPC), and modern platform services....  ..., genomics, bioinformatics, AI/ML, or other scientific computing... 
    Full time
    Work from home
    Monday to Friday

    Pacific Biosciences

    Menlo Park, CA
    8 hours ago
  •  ...looking for a Core Systems Engineer with a deep mastery of Rust...  ...’s underlying model serving infrastructure and real-time performance engines...  .... ~Debug complex distributed system deadlocks, race conditions...  ...systems, work with bleeding-edge AI workflows, and fundamentally... 
    Full time

    Sequen AI

    Remote
    a month ago
  •  ...leading technology company for AI and Bitcoin mining infrastructure. Bitdeer is committed to...  ...Overview ~ We are seeking a Staff Slurm Cluster & HPC Scheduling Engineer to own Slurm as a first-class...  ...delivery, shared parallel-storage mounts, and login pods running... 
    Full time
    Local area
    Shift work

    Bitdeer Technologies Group

    Remote
    4 days ago
  •  ...interconnects and storage fabrics. ~Modernize...  ...with Principal Engineers, hardware teams,...  ...Strong knowledge of distributed systems, parallel...  ...Requirements ~7+ years in HPC or storage...  ...storage, or AI/ML I/O workloads (...  ...~Want to work on infrastructure that operates at real... 
    Full time

    DDN

    Remote
    a month ago
  •  ...the cloud platform that every engineer at Headway deploys on. Make...  ...deploys boring, scaling automatic, infrastructure self-serve, and cost...  ...and Eddy, Headway's internal AI platform. We own the infrastructure...  ...platform, and bring Staff-level influence to an area that... 
    Remote work
    Shift work

    Headway - Design & Development

    United States
    1 day ago
  •  ...partner is looking for a Staff Engineer, Platform & Infrastructure based in United States....  ...including compute, networking, storage, and core data services....  ...in a remote-first, distributed environment. ~ Proactive...  ...Jobgether works: We use an AI-powered matching process... 
    Full time
    Remote work

    jobgether

    United States
    4 days ago
  •  ...Staff Platform Engineer (Infrastructure) We are seeking an experienced Staff Platform Engineer to help build...  ...encoded as self-verifying runbooks that AI agents execute. We are not hiring...  ...incident-response experience on distributed-systems failures — databases, streaming... 
    Full time

    Postscript

    Remote
    a month ago
  •  ...Bay Area; Austin, TX; Washington, DC Position Summary As a Staff Infrastructure Engineer, you will work with your team on various challenging...  ...GitHub Actions, pipeline optimization Experience with emerging AI technologies (model deployment, data compliance) Benefits... 
    Apprenticeship
    Local area
    Remote work
    Flexible hours

    CloudDevs

    Seattle, WA
    1 day ago
  • $180k - $250k

     ...applications and next steps. Our partner is looking for a Senior/Staff Network Infrastructure Engineer based in United States. This is a senior-level...  ...operating the network foundation behind a rapidly scaling AI and generative media platform. You will build... 
    Full time
    Remote work

    jobgether

    United States
    2 days ago
  • $127k - $249k

     ...an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build...  ...redefined the data platform for the AI era, enabling builders to create,...  ...the most widely available, globally distributed data platform on the market, helps... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  •  ...highly experienced Senior Staff Engineer specializing in AI Data Path & Storage to lead hands-on...  ...across GPU, memory, and distributed storage layers. ~Involve...  ...production-grade AI infrastructure. Qualifications...  ...performance computing (HPC) or hyperscale distributed... 
    Full time

    DDN

    Remote
    a month ago
  • $157k - $234k

     ...Description Job Description Waabi, founded by AI visionary Raquel Urtasun, is the leader...  ...- Collaborate with data scientists and engineers to understand model requirements and...  ...Bonus/nice to have:  - Experience with infrastructure-as-code (IaC) tools such as Terraform or... 
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    Dallas, TX
    2 days ago
  •  ...Staff Software Engineer We are looking for a hands-on Staff Software Engineer...  ...and transformed the storage industry. What Will You...  ...software-defined clustered, distributed, or cloud-native solutions...  ...between our proprietary Go infrastructure and open-source tools.... 
    Remote work

    DDN Storage

    United States
    1 day ago
  • $180k - $220k

     ....Learn about the Danaher Business System which makes everything possible.The Staff Platform Infrastructure Engineer will be instrumental in building and scaling our next-generation data and AI infrastructure — driving technical innovation across all of Danaher. In this... 
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours

    Danaher Corporation

    New York, NY
    8 hours ago
  • $252k - $315k

    About Scale AIScale AI is the data foundation for AI,...  ...remains one of the hardest engineering challenges.As a Staff Frontier Agent Engineer (...  ....Collaborate with infrastructure engineers to deploy AI systems...  ...EngineeringExperience building distributed production systems.... 
    Full time

    Scale AI

    New York, NY
    3 days ago
  •  ...TXTechnology - Security /RemoteThe Staff Security Engineer will be responsible for...  ...hands-on proficiency with Infrastructure as Code (Terraform or...  ...and deploying large-scale distributed systems at scaleExperience...  ...and cloud-native workloads. AI/ML Security & Threat Modeling... 
    Temporary work
    Remote work
    Flexible hours

    Aledade

    Austin, TX
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Engineer, Distributed Storage and HPC & AI Infrastructure. Be the first to apply!