Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Engineer, Distributed Storage & AI Infra

$250k - $300k
Full-time

Together AI

Salary: $250,000 - 300,000 per year Requirements:

  • We need 8+ years of experience in storage engineering, including operating distributed storage at multi-petabyte scale
  • We look for a proven background in deploying and running high-performance storage for GPU or HPC clusters
  • We require deep production experience with Kubernetes and cloud-native storage
  • We expect strong Go and Python programming skills with the ability to build production-grade systems and tooling
  • We require a BS/MS in Computer Science, Engineering, or equivalent practical experience
  • We value technical leadership that has improved performance, reliability, or cost efficiency at scale
  • We need deep expertise in distributed storage systems such as Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems
  • We expect production experience with object storage such as S3, MinIO, Ceph, or R2, including performance tuning and cost control
  • We require experience with Kubernetes storage components such as CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers
  • We need experience optimizing storage for GPU workloads and working with RDMA/InfiniBand networking and parallel filesystem performance
  • We require automation and tooling experience with Go and Python
  • We expect infrastructure-as-code experience with Terraform, Ansible, Helm, and GitOps using ArgoCD
  • We need advanced knowledge of the Linux storage stack, including ext4, xfs, LVM, NVMe tuning, and RAID configurations
  • We expect observability experience with Prometheus, Grafana, and Thanos
  • Nice-to-have experience includes GPU Direct Storage, NVMe-oF, storage networking, and RDMA implementations
  • Nice-to-have familiarity with ML/AI storage patterns such as model weights, checkpointing, and dataset caching
  • Nice-to-have experience with storage benchmarking and profiling tools such as fio, iperf3, iostat, and blktrace
Responsibilities:
  • We will define and drive the technical strategy and storage roadmap for Together AI as we expand our GPU fleet
  • We will engineer and scale multi-petabyte AI/ML storage systems by integrating Vast, Weka, and Ceph while reducing cost through automated tiering and lifecycle policies
  • We will design intelligent caching and tiered storage architectures to deliver extreme IOPS and cluster-wide throughput for training and inference workloads
  • We will tune storage isolation at the L2/L3 network layers to provide secure, production-ready multi-tenancy for storage clients
  • We will build Kubernetes storage operators and controllers that enable automated provisioning, self-service workflows, and quota enforcement
  • We will optimize end-to-end data paths to achieve 10+ GB/s per GPU node, architect multi-tier caching for model weights and datasets, tune parallel filesystems with advanced profiling, and scale storage across thousands of nodes
  • We will improve data paths through advanced benchmarking and profiling, while contributing high-impact code to open-source storage projects and internal tooling
Technologies:
  • AI
  • Ansible
  • Architect
  • ArgoCD
  • Cloud
  • Ceph
  • GitOps
  • Grafana
  • Hardware
  • Helm
  • InfiniBand
  • Kubernetes
  • Linux
  • Network
  • Prometheus
  • Python
  • RDMA
  • Terraform
  • NodeJS
  • DevOps

More:

We are Together AI, a research-driven artificial intelligence company focused on open and transparent AI systems that expand innovation and deliver better outcomes for society. Our mission is to significantly reduce the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets, and our team has helped drive advances such as FlashAttention, Hyena, FlexGen, and RedPajama. We offer competitive compensation, startup equity, health insurance, additional benefits, and flexibility for remote work. The US base salary range for this full-time role is $250,000 to $300,000 plus equity and benefits, with compensation determined by location, level, experience, skills, and role-related knowledge. We are an equal opportunity employer and welcome applicants from all backgrounds.

last updated 33 week of 2026

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Staff Engineer, Distributed Storage & AI Infra in San Francisco, CA vacancy
  • $150k - $250k

    Asari AI in San Francisco is looking for a skilled individual to build the supercomputing infrastructure that runs AI agents, tackling...  ...workloads. Your role will involve designing cloud compute, distributed systems, and sandboxed tooling to ensure efficiency and... 
    Suggested

    Asari AI

    San Francisco, CA
    3 days ago
  • A leading technology firm in San Francisco is seeking a candidate to build and scale distributed training systems for large model pre-training. You will collaborate with research teams to design and operate training runs and enhance performance across distributed training... 
    Suggested

    Reflection

    San Francisco, CA
    2 days ago
  •  ...Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be...  ...candidate has a strong background in distributed systems and is eager to engage in complex... 
    Suggested

    Sail Research

    San Francisco, CA
    4 days ago
  •  ...Francisco, CA, is seeking a Member of Technical Staff for distributed systems to design, build, and operate the platform that schedules AI workloads across thousands of nodes. This...  .... You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie... 
    Suggested

    Acceler8 Talent

    San Francisco, CA
    3 days ago
  •  ...in San Francisco is seeking a Member of Technical Staff to design and build distributed systems for AI workloads. The role involves developing scheduling...  ...APIs. Ideal candidates should have strong software engineering skills and experience with distributed systems. This... 
    Suggested

    Gimlet Labs

    San Francisco, CA
    2 days ago
  • Eventual is seeking a Member of Technical Staff to build Eventual's core products and...  ...autonomously solving problems. We value engineers who can scope tasks and implement efficient...  ..., with a focus on performance and reliability across distributed #J-18808-Ljbffr Mixpeek

    Mixpeek

    San Francisco, CA
    4 days ago
  • $150k - $350k

     ...Inc. is seeking a Member of Technical Staff to focus on distributed systems in San Francisco, California....  ...and building the core platform for AI workloads, developing resource management...  ...should have strong software engineering fundamentals and experience with distributed... 

    Gimlet Labs, Inc.

    San Francisco, CA
    2 days ago
  • Acceler8 Talent in San Francisco is seeking a Member of Technical Staff to lead distributed systems at the core of our AI infrastructure. You’ll design, build, and operate scheduling, routing, and coordination for AI workloads across CPUs, GPUs, and accelerators. You’ll... 

    Acceler8 Talent

    San Francisco, CA
    3 days ago
  • $160k - $300k

     ...pioneering foundational AI company for physical product...  ...is to revolutionize how engineering decisions are made,...  ...the Role As a Senior / Staff Infrastructure Engineer at...  ...experience (Python, APIs, distributed systems) Exposure to ML infra Personality & Values:... 
    Work at office
    Visa sponsorship
    Flexible hours

    Apiphany

    San Francisco, CA
    2 days ago
  •  ...Francisco is hiring Members of Technical Staff to build systems that accelerate LLM inference...  ...on high-performance kernels, inference engine internals, and production infrastructure for...  ...while collaborating with a fast-growing AI inference company. #J-18808-Ljbffr Simplify

    Simplify

    San Francisco, CA
    2 days ago
  • A cutting-edge AI research firm in San Francisco is seeking talent to build and optimize GPU infrastructure for large-scale model inference and training workloads. The ideal candidate will have hands-on experience with GPU systems and optimization techniques, actively... 

    Reflection

    San Francisco, CA
    2 days ago
  • Causal Labs is building a Large Physics foundation Model and GPU-driven compute environment to enable rapid research iteration at scale. You will design, deploy, and operate massive GPU clusters, extending Kubernetes and Slurm for efficient, multi-tenant workloads. You ...

    Causal Labs

    San Francisco, CA
    6 days ago
  • David Joseph & Company is seeking a Member of Technical Staff to architect and build systems that simulate, analyze, and evaluate conversational AI agents. You will design scalable cloud infrastructure and own end-to-end features while collaborating with major customers... 

    David Joseph & Company

    San Francisco, CA
    4 days ago
  •  ...Harrison Clarke is working with a high-growth startup in San Francisco seeking a Staff Distributed Systems Engineer. This hands-on role involves designing and building core systems for a cutting-edge AI code generation product. The ideal candidate will have excellent coding... 

    Harrison Clarke

    San Francisco, CA
    4 hours ago
  •  ...are the Data Foundation & AI team within Plaid’s Data organization...  ..., applied AI, and distributed systems, helping establish...  ...innovation across Plaid. As a Staff Machine Learning Engineer, you will lead the...  ...platform capabilities (serving infra, feature stores) used... 
    Full time
    Work experience placement
    Local area
    Immediate start

    Plaid Inc.

    San Francisco, CA
    1 day ago
  • Claryo is seeking a Staff Software Engineer with a focus on Computer Vision Deployment based in San...  ...develop robust infrastructures that power AI-driven warehouse intelligence....  ...Responsibilities include creating and managing distributed cloud GPU infrastructures and building... 
    Work at office
    3 days per week

    Claryo

    San Francisco, CA
    3 days ago
  • $150k - $300k

     ...developer to build open superintelligence stacks. This hybrid role involves designing and implementing distributed orchestration infrastructure while developing tools for AI workload management. Candidates should have strong systems programming experience, particularly... 
    Flexible hours

    Prime Intellect, Inc.

    San Francisco, CA
    5 days ago
  •  ...TrueFoundry is hiring a Staff Engineer – Core Engineering to help build a unified compute layer and governance for AI production systems. You will own system design, drive architectural...  .... The role requires deep expertise in distributed systems, cloud-native architectures, and... 

    TrueFoundry

    San Francisco, CA
    5 hours ago
  •  ...run our fleet orchestration, distribution center automation, telemetry...  ...field-ops teams. Expect hands-on engineering work, prioritized ownership...  ...action tracking. Secure AI/agent-assisted development and...  ...multiple domains (cloud infra, web services, and embedded/autonomy... 
    Local area

    Zipline

    San Francisco, CA
    6 days ago
  • $227.87k

     ...ecosystem, including supply, distribution, and engagement/utility outcomes...  ...teams, Trust & Safety, Ads, Infra) to translate ambiguous...  ...Mentor and grow junior ML engineers through technical coaching, design...  ..., Copilot, Codex, or similar AI coding assistants for development... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    4 hours ago
  • Cohere is a leading security-first enterprise AI company with offices worldwide, including...  .... We’re seeking a Member of Technical Staff to design and deploy scalable training...  ...’ll work with Python, ML frameworks, and distributed training tools. Join a world-class team with... 
    Remote work
    Worldwide

    Cohere

    San Francisco, CA
    5 days ago
  • Fluidstack is seeking a senior network deployment engineer to lead end-to-end fabric turn-ups across data centers. You will own technical execution, build playbooks for ZTP and validation, and tackle escalations that others cannot close. Candidates should have proven experience... 

    FluidStack

    San Francisco, CA
    5 days ago
  • David Joseph & Company is partnering with a seed-stage startup building the simulation, observability, and evaluation layer for Voice AI agents. The role entails architecting systems that simulate and analyze conversational AI, plus owning scalable real-time... 

    David Joseph & Company

    San Francisco, CA
    4 days ago
  • $200k - $245k

     ...an artificial intelligence (AI) powered technology stack purpose...  ...question. It's the core engineering challenge of this role, and there...  ...with skewed and heavy-tailed distributions — not just the Gaussian/...  ...Terraform), and large-scale result storage and query (Elasticsearch or... 
    Temporary work
    Work at office
    Visa sponsorship
    Flexible hours

    Kodiak

    San Francisco, CA
    14 days ago
  • $252k - $315k

    About Scale AIScale AI is the data foundation for AI, helping organizations...  ...problems remains one of the hardest engineering challenges.As a Staff Frontier Agent Engineer (Applied AI)...  ...Software EngineeringExperience building distributed production systems.Experience with... 
    Full time

    Scale AI

    San Francisco, CA
    3 days ago
  • Sieve, a multi-modal lab, seeks an infrastructure engineer to design and run ML data pipelines and ETL systems at scale in San Francisco...  ...workflows. You’ll work with cloud architectures, build fast distributed systems, and create internal tooling and CI/CD for rapid... 

    Sieve

    San Francisco, CA
    4 days ago
  • $141.7k - $250.8k

     ...San Francisco Bay Area, CA As a Sr. Staff Technical Solutions Engineer and tech subject matter expert, you...  ...tools like Spark UI metrics, Mosaic AI Model Service, DAGs, and event logs....  ...designing, building, and troubleshooting distributed computing applications, with 4+... 
    Local area
    Worldwide
    Night shift

    DataBricks

    San Francisco, CA
    2 days ago
  •  ...metrics, alerting, and on-call practices across distributed deployments What we look for Shipped real infra at a startup and lived with the consequences Strong...  ...operational runbooks Bachelors degree in CS, Engineering, or related field (or equivalent practical experience... 
    Visa sponsorship
    Flexible hours

    NeoSigma Telecom Solutions Private Limited

    San Francisco, CA
    5 days ago
  • $200k - $300k

     ...the leading provider of cloud-based AI solutions to understand, search, and...  ...interested in joining the future of AI! Staff Machine Learning Engineer In order to execute our vision,...  ...implemented highly-available distributed systems/microservices You have delivered... 
    Full time

    Hive

    San Francisco, CA
    1 day ago
  •  ...by providing trusted decision-ready AI to the world's most critical organizations...  ...stake real consequences on. As a Staff Machine Learning Engineer, you’ll own AI-driven products end to...  ...reliable at scale, drawing on real distributed-systems experience. You’re energized... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Primer.ai

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Engineer, Distributed Storage & AI Infra. Be the first to apply!