Staff Engineer, Distributed Storage & AI Infra
$250k - $300kTogether AI
Salary: $250,000 - 300,000 per year Requirements:
- We need 8+ years of experience in storage engineering, including operating distributed storage at multi-petabyte scale
- We look for a proven background in deploying and running high-performance storage for GPU or HPC clusters
- We require deep production experience with Kubernetes and cloud-native storage
- We expect strong Go and Python programming skills with the ability to build production-grade systems and tooling
- We require a BS/MS in Computer Science, Engineering, or equivalent practical experience
- We value technical leadership that has improved performance, reliability, or cost efficiency at scale
- We need deep expertise in distributed storage systems such as Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems
- We expect production experience with object storage such as S3, MinIO, Ceph, or R2, including performance tuning and cost control
- We require experience with Kubernetes storage components such as CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers
- We need experience optimizing storage for GPU workloads and working with RDMA/InfiniBand networking and parallel filesystem performance
- We require automation and tooling experience with Go and Python
- We expect infrastructure-as-code experience with Terraform, Ansible, Helm, and GitOps using ArgoCD
- We need advanced knowledge of the Linux storage stack, including ext4, xfs, LVM, NVMe tuning, and RAID configurations
- We expect observability experience with Prometheus, Grafana, and Thanos
- Nice-to-have experience includes GPU Direct Storage, NVMe-oF, storage networking, and RDMA implementations
- Nice-to-have familiarity with ML/AI storage patterns such as model weights, checkpointing, and dataset caching
- Nice-to-have experience with storage benchmarking and profiling tools such as fio, iperf3, iostat, and blktrace
- We will define and drive the technical strategy and storage roadmap for Together AI as we expand our GPU fleet
- We will engineer and scale multi-petabyte AI/ML storage systems by integrating Vast, Weka, and Ceph while reducing cost through automated tiering and lifecycle policies
- We will design intelligent caching and tiered storage architectures to deliver extreme IOPS and cluster-wide throughput for training and inference workloads
- We will tune storage isolation at the L2/L3 network layers to provide secure, production-ready multi-tenancy for storage clients
- We will build Kubernetes storage operators and controllers that enable automated provisioning, self-service workflows, and quota enforcement
- We will optimize end-to-end data paths to achieve 10+ GB/s per GPU node, architect multi-tier caching for model weights and datasets, tune parallel filesystems with advanced profiling, and scale storage across thousands of nodes
- We will improve data paths through advanced benchmarking and profiling, while contributing high-impact code to open-source storage projects and internal tooling
- AI
- Ansible
- Architect
- ArgoCD
- Cloud
- Ceph
- GitOps
- Grafana
- Hardware
- Helm
- InfiniBand
- Kubernetes
- Linux
- Network
- Prometheus
- Python
- RDMA
- Terraform
- NodeJS
- DevOps
More:
We are Together AI, a research-driven artificial intelligence company focused on open and transparent AI systems that expand innovation and deliver better outcomes for society. Our mission is to significantly reduce the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets, and our team has helped drive advances such as FlashAttention, Hyena, FlexGen, and RedPajama. We offer competitive compensation, startup equity, health insurance, additional benefits, and flexibility for remote work. The US base salary range for this full-time role is $250,000 to $300,000 plus equity and benefits, with compensation determined by location, level, experience, skills, and role-related knowledge. We are an equal opportunity employer and welcome applicants from all backgrounds.
last updated 33 week of 2026
$150k - $250k
Asari AI in San Francisco is looking for a skilled individual to build the supercomputing infrastructure that runs AI agents, tackling... ...workloads. Your role will involve designing cloud compute, distributed systems, and sandboxed tooling to ensure efficiency and...Suggested- A leading technology firm in San Francisco is seeking a candidate to build and scale distributed training systems for large model pre-training. You will collaborate with research teams to design and operate training runs and enhance performance across distributed training...Suggested
- ...Research in San Francisco is seeking a talented engineer to design and implement robust systems that ensure fast and cost-efficient AI inference at global scale. You will be... ...candidate has a strong background in distributed systems and is eager to engage in complex...Suggested
- ...Francisco, CA, is seeking a Member of Technical Staff for distributed systems to design, build, and operate the platform that schedules AI workloads across thousands of nodes. This... .... You will collaborate with founders and engineers from Nvidia, Google AI, Intel, and Pixie...Suggested
- ...in San Francisco is seeking a Member of Technical Staff to design and build distributed systems for AI workloads. The role involves developing scheduling... ...APIs. Ideal candidates should have strong software engineering skills and experience with distributed systems. This...Suggested
- Eventual is seeking a Member of Technical Staff to build Eventual's core products and... ...autonomously solving problems. We value engineers who can scope tasks and implement efficient... ..., with a focus on performance and reliability across distributed #J-18808-Ljbffr Mixpeek
$150k - $350k
...Inc. is seeking a Member of Technical Staff to focus on distributed systems in San Francisco, California.... ...and building the core platform for AI workloads, developing resource management... ...should have strong software engineering fundamentals and experience with distributed...- Acceler8 Talent in San Francisco is seeking a Member of Technical Staff to lead distributed systems at the core of our AI infrastructure. You’ll design, build, and operate scheduling, routing, and coordination for AI workloads across CPUs, GPUs, and accelerators. You’ll...
$160k - $300k
...pioneering foundational AI company for physical product... ...is to revolutionize how engineering decisions are made,... ...the Role As a Senior / Staff Infrastructure Engineer at... ...experience (Python, APIs, distributed systems) Exposure to ML infra Personality & Values:...Work at officeVisa sponsorshipFlexible hours- ...Francisco is hiring Members of Technical Staff to build systems that accelerate LLM inference... ...on high-performance kernels, inference engine internals, and production infrastructure for... ...while collaborating with a fast-growing AI inference company. #J-18808-Ljbffr Simplify
- A cutting-edge AI research firm in San Francisco is seeking talent to build and optimize GPU infrastructure for large-scale model inference and training workloads. The ideal candidate will have hands-on experience with GPU systems and optimization techniques, actively...
- Causal Labs is building a Large Physics foundation Model and GPU-driven compute environment to enable rapid research iteration at scale. You will design, deploy, and operate massive GPU clusters, extending Kubernetes and Slurm for efficient, multi-tenant workloads. You ...
- David Joseph & Company is seeking a Member of Technical Staff to architect and build systems that simulate, analyze, and evaluate conversational AI agents. You will design scalable cloud infrastructure and own end-to-end features while collaborating with major customers...
- ...Harrison Clarke is working with a high-growth startup in San Francisco seeking a Staff Distributed Systems Engineer. This hands-on role involves designing and building core systems for a cutting-edge AI code generation product. The ideal candidate will have excellent coding...
- ...are the Data Foundation & AI team within Plaid’s Data organization... ..., applied AI, and distributed systems, helping establish... ...innovation across Plaid. As a Staff Machine Learning Engineer, you will lead the... ...platform capabilities (serving infra, feature stores) used...Full timeWork experience placementLocal areaImmediate start
- Claryo is seeking a Staff Software Engineer with a focus on Computer Vision Deployment based in San... ...develop robust infrastructures that power AI-driven warehouse intelligence.... ...Responsibilities include creating and managing distributed cloud GPU infrastructures and building...Work at office3 days per week
$150k - $300k
...developer to build open superintelligence stacks. This hybrid role involves designing and implementing distributed orchestration infrastructure while developing tools for AI workload management. Candidates should have strong systems programming experience, particularly...Flexible hours- ...TrueFoundry is hiring a Staff Engineer – Core Engineering to help build a unified compute layer and governance for AI production systems. You will own system design, drive architectural... .... The role requires deep expertise in distributed systems, cloud-native architectures, and...
- ...run our fleet orchestration, distribution center automation, telemetry... ...field-ops teams. Expect hands-on engineering work, prioritized ownership... ...action tracking. Secure AI/agent-assisted development and... ...multiple domains (cloud infra, web services, and embedded/autonomy...Local area
$227.87k
...ecosystem, including supply, distribution, and engagement/utility outcomes... ...teams, Trust & Safety, Ads, Infra) to translate ambiguous... ...Mentor and grow junior ML engineers through technical coaching, design... ..., Copilot, Codex, or similar AI coding assistants for development...Work at officeLocal areaRelocationRelocation package- Cohere is a leading security-first enterprise AI company with offices worldwide, including... .... We’re seeking a Member of Technical Staff to design and deploy scalable training... ...’ll work with Python, ML frameworks, and distributed training tools. Join a world-class team with...Remote workWorldwide
- Fluidstack is seeking a senior network deployment engineer to lead end-to-end fabric turn-ups across data centers. You will own technical execution, build playbooks for ZTP and validation, and tackle escalations that others cannot close. Candidates should have proven experience...
- David Joseph & Company is partnering with a seed-stage startup building the simulation, observability, and evaluation layer for Voice AI agents. The role entails architecting systems that simulate and analyze conversational AI, plus owning scalable real-time...
$200k - $245k
...an artificial intelligence (AI) powered technology stack purpose... ...question. It's the core engineering challenge of this role, and there... ...with skewed and heavy-tailed distributions — not just the Gaussian/... ...Terraform), and large-scale result storage and query (Elasticsearch or...Temporary workWork at officeVisa sponsorshipFlexible hours$252k - $315k
About Scale AIScale AI is the data foundation for AI, helping organizations... ...problems remains one of the hardest engineering challenges.As a Staff Frontier Agent Engineer (Applied AI)... ...Software EngineeringExperience building distributed production systems.Experience with...Full time- Sieve, a multi-modal lab, seeks an infrastructure engineer to design and run ML data pipelines and ETL systems at scale in San Francisco... ...workflows. You’ll work with cloud architectures, build fast distributed systems, and create internal tooling and CI/CD for rapid...
$141.7k - $250.8k
...San Francisco Bay Area, CA As a Sr. Staff Technical Solutions Engineer and tech subject matter expert, you... ...tools like Spark UI metrics, Mosaic AI Model Service, DAGs, and event logs.... ...designing, building, and troubleshooting distributed computing applications, with 4+...Local areaWorldwideNight shift- ...metrics, alerting, and on-call practices across distributed deployments What we look for Shipped real infra at a startup and lived with the consequences Strong... ...operational runbooks Bachelors degree in CS, Engineering, or related field (or equivalent practical experience...Visa sponsorshipFlexible hours
$200k - $300k
...the leading provider of cloud-based AI solutions to understand, search, and... ...interested in joining the future of AI! Staff Machine Learning Engineer In order to execute our vision,... ...implemented highly-available distributed systems/microservices You have delivered...Full time- ...by providing trusted decision-ready AI to the world's most critical organizations... ...stake real consequences on. As a Staff Machine Learning Engineer, you’ll own AI-driven products end to... ...reliable at scale, drawing on real distributed-systems experience. You’re energized...Full timeContract workRemote workFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Engineer, Distributed Storage & AI Infra. Be the first to apply!
- assistant engineering manager San Francisco, CA
- assistant civil engineer San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineer San Francisco, CA
- staff engineer San Francisco, CA
- staff data engineer San Francisco, CA
- software engineer staff San Francisco, CA
- assistant electrical engineer San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA




