Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer - Infrastructure Storage

Full-time

Lambda

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU. If you'd like to build the world's best AI cloud, join us. *Note: This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday. In the world of distributed AI, raw GPU and CPU horsepower is just a part of the story. High-performance networking and storage are the critical components that enable and unite these systems, making groundbreaking AI training and inference possible. The Lambda Infrastructure Engineering organization forges the foundation of high-performance AI clusters by welding together the latest in AI storage, networking, GPU and CPU hardware. Our expertise lies at the intersection of: High-Performance Distributed Storage Solutions and Protocols: We engineer the protocols and systems that serve massive datasets at the speeds demanded by modern clustered GPUs. Dynamic Networking: We design advanced networks that provide multi-tenant security and intelligent routing without compromising performance, using the latest in AI networking hardware. Compute Virtualization: We enable cutting-edge virtualization and clustering that allows AI researchers and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs. About the Role: We are seeking a seasoned Staff Storage Software Engineer with deep experience designing and deploying storage protocol solutions at scale across object, block, and file paradigms. This is a unique opportunity to work at the intersection of large-scale distributed systems and the rapidly evolving field of artificial intelligence infrastructure. This is an opportunity to have a significant impact on the future of AI. You will be building the foundational infrastructure that powers some of the most advanced AI research and products in the world. What You’ll Do Technical Leadership: Set technical direction for storage software architecture across the Infrastructure Engineering organization, influencing decisions that span petabyte-scale deployments. Author and review design documents for new storage systems, protocols, and integrations; raise the technical bar across the team. Mentor and develop senior engineers, providing guidance on systems design, debugging complex distributed systems issues, and navigating technical tradeoffs. Serve as a technical anchor for cross-functional initiatives involving storage, networking, compute, and control plane teams. Represent the storage software team in architectural reviews, roadmap planning, and customer-facing technical discussions where needed. Execution: Design, develop, and maintain high-performance storage systems software with a focus on performance, scalability, reliability, and operational simplicity. Implement and optimize storage protocol APIs across file (NFS, SMB, Lustre), block (NVMe-oF, iSCSI, Fibre Channel), and object (S3) access patterns. Develop distributed systems for managing and orchestrating storage resources across multiple solutions and redundant arrays. Collaborate with hardware and system architects to integrate software with storage solutions including NVMe, GPU-direct storage, and DPU-accelerated data paths. Troubleshoot and resolve complex issues in production data center environments, including performance regressions, protocol mismatches, and hardware failures. Contribute across the full software development lifecycle — from requirements gathering and system design through deployment, monitoring, and long-term maintenance. Build and maintain tooling for storage benchmarking, performance profiling, and capacity planning. Collaboration Work closely with storage software and networking teams to execute cross-functional infrastructure initiatives and new data center deployments, including integration of storage protocols across a variety of on-prem solutions. Partner with the control plane and Kubernetes teams to meet customer and product requirements for usability, reliability, and telemetry. Work with the observability team to define, build, and track SLOs/SLIs for storage systems. Coordinate with Networking, Compute, and Storage Engineering teams to deploy high-performance distributed storage solutions that serve AI/ML workloads. Partner with the Fleet Engineering team to ensure seamless deployment, monitoring, and ongoing maintenance of distributed storage infrastructure. Innovate: Stay current with the latest research and developments in AI and HPC storage technologies, and bring relevant advances into Lambda's infrastructure. Work with the Lambda product team to identify emerging trends in AI inference and training that will shape next-generation storage requirements. Evaluate and prototype new storage solutions, protocols, and hardware integrations - from open-source distributed filesystems to vendor-specific accelerated storage products. Optimize storage protocol solutions for AI workloads, including checkpoint I/O for training, high-throughput dataset serving, and latency-sensitive inference pipelines. You Have: Experience 10+ years of experience in storage systems engineering, with at least 5 years in a technical lead or Staff+ IC role. Proven track record designing and operating storage infrastructure at scale (multi-petabyte environments preferred) in production data center or cloud settings. Experience leading technical projects end-to-end, from architecture through delivery with cross-functional stakeholders. Background working in high-performance computing, AI/ML infrastructure, or large-scale cloud storage environments. Systems-Level Programming Strong proficiency in one or more low-level systems programming languages: C, C++, Rust, or Go. Demonstrated ability to write high-performance, concurrent, production-grade systems code and conduct thorough code reviews. Experience with kernel-level storage drivers, user-space I/O frameworks, or storage daemon development is a strong plus. Familiarity with DPDK and SPDK and their role in building high-performance, kernel-bypass storage and networking data paths. Storage Protocol & API Expertise Deep hands-on experience with two or more storage protocols: object (S3 or similar), block (iSCSI, Fibre Channel, NVMe-oF), or file (NFS, SMB, Lustre, DAOS). Experience implementing or maintaining storage protocol servers or clients in production, not just consuming them. Familiarity with storage API performance characteristics such as latency, throughput, IOPS and the ability to diagnose and resolve bottlenecks at the protocol level. Storage Performance Optimization Experience profiling and tuning storage systems for throughput, latency, and IOPS under real production workloads. Familiarity with tools such as fio, blktrace, perf, eBPF/bpftrace, or equivalent for storage performance analysis. Understanding of I/O scheduling, caching layers, write amplification, and related performance tradeoffs. Modern Storage Technologies Familiarity with NVMe, NVMe-oF, and RDMA (RoCE or InfiniBand) and their impact on storage system architecture. Working knowledge of DPUs (e.g., NVIDIA BlueField) and their role in offloading storage and networking data paths. Experience with GPU-direct storage or similar zero-copy data paths is a plus. Physical Infrastructure & Operational Acumen Comfort working in a physical data center environment — understanding rack-scale infrastructure, storage array hardware, cabling, and failure domains. Experience building and operating storage systems with strong reliability expectations: designing for failure, building runbooks, and driving incident response. Familiarity with storage observability tooling — metrics pipelines (Prometheus, Grafana), log aggregation, and tracing in distributed storage environments. Nice to Have Experience with NVIDIA BlueField DPUs or SuperNICs for accelerated storage data paths, including GPUDirect Storage implementation. Deep production experience with enterprise or HPC storage platforms: Vast Data, Weka, NetApp, or IBM Spectrum Scale. Experience deploying and operating Ceph at scale (100PB+) in an HPC or AI infrastructure environment. Familiarity with emerging storage technologies such as CXL memory pooling, computational storage, or ZNS (Zoned Namespace) SSDs. Experience contributing to or maintaining open-source storage projects (e.g., Ceph, DAOS, Lustre, MinIO). Salary Range Information The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description. About Lambda Founded in 2012, with 500+ employees, and growing fast Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG Our values are publicly available: We offer generous cash & equity compensation Health, dental, and vision coverage for you and your dependents Wellness and commuter stipends for select roles 401k Plan with 2% company match (USA employees) Flexible paid time off plan that we all actually use A Final Note: You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills. Equal Opportunity Employer Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer - Infrastructure Storage in San Francisco, CA vacancy
  • $240k - $310k

     ...intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and...  ...us at Crusoe. About This Role The Cloud Storage team at Crusoe is seeking a Staff Software Engineer to serve as a primary architect and visionary... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    20 days ago
  • $155k - $250k

     ...future of global computing infrastructure. As data centers consume...  ...We are seeking a seasoned software architect/engineer with a deep passion for building...  ..., high performance storage systems with significant experience...  ...the Role: As a Senior/Staff Software Engineer on the... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    more than 2 months ago
  • $245k - $290k

     ...Crusoe is building the World’s Favorite AI-first Cloud infrastructure company. We’re pioneering vertically integrated,...  ...infrastructure. About This Role: As a Senior Staff Software Engineer on the Cloud Storage team, you will lead the development and execution of... 
    Suggested
    Temporary work

    Crusoe

    San Francisco, CA
    more than 2 months ago
  • $246k - $370k

     ...invest heavily in building the best creator platform with the best team in the creator economy and are looking for a Staff Storage Platform Software Engineer to support our mission. This role is based in San Francisco and open to those who are able to be in-office 2... 
    Suggested
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    2 days per week

    Patreon

    San Francisco, CA
    more than 2 months ago
  • $180k - $200k

     ...follow us on LinkedIn. AI Engineering @ Ironclad Ironclad is...  ...reliable, highly scalable software and services designed for a...  ...Deliver and Optimize AI Infrastructure: Work with platform teams to...  ...0,000 Base Salary Range - Staff: $210,000 - $235,000 The base... 
    Suggested
    Contract work

    Ironclad

    San Francisco, CA
    4 days ago
  • $177.19k - $364.8k

     ...use AI in our recruiting process here. The Trends & Insights Engineering team builds the products that turn Pinterest's unique signal...  ...for into actionable intelligence for advertisers. As a Staff Software Engineer on this team, you'll lead the technical evolution of... 
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Relocation package
    Day shift

    Pinterest

    San Francisco, CA
    18 hours ago
  • $230k - $285k

     ...Connor was a machine learning research engineer at Scale AI . The rest of our team comes...  ...Who you are: You have 8+ years of software engineering experience and have a...  ...have experience standing up and managing infrastructure. You are comfortable and excited to mentor... 

    Unify

    San Francisco, CA
    4 days ago
  •  ...Staff Software Engineer, Listings & Host Tools and AI Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across... 
    Work experience placement

    airbnb, Inc.

    San Francisco, CA
    4 days ago
  • $245k - $295k

     ...As the only vertically integrated AI infrastructure company built from the ground up, we...  ...About the Role: We are seeking a Sr Staff Software Engineer to anchor distributed systems depth across...  ...decisions across Control Plane, Storage & State, Edge & Agents, Data Pipeline... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  •  ...improve defense readiness through software products that streamline and...  ...-office in one of our core engineering locations: El Segundo, CA;...  ..., CA About the Role The Infrastructure team owns the platform that...  ...cloud — operators, networking, storage, and multi-cluster... 
    Full time
    Work at office
    Local area

    Gallatin

    San Francisco, CA
    2 days ago
  • $200k - $230k

     ...looking for an experienced engineer with deep expertise in distributed...  ...shape the future of Gusto's storage layer. You'll manage complex...  ...the Team: The Datastores Infrastructure Engineering team designs,...  ...for: ~12+ years of software engineering experience building... 
    Work at office
    Local area
    Remote work
    2 days per week
    3 days per week

    gusto

    San Francisco, CA
    more than 2 months ago
  • $188k - $275k

     ...CoreWeave combines superior infrastructure performance with deep...  ...tackling challenging engineering problems, and are...  ...About the Role: As a Staff Engineer on Marimo's molab...  ...with CoreWeave object storage, and will solve for...  ...years of experience in software engineering ~ Strong... 
    Permanent employment
    Temporary work
    Casual work
    Work at office
    Worldwide
    Flexible hours

    CoreWeave

    San Francisco, CA
    27 days ago
  • $320k

     ...group of committed researchers, engineers, policy experts, and...  ...technical direction for caching infrastructure used across Product and Research...  ...Significant experience as a software engineer building and...  ...policy: Currently, we expect all staff to be in one of our offices... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    3 days ago
  • $224k - $264k

     ...customer satisfaction, across a complex, regulated domain. As a Staff Engineer in this organization, you'll operate across several product...  ...field, or equivalent practical experience 10+ years of software development experience and 3+ years experience in a project leadership... 
    Full time
    Work at office
    Local area
    Remote work
    Relocation
    Flexible hours
    3 days per week

    Checkr

    San Francisco, CA
    3 days ago
  • $185k - $224k

     ...energy and intelligence. We’re crafting the engine that powers a world where people can...  ...for responsible, transformative cloud infrastructure. About This Role: Crusoe Cloud seeks a highly skilled and experienced Staff Software Engineer to lead the development and... 
    Temporary work

    Crusoe

    San Francisco, CA
    more than 2 months ago
  • $202k - $269k

     ...and convert high-quality pipeline to revenue. Role Summary We are looking for a highly skilled and experienced Staff Software Engineer, Infrastructure to elevate our infrastructure. You will be a key driver in shaping the future of our data and compute platforms,... 
    Remote job
    Full time
    Work experience placement

    6sense

    San Francisco, CA
    more than 2 months ago
  • $185k - $224k

     ...and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each...  ...: Crusoe Cloud seeks a highly skilled and experienced Staff Software Engineer to lead the development and execution of our cutting-edge... 
    Temporary work

    Crusoe

    San Francisco, CA
    6 days ago
  • $300 per month

     ...and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each...  ...Systems is seeking a highly skilled and motivated Senior Staff Software Engineer - Software Defined Networking to lead the development and... 
    Temporary work

    Crusoe

    San Francisco, CA
    19 days ago
  • $204k - $247k

     ...s Favorite AI-first Cloud infrastructure company. We’re pioneering...  ...Role: The Crusoe Cloud Software Development team is seeking...  ...passionate and experienced Senior Staff Software Engineer specializing in Hypervisor...  ...accelerating AI compute, storage, and networking resources.... 
    Full time
    Temporary work

    Crusoe

    San Francisco, CA
    more than 2 months ago
  • $192k - $215k

     ...chapter: Unwinding years of Salesforce-bound logic into modern, engineering-owned backend services. Integrating best-in-class vendors...  ..., observable, and reliable across products. If you’re a Staff-level engineer who thrives on modernizing fragile systems, building... 
    Full time
    Work at office
    Local area

    Ripple

    San Francisco, CA
    more than 2 months ago
  • $159k - $268k

     ...to autonomous vehicle development. Our closed-loop simulation engine built with the latest in generative AI technologies, Waabi World...  ...extremely large scale.  - Design and implement orchestration software between simulation subcomponents including the autonomy system,... 
    Full time
    Work at office
    Work from home
    Flexible hours

    Waabi

    San Francisco, CA
    26 days ago
  • $131k - $285k

     ...verticals, including Retail, Discovery (Feed), and Ads, with several engineers building new product experiences on top of the platform. This...  .... Collaborate across the Company - Work closely with the Infrastructure, Product, and AI teams across the company to tailor the... 
    Hourly pay
    Work at office
    Local area
    Flexible hours

    DoorDash

    San Francisco, CA
    more than 2 months ago
  • $160k - $250k

     ...automation startup serving US manufacturers. As a Senior/Staff Platform Engineer, you will build the software layer that powers autonomous manufacturing...  ...Develop and productionize LLM-based agents, AI/ML infrastructure, and RAG patterns for intelligent quote analysis and... 
    Visa sponsorship

    Clera

    San Francisco, CA
    26 days ago
  • $163k - $204k

     ...interview process. We’re looking for seasoned full-stack software engineers to join Gusto’s Growth team building the systems and workflows...  ..., relevant engagement experiences across the funnel. As a Staff Software Engineer, you’ll operate across the full product lifecycle... 
    Full time
    Work at office
    Local area
    Remote work
    2 days per week
    3 days per week

    gusto

    San Francisco, CA
    11 days ago
  • $163k - $247k

     ...more about our Total Rewards philosophy .  About the Role: We’re hiring a Staff Software Engineer to join Gusto’s Payroll Platform team, where you’ll build backend product infrastructure that powers payroll experiences across Gusto. This is a fully backend role... 
    For contractors
    Work at office
    Local area
    Remote work
    2 days per week
    3 days per week

    Gusto

    San Francisco, CA
    more than 2 months ago
  • $163k - $204k

     ...Role: Gusto's Platform Orchestration team is rebuilding how engineers ship at Gusto — turning operational toil into a product-grade...  ...agents can drive end-to-end. We sit at the intersection of Infrastructure Engineering, Developer Productivity, Security, and Product Engineering... 
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Shift work
    2 days per week
    3 days per week

    gusto

    San Francisco, CA
    a month ago
  • $237.6k - $288k

     ...and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each...  ...Crusoe. About the Role: We are seeking a Senior Staff Software Engineer for our Billing team to design, build, and scale Crusoe Cloud... 
    Temporary work
    Flexible hours

    Crusoe

    San Francisco, CA
    18 days ago
  • $230k - $270k

     ...LGBTQIA+ Advocacy Award (2022) About the Role As a staff level software engineer at Maven Clinic, you will be responsible for driving the...  ...ensure the agility, flexibility, and scalability of our auth infrastructure including SSO, MFA, and federated identity integrations.... 
    Full time
    Contract work
    Work at office
    Immediate start
    Remote work
    Flexible hours
    3 days per week

    Maven Clinic

    San Francisco, CA
    8 days ago
  • $220k - $240k

     ...are on a mission to revolutionize employment by building the infrastructure that powers every facet of work. To do this, we're looking for...  ...and opportunity cost. You get energy from unblocking other engineers and improving their day-to-day experience. You're equally comfortable... 
    Full time
    Work at office
    Immediate start
    2 days per week

    Finch

    San Francisco, CA
    22 days ago
  • $255k - $405k

     ...About the Team The Agent Infrastructure team at OpenAI is responsible for building systems...  ...execute code, debug issues, and develop software just as human SWEs do. Our training...  ...world. About the Role As a Software Engineer on the Agent Infrastructure team, you will... 
    Work at office
    Worldwide
    Relocation package

    OpenAI

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer - Infrastructure Storage. Be the first to apply!