Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer - Storage

$267k - $356k
Full-time

Lambda

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


Lambda's Storage Engineering team is the backbone behind our world-class storage offerings, operating at a scale seldom seen in today's markets. We own the full spectrum of Lambda's data platform services—from low-level storage systems to the APIs and tooling our customers build on every day.

Our work directly powers some of the most demanding compute workloads in the industry, which means reliability and performance aren't just goals—they're the baseline. We're looking for engineers who want to solve problems at scale, own critical systems end to end, and help shape the future of infrastructure for AI and beyond

What You’ll Do

  • Own the reliability, performance, and capacity health of Lambda's production storage fleet across all data centers, operating behind Lambda's own software-defined data plane.

  • Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.

  • Investigate and resolve storage-related incidents using deep telemetry, logs, and performance profiling — from a single flapping NIC to a cluster-wide rebuild.

  • Automate ticketing, escalation, and incident-response workflows so the team spends less time on repetitive triage and more time on root cause.

  • Design and maintain self-healing automation for common failure modes: drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.

  • Implement CI/CD pipelines for storage automation and tooling.

  • Partner with Storage Engineers, Fleet Orchestration, and Release Engineering to automate the deployment and configuration of software-defined storage across new and existing sites using tools such as Ansible, Jenkins etc.

  • Work with hardware and networking teams to diagnose low-level I/O and network issues — NIC errors, RDMA/RoCE/InfiniBand fabric health, path multipathing — that surface as storage-layer symptoms.

  • Participate in an on-call rotation supporting Lambda's storage fleet, with a focus on driving down MTTR and building the automation that keeps you from getting paged for the same thing twice.

You Have

  • 5+ years of experience operating Linux systems in production or HPC environments, with hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS, or similar).

  • Hands-on experience operating Software-Defined Storage (SDS) platforms at scale, including integrating with their management and data-plane APIs.

  • Strong incident-response instincts: comfortable owning a production storage incident end to end, from first alert through root cause to postmortem.

  • Working experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic — including building dashboards and alert/pager routing for multiple audiences.

  • Working experience with Kubernetes (GitOps tooling such as ArgoCD, Helm/Kustomize) and hands-on troubleshooting.

  • Working experience with CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming in Python or Go.

  • Working experience with Infrastructure as Code (Terraform, Ansible).

  • Solid understanding of core storage protocols across file (NFS, SMB), object (S3), block (NVMe-oF/TCP), and structured (vector DB, SQL) storage.

Nice to Have

  • Experience with cutting-edge software-defined storage solutions such as VAST or Weka.

  • Enterprise storage expertise in solutions such as NetApp, Dell PowerScale, GPFS, or Lustre.

  • Experience writing or operating Kubernetes CSI drivers.

  • Experience with SR-IOV and virtualization (KVM/QEMU).

  • Experience with GPUDirect Storage, RDMA, InfiniBand, or RoCE networking.

  • Familiarity with NIC-level diagnostics (ethtool, mlxlink) and fleet-wide operations tooling (clush or similar).

  • Contributions to open-source storage projects.

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available:

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer - Storage in San Francisco, CA vacancy
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    4 days ago
  • $152.5k - $205k

     ...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    3 days ago
  •  ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of...  ...ownership, working together to build scalable, reliable, and secure products that empower...  ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    4 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $148.5k - $223.9k

     ...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    3 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    1 day ago
  • $117k - $209.33k

     ...3Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,...  ..., or similar technologiesExperience operating databases, storage platforms, messaging systems, caching technologiesExperience... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    22 hours ago
  • $167.7k - $245.2k

     ...approximately 2 days per week on-site at Cisco offices in either...  ...as intended, improving reliability and reducing risks. This...  ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate...  ...spanning Kubernetes, networking, storage, and application layers•... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $165k - $241.4k

     ...very effective.We’re looking for talented engineers with a software or operations background...  ...development teams to ensure the reliability, performance and security of our infrastructure...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    3 days ago
  • $220k - $235k

     ...strategic, high-output Staff/Senior Staff SRE to define the...  ...cloud platform and champion engineering excellence across Ironclad....  ...strategic direction for the Site Reliability Engineering team and our broader...  ...backup/recovery systems, cloud storage architecture, and... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    3 days ago
  •  ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services...  ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role... 
    Senior

    Socket

    San Francisco, CA
    1 day ago
  •  ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves... 
    Senior

    Breakout Tools

    San Francisco, CA
    4 days ago
  •  ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This...  ...intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).... 
    Senior

    Alembic

    San Francisco, CA
    1 day ago
  • $127k - $249k

    We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates... 
    Senior
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    4 days ago
  • $215k - $275k

     ...by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing... 
    Senior
    Work at office

    Anyscale

    San Francisco, CA
    1 day ago
  • $165k - $241.4k

     ...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    22 hours ago
  •  ...Partner with software developers, platform engineers, and IT staff to improve system design,...  ...requirements, service quality, reliability, security, and compliance needs. Drive continuous...  ...Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform... 
    Senior
    Work at office
    Remote work

    GrabJobs

    San Francisco, CA
    2 days ago
  • $175k - $250k

     ...,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco...  .... Modality: On-Site only. Must live within...  ...scalability, performance, and reliability across environments. What...  ...observability, distributed storage, and networking Ensure... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    San Francisco, CA
    14 hours ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient... 
    Senior

    TechChain Talent

    San Francisco, CA
    4 days ago
  • $240k - $310k

     ...high-performing team that believes in each other, come build with us at Crusoe.About This RoleThe Cloud Storage team at Crusoe is seeking a Senior Staff Software Engineer to serve as a primary architect and visionary for our storage strategy. While a Staff Engineer leads... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $210.8k - $272.8k

    About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user... 
    Senior
    Local area

    Thumbtack

    San Francisco, CA
    4 days ago
  • $250k

     ...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves... 
    Senior
    Full time
    Remote work
    San Francisco, CA
    more than 2 months ago
  •  ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems...  ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes... 
    Senior

    Saviynt

    San Francisco, CA
    a month ago
  • $300 per month

     ...industry's most high value workloads. As a Senior Engineering Manager, you will lead the team...  ...building the next generation, bespoke cloud storage platform optimized for AI and HPC...  ..., and CCX teams to ensure storage is a reliable, well-integrated part of the overall platformEngage... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    4 days ago
  • $15k

     ...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster...  ..., Loki, ELK, OpenTelemetry)Experience with distributed storage technologies (Lustre, Ceph, S3)Embodies a "system... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    2 days ago
  • $139.76k - $287.75k

     ...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,...  ...will be instrumental in advancing the reliability, scalability, automation, observability...  ...The ideal candidate is a highly hands-on engineer with strong production experience and a... 
    Senior
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    San Francisco, CA
    2 days ago
  • $153k - $248.5k

    A leading battery materials company in San Francisco seeks a Power Electronics Controls Engineer to design and launch innovative power converters for energy storage. Ideal candidates will have 7+ years in electrical engineering with expertise in high-voltage design and... 
    Senior

    Redwood Materials

    San Francisco, CA
    1 day ago
  • $300k

     ...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will... 
    Senior

    Kody

    San Francisco, CA
    21 days ago
  •  ...Creek Renewables in San Francisco or NYC is seeking a Senior Analyst, Structured Finance, to support financing late‑stage solar and battery storage projects. You will gain exposure to project development, engineering, and EPC workstreams while collaborating with senior... 
    Senior

    Cypress Creek Renewables

    San Francisco, CA
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Storage. Be the first to apply!