Senior Site Reliability Engineer - Storage
$267k - $356kLambda
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
Lambda's Storage Engineering team is the backbone behind our world-class storage offerings, operating at a scale seldom seen in today's markets. We own the full spectrum of Lambda's data platform services—from low-level storage systems to the APIs and tooling our customers build on every day.
Our work directly powers some of the most demanding compute workloads in the industry, which means reliability and performance aren't just goals—they're the baseline. We're looking for engineers who want to solve problems at scale, own critical systems end to end, and help shape the future of infrastructure for AI and beyond
What You’ll Do
Own the reliability, performance, and capacity health of Lambda's production storage fleet across all data centers, operating behind Lambda's own software-defined data plane.
Build and maintain monitoring, dashboards, and alerting for storage performance, capacity, and hardware failures.
Investigate and resolve storage-related incidents using deep telemetry, logs, and performance profiling — from a single flapping NIC to a cluster-wide rebuild.
Automate ticketing, escalation, and incident-response workflows so the team spends less time on repetitive triage and more time on root cause.
Design and maintain self-healing automation for common failure modes: drive replacement, node swaps, rebuild monitoring, and capacity rebalancing.
Implement CI/CD pipelines for storage automation and tooling.
Partner with Storage Engineers, Fleet Orchestration, and Release Engineering to automate the deployment and configuration of software-defined storage across new and existing sites using tools such as Ansible, Jenkins etc.
Work with hardware and networking teams to diagnose low-level I/O and network issues — NIC errors, RDMA/RoCE/InfiniBand fabric health, path multipathing — that surface as storage-layer symptoms.
Participate in an on-call rotation supporting Lambda's storage fleet, with a focus on driving down MTTR and building the automation that keeps you from getting paged for the same thing twice.
You Have
5+ years of experience operating Linux systems in production or HPC environments, with hands-on storage experience at scale on scale-out or software-defined platforms (e.g., CEPH, Lustre, GPFS, or similar).
Hands-on experience operating Software-Defined Storage (SDS) platforms at scale, including integrating with their management and data-plane APIs.
Strong incident-response instincts: comfortable owning a production storage incident end to end, from first alert through root cause to postmortem.
Working experience with monitoring and logging platforms such as Prometheus, Grafana, Alertmanager, Datadog, or SumoLogic — including building dashboards and alert/pager routing for multiple audiences.
Working experience with Kubernetes (GitOps tooling such as ArgoCD, Helm/Kustomize) and hands-on troubleshooting.
Working experience with CI/CD tooling (GitHub Actions, Jenkins, BuildKite), containerization (Docker/Podman), and systems programming in Python or Go.
Working experience with Infrastructure as Code (Terraform, Ansible).
Solid understanding of core storage protocols across file (NFS, SMB), object (S3), block (NVMe-oF/TCP), and structured (vector DB, SQL) storage.
Nice to Have
Experience with cutting-edge software-defined storage solutions such as VAST or Weka.
Enterprise storage expertise in solutions such as NetApp, Dell PowerScale, GPFS, or Lustre.
Experience writing or operating Kubernetes CSI drivers.
Experience with SR-IOV and virtualization (KVM/QEMU).
Experience with GPUDirect Storage, RDMA, InfiniBand, or RoCE networking.
Familiarity with NIC-level diagnostics (ethtool, mlxlink) and fleet-wide operations tooling (clush or similar).
Contributions to open-source storage projects.
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
Founded in 2012, with 500+ employees, and growing fast
Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
Our values are publicly available:
We offer generous cash & equity compensation
Health, dental, and vision coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
- About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure...Senior
$152.5k - $205k
...flexible work environment where new ideas are encouraged and everyone is a stakeholder.What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design, build, and operate the secure, scalable platform infrastructure behind...SeniorFlexible hours- ...’s build what’s next.About the teamThe Engineering team at Airwallex is a diverse group of... ...ownership, working together to build scalable, reliable, and secure products that empower... ...our Global services.What you’ll doAs a Senior Site Reliability Engineer, you’ll work...SeniorTemporary workLocal areaWorldwide
$165k - $225.6k
...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,...SeniorPermanent employmentLocal areaWorldwideFlexible hours$148.5k - $223.9k
...right place! Agentforce is the future of AI, and you are the future of Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco. Working closely with counterparts in the Infrastructure and R&D organizations...SeniorFull timeWorldwideWeekend work$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$117k - $209.33k
...3Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ..., or similar technologiesExperience operating databases, storage platforms, messaging systems, caching technologiesExperience...SeniorFull timeFor contractors$167.7k - $245.2k
...approximately 2 days per week on-site at Cisco offices in either... ...as intended, improving reliability and reducing risks. This... ...and control.As a Senior Site Reliability Engineer (SRE), you will build, operate... ...spanning Kubernetes, networking, storage, and application layers•...SeniorFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $241.4k
...very effective.We’re looking for talented engineers with a software or operations background... ...development teams to ensure the reliability, performance and security of our infrastructure... ...insurance. Please see the Cisco careers site to discover more benefits and perks....SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$220k - $235k
...strategic, high-output Staff/Senior Staff SRE to define the... ...cloud platform and champion engineering excellence across Ironclad.... ...strategic direction for the Site Reliability Engineering team and our broader... ...backup/recovery systems, cloud storage architecture, and...SeniorFull timeContract workWork at office- ...Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services... ...will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The role...Senior
- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have over 5 years of relevant experience and strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves...Senior
- ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This... ...intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI)....Senior
$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$215k - $275k
...by Andreessen Horowitz, NEA, and Addition with $250+ million raised to date.About the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale aims to provide the next generation of tools and infrastructure to make developing...SeniorWork at office$165k - $241.4k
...within Cisco’s Networking, Security, Collaboration, and Observability portfolios.Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...Partner with software developers, platform engineers, and IT staff to improve system design,... ...requirements, service quality, reliability, security, and compliance needs. Drive continuous... ...Required: 8+ years of experience in Site Reliability Engineering, DevOps, Platform...SeniorWork at officeRemote work
$175k - $250k
...,000.00/yr Job Title: Senior Cloud Infrastructure Engineer Location: San Francisco... .... Modality: On-Site only. Must live within... ...scalability, performance, and reliability across environments. What... ...observability, distributed storage, and networking Ensure...SeniorFull timeRemote workRelocationRelocation package- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public blockchain that gives developers the tools to create experiences that are more like cash than crypto. The network is faster, cheaper, and far more energy-efficient...Senior
$240k - $310k
...high-performing team that believes in each other, come build with us at Crusoe.About This RoleThe Cloud Storage team at Crusoe is seeking a Senior Staff Software Engineer to serve as a primary architect and visionary for our storage strategy. While a Staff Engineer leads...SeniorTemporary work$210.8k - $272.8k
About Thumbtack Thumbtack helps millions of people confidently care for their homes. About the Site Reliability Engineering Team The Site Reliability Engineering team focuses on creating and maintaining a reliable, secure, and scalable platform vital for a seamless user...SeniorLocal area$250k
...across Europe, while now significantly expanding its footprint in the United States. The company is looking for a Senior / Staff Site Reliability Engineer to support and scale large-scale HPC and cloud environments powering GPU-intensive workloads. The role involves...SeniorFull timeRemote work- ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical role in ensuring these systems... ...hands-on engineering and technical leadership role. You will own reliability for major platform domains, design scalable solutions on Kubernetes...Senior
$300 per month
...industry's most high value workloads. As a Senior Engineering Manager, you will lead the team... ...building the next generation, bespoke cloud storage platform optimized for AI and HPC... ..., and CCX teams to ensure storage is a reliable, well-integrated part of the overall platformEngage...SeniorTemporary work$15k
...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster... ..., Loki, ELK, OpenTelemetry)Experience with distributed storage technologies (Lustre, Ceph, S3)Embodies a "system...SeniorWork at officeLocal areaRemote work$139.76k - $287.75k
...to grow their business.We are seeking a Senior Site ReliabilityEngineer to help operate,... ...will be instrumental in advancing the reliability, scalability, automation, observability... ...The ideal candidate is a highly hands-on engineer with strong production experience and a...SeniorWork at officeLocal areaRelocationRelocation package$153k - $248.5k
A leading battery materials company in San Francisco seeks a Power Electronics Controls Engineer to design and launch innovative power converters for energy storage. Ideal candidates will have 7+ years in electrical engineering with expertise in high-voltage design and...Senior$300k
...thousands of H100s, H200s, and B200s, ready for experimentation, full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance, and automation of this GPU-powered infrastructure, ensuring...SeniorPermanent employment- ...Job Description Job Description Senior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will...Senior
- ...Creek Renewables in San Francisco or NYC is seeking a Senior Analyst, Structured Finance, to support financing late‑stage solar and battery storage projects. You will gain exposure to project development, engineering, and EPC workstreams while collaborating with senior...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Storage. Be the first to apply!
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- senior operations associate San Francisco, CA
- senior safety specialist San Francisco, CA
- senior technology project manager San Francisco, CA
- remote senior business analyst San Francisco, CA
- senior director fp&a San Francisco, CA
- senior manager clinical operations San Francisco, CA
- senior supervisor San Francisco, CA

