Senior Site Reliability Engineer - Core Cloud Platform
Lambda Labs
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.If you'd like to build the world's best AI cloud, join us.*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.About the RoleLambda’s Core Cloud Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base grow.You will work across Kubernetes, infrastructure automation, observability, deployment systems, and incident response to build a resilient foundation for customer AI workloads.What You’ll DoOperate and scale critical platform services across Lambda’s data centers.Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures.Define SLIs, SLOs, error budgets, and operational readiness standards.Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures.Lead production incident response, postmortems, and durable corrective actions.Partner with Compute, Networking, Storage, Security, and Support teams.Participate in on-call and improve its sustainability through automation and better tooling.Mentor engineers and raise the reliability bar across the organization.YouHave 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering.Have deep experience operating Kubernetes in production.Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes.Have experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services.Are proficient with Terraform or similar infrastructure-as-code tools.Have built CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize.Have experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.Can build production-quality tooling in Go, Python, or a similar language.Understand distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure.Have experience defining and operating against SLIs and SLOs.Can lead effectively during high-severity incidents.Approach recurring operational issues as engineering and automation problems.Communicate clearly and work effectively across teams.Bring strong ownership, sound judgment, and low ego.Nice to HaveExperience with AI infrastructure, GPU platforms, or high-performance computing.Experience operating distributed systems across multiple regions or data centers.Experience with Kubernetes controllers, operators, CRDs, admission control, or scheduler extensions.Experience with etcd performance, backup, restore, or disaster recovery.Experience with Linux systems, container runtimes, cgroups, storage, or networking.Experience with chaos engineering, fault injection, or automated remediation.Experience with Kubernetes RBAC, OIDC, workload identity, or certificate management.Familiarity with SOC 2, ISO 27001, or similar compliance frameworks.Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.About LambdaFounded in 2012, with 500+ employees, and growing fastOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent CoveWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOGOur values are publicly available: We offer generous cash & equity compensationHealth, dental, and vision coverage for you and your dependentsWellness and commuter stipends for select roles401k Plan with 2% company match (USA employees)Flexible paid time off plan that we all actually useEqual Opportunity EmployerLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $240K - $356KLocationSan Francisco Office (Fremont St); Bellevue Office; San Jose Office (First St)Employment TypeFull timeLocation TypeHybridDepartmentData Center BusinessCompensationSan Francisco / San JoseSan Francisco / San Jose $267K – $356KBellevueBellevue $240K – $320K
$168k - $264.5k
...Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA... ...deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and... ...applications.Strong knowledge of the Kubernetes Platform, deployments, and cloud-native...PlatformSeniorFull time- Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...upgrades, and scaling.Own the reliability, performance, and security... ...networking, and RBAC across the platform.Lead incident response, root...PlatformSeniorWork at officeLocal areaWork from homeFlexible hours
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the... ...operations, increase efficiency in our use of cloud resources and our developer’s time,... ...through its Intelligent Operations Platform. Built to address the increasing...PlatformSeniorFlexible hours
$182.5k - $220k
...our team members.Role OverviewAs a Senior Staff Backend Engineer on the Platform Core team, you will build the services... ...to observability, reliability, and CI/CD standardsCollaborate with... ...typed service contractsAWS or other cloud-native infrastructure experienceStrong...PlatformSeniorLocal area- ...are.Disruption is at the core of our technology and on... ...DescriptionThe TeamPrisma Cloud PlatformWhy is the Prisma Cloud Platform team so special?Prisma Cloud... ...experienceYour Career Senior Principal Backend Developer... ...a Sr.Principal Software Engineer, you will break monoliths...PlatformSeniorFlexible hours
- ...The Superintelligence Cloud, is a leader in AI cloud... ...is currently Tuesday.Engineering at Lambda is... ...tenant cloud networking platform and SDN infrastructureOperate... ...teams to improve service reliability and deployment... ...years of experience in Site Reliability Engineering...PlatformSeniorWork at officeLocal areaWork from homeFlexible hours
- ...and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP... ...experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs...PlatformSeniorFull timeWork at office2 days per week
$262k - $365k
...design consulting, developing software platforms and frameworks, capacity planning and... ...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...'s degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines...PlatformSenior$101k - $161k
...data-driven, client-to-cloud networking for large... ...awards, such as Best Engineering Team, Best Company... ...WithWe’re looking for Site Reliability Engineers to join our... ...with GCP (Google Cloud Platform) and GKE (Google... ...EngineeringExperience level: Mid-Senior LevelIndustry:...PlatformSenior$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our... ...our global services platform. At NVIDIA, you’ll... ...globally distributed, multi‑cloud hybrid environment -... ...management, fleet reliability/auto-healing, E2E observability... ...Ruby.Mentored other engineers and influenced...PlatformSeniorFull time$159.2k - $301.6k
...control and search, (2) async compute platform for running Graphs on the cloud. In this reliability-focused role, you will own the... ...You'll partner with the backend engineers building these APIs to make sure... .... 5-10 years of experience in site reliability engineering,...PlatformSeniorFull timeTemporary workLocal areaWorldwide$210.6k - $305.1k
...and operates our US GovCloud platform. This team is responsible for... ...by AI and an unmatched set of cloud, internet and enterprise network... ...led a distributed team of 5+ engineers, can demonstrate strong technical... ...Please see the Cisco careers site to discover more benefits and...PlatformSeniorFull timeTemporary workLocal areaFlexible hours$187.04k - $359.72k
...changes that improve reliability and velocity. Qualifications... ...Science, Electrical Engineering, Computer Engineering... ...of the TikTok platform and U.S. user data, so... ...Functions and more. On-site presence across teams... ...creativity is at the core of TikTok's mission. Our...PlatformSeniorTemporary workLocal areaOverseasShift work- ...delivering an AI-powered platform that governs and secures... ...complex, distributed, cloud-native systems. As a Staff Platform Engineer, you will play a critical... ...leadership role. You will own reliability for major platform... ...teams to focus on their core business logic and deliver...PlatformSenior
$184k - $287.5k
NVIDIA’s accelerated computing platform is foundational to modern HPC... ...of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software. We are hiring a Senior Software Engineer to develop the Rust experience...PlatformSeniorFull time$184k - $287.5k
NVIDIA’s accelerated computing platform is foundational to modern... ...of this platform are CUDA Core Libraries that provide the... ...capabilities needed to build fast, reliable, and scalable GPU-... ...accelerated software.We are hiring a Senior Software Engineer to advance the C++...PlatformSeniorFull time$184k - $287.5k
NVIDIA’s accelerated computing platform is foundational to modern HPC... ...of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software.We are hiring a Senior Software Engineer to advance the Python experience...PlatformSeniorFull time$174k - $253k
...consulting, developing software platforms and frameworks, capacity... ...pushing for changes that improve reliability and velocity.Practice... ...degree in Computer Science, Engineering, a related field, or equivalent... ...Science or Engineering.Site Reliability Engineering (SRE...PlatformSenior$174k - $253k
...accessible technologies.Google's software engineers develop the next-generation technologies... ..., and enhance software solutions.The Core team builds the technical foundation behind... ...underlying design elements, developer platforms, product components, and infrastructure...PlatformSenior$184k - $287.5k
We are looking for an experienced Senior System Software Engineer to help build NVIDIA Halos. Halos is our full-stack safety platform for physical AI: autonomous vehicles, humanoids, industrial robots, and intelligent machines that sense, decide, and act in the real world...PlatformSeniorFull time$90k - $180k
...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our... ...of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists... ...demands, including distributed systems and cloud deployments in Azure. Work closely...PlatformSeniorRemote work$207.4k - $259.2k
...experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing... ..., and security of our core systems and services. You... ..., potentially leveraging platforms like OpenRouter or similar... ...available, scalable, and secure cloud-native infrastructure on...PlatformSeniorPermanent employmentLocal area$184k - $287.5k
Our Autonomous Vehicles Platform team is searching for engineers to develop and bring NVIDIA's automotive platform out to the world. You will participate in a focused effort to develop and productize ground-breaking solutions that will revolutionize the world of transportation...PlatformSeniorFull time$262k - $365k
...a team of Software/Systems Engineers on projects for users and be... ...Computer Science or Engineering.Site Reliability Engineering (SRE) combines... ...chose to join SRE.As the Senior Engineering Manager for Collaboration... ...’s flagship collaboration platform.Behind everything our users...PlatformSenior$152k - $241.5k
The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to new markets! This role involves developing and productizing innovative solutions that will transform transportation and the field of...PlatformSeniorFull time- DDN is seeking a Senior Software Engineering Manager to lead the engineering organization... ...for our KV Cache Platform—a distributed memory and storage... ..., ensuring scalability, reliability, security, and operational... ...distributed systems, cloud infrastructure, storage platforms...PlatformSenior
$224k - $356.5k
...their best work. Come join the team and see how you can make a lasting impact on the world.We are now looking for a Senior Software Performance Engineer for Autonomous Vehicles! Our team builds NVIDIA’s end-to-end autonomous driving applications. We are seeking senior...PlatformSeniorFull time$184k - $287.5k
The Autonomous Vehicles Platform team is looking for a hands-on System Software Engineer. As part of our team, you will work on our Autonomous Driving Platform software... ...stack, from platform and embedded software to cloud infrastructure, underpinned by safety and...PlatformSeniorFull time$184k - $287.5k
We are looking for a senior systems software engineer to improve the operation and user experience of distributed... ...the future of the NVIDIA software platform! We want to develop next-generation... ...thousands of servers efficiently, reliably, and securely.What you will be doing...PlatformSeniorFull time$152k - $241.5k
...creative Factory System Software and Diagnostics Integration engineer to join the Datacenter Platform Software team. You will play a crucial role in... ...address issues, and propose solutions to enhance system reliability and efficiency. Collaborate with vendors and external...PlatformSeniorFull time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Core Cloud Platform. Be the first to apply!
- senior cloud solutions architect San Jose, CA
- senior cloud engineer San Jose, CA
- software engineer - cloud services San Jose, CA
- cloud engineer remote San Jose, CA
- remote cloud architect San Jose, CA
- big data cloud engineer San Jose, CA
- cloud developer San Jose, CA
- senior cloud data engineer San Jose, CA
- senior principal cloud computing engineer San Jose, CA
- aws cloud infrastructure engineer San Jose, CA

