Senior Site Reliability Engineer - Core Cloud Platform
Lambda Labs
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.If you'd like to build the world's best AI cloud, join us.*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.About the RoleLambda’s Core Cloud Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base grow.You will work across Kubernetes, infrastructure automation, observability, deployment systems, and incident response to build a resilient foundation for customer AI workloads.What You’ll DoOperate and scale critical platform services across Lambda’s data centers.Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures.Define SLIs, SLOs, error budgets, and operational readiness standards.Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures.Lead production incident response, postmortems, and durable corrective actions.Partner with Compute, Networking, Storage, Security, and Support teams.Participate in on-call and improve its sustainability through automation and better tooling.Mentor engineers and raise the reliability bar across the organization.YouHave 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering.Have deep experience operating Kubernetes in production.Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes.Have experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services.Are proficient with Terraform or similar infrastructure-as-code tools.Have built CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize.Have experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog.Can build production-quality tooling in Go, Python, or a similar language.Understand distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure.Have experience defining and operating against SLIs and SLOs.Can lead effectively during high-severity incidents.Approach recurring operational issues as engineering and automation problems.Communicate clearly and work effectively across teams.Bring strong ownership, sound judgment, and low ego.Nice to HaveExperience with AI infrastructure, GPU platforms, or high-performance computing.Experience operating distributed systems across multiple regions or data centers.Experience with Kubernetes controllers, operators, CRDs, admission control, or scheduler extensions.Experience with etcd performance, backup, restore, or disaster recovery.Experience with Linux systems, container runtimes, cgroups, storage, or networking.Experience with chaos engineering, fault injection, or automated remediation.Experience with Kubernetes RBAC, OIDC, workload identity, or certificate management.Familiarity with SOC 2, ISO 27001, or similar compliance frameworks.Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.About LambdaFounded in 2012, with 500+ employees, and growing fastOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent CoveWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOGOur values are publicly available: We offer generous cash & equity compensationHealth, dental, and vision coverage for you and your dependentsWellness and commuter stipends for select roles401k Plan with 2% company match (USA employees)Flexible paid time off plan that we all actually useEqual Opportunity EmployerLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $240K - $356KLocationSan Francisco Office (Fremont St); Bellevue Office; San Jose Office (First St)Employment TypeFull timeLocation TypeHybridDepartmentData Center BusinessCompensationSan Francisco / San JoseSan Francisco / San Jose $267K – $356KBellevueBellevue $240K – $320K
$168k - $264.5k
...Digital Marketing Organization seeks a senior Site Reliability Engineer (SRE) to join our Santa Clara, CA... ...deployment pipelines, Akamai CDN, WAF, and cloud infrastructure.Author, test, and... ...applications.Strong knowledge of the Kubernetes Platform, deployments, and cloud-native...PlatformSeniorFull time- Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...upgrades, and scaling.Own the reliability, performance, and security... ...networking, and RBAC across the platform.Lead incident response, root...PlatformSeniorWork at officeLocal areaWork from homeFlexible hours
- Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the... ...operations, increase efficiency in our use of cloud resources and our developer’s time,... ...through its Intelligent Operations Platform. Built to address the increasing...PlatformSeniorFlexible hours
- ...are.Disruption is at the core of our technology and on... ...DescriptionThe TeamPrisma Cloud PlatformWhy is the Prisma Cloud Platform team so special?Prisma Cloud... ...experienceYour Career Senior Principal Backend Developer... ...a Sr.Principal Software Engineer, you will break monoliths...PlatformSeniorFlexible hours
$168k - $270.25k
NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at... ...our internal and external-facing GPU cloud gaming services have reliability and... ...consulting, developing software platforms and frameworks, capacity management...PlatformSeniorFull time- ...The Superintelligence Cloud, is a leader in AI cloud... ...is currently Tuesday.Engineering at Lambda is... ...tenant cloud networking platform and SDN infrastructureOperate... ...teams to improve service reliability and deployment... ...years of experience in Site Reliability Engineering...PlatformSeniorWork at officeLocal areaWork from homeFlexible hours
- Lambda, The Superintelligence Cloud, is a leader in AI cloud... ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...automate the validation of platform quality.Design, build, and maintain... ..., workloads, and platform reliability.You6+ years of experience in...PlatformSeniorWork at officeLocal areaWork from homeFlexible hours
- ...and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP... ...experience configuring New Relic (or similar platforms) to create meaningful dashboards, SLIs, and SLOs...PlatformSeniorFull timeWork at office2 days per week
$267k - $356k
...Superintelligence Cloud, is a leader in AI... ....Lambda's Storage Engineering team is the backbone... ...of Lambda's data platform services—from low-... ..., which means reliability and performance aren... ...new and existing sites using tools such as... ...understanding of core storage protocols...PlatformSeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$101k - $161k
...data-driven, client-to-cloud networking for large... ...awards, such as Best Engineering Team, Best Company... ...WithWe’re looking for Site Reliability Engineers to join our... ...with GCP (Google Cloud Platform) and GKE (Google... ...EngineeringExperience level: Mid-Senior LevelIndustry:...PlatformSenior$262k - $364k
...design consulting, developing software platforms and frameworks, capacity planning and... ...systems by pushing for changes that improve reliability and velocity.Practice sustainable... ...'s degree in Computer Science or Engineering.Site Reliability Engineering (SRE) combines...PlatformSenior$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our... ...our global services platform. At NVIDIA, you’ll... ...globally distributed, multi‑cloud hybrid environment -... ...management, fleet reliability/auto-healing, E2E observability... ...Ruby.Mentored other engineers and influenced...PlatformSeniorFull time$167.7k - $245.2k
...2 days per week on-site at Cisco offices in... ...intended, improving reliability and reducing risks.... ...observability and control.As a Senior Site Reliability Engineer (SRE), you will... ...'s deployment platform and production infrastructure... ...supporting both cloud and air-gapped...PlatformSeniorFull timeTemporary workLocal areaFlexible hours2 days per week$146.7k - $339.3k
...this positionWhat you can expect As a Senior Lead Site Reliability Engineer, you can anticipate opportunities to... ...maintain thousands of physical and cloud systems worldwide. To streamline... ...Partner with Security, Networking, and Platform teams on architecture roadmaps. Influence...PlatformSeniorFull timeWork at officeRemote workWorldwideShift workWeekend work$210.6k - $305.1k
...and operates our US GovCloud platform. This team is responsible for... ...by AI and an unmatched set of cloud, internet and enterprise network... ...led a distributed team of 5+ engineers, can demonstrate strong technical... ...Please see the Cisco careers site to discover more benefits and...PlatformSeniorFull timeTemporary workLocal areaFlexible hours$144k - $216k
# Senior Software Engineer, Core Platform## FloQastPublished 16 Jul 2026### Share this jobSan Jose, CA, USA144K - 216K USD AnnualFull Time## Role Highlights###... ...The position requires maintaining high standards for reliability, security, and scalability to support product...PlatformSenior- Speechify is seeking a Senior Software Engineer to join the Core Experiences Team. This role... ...powering Speechify across platforms, blending product and... ...infrastructure expertise to design reliable APIs and simple systems... ...performance, and ship cloud functions, lightweight...PlatformSeniorRemote work
$101k - $161k
...experience. We look for 5+ years of software engineering experience. We need experience... ...lead projects across areas such as data platform architecture and performance, capacity... ...network architecture, cost optimization, and cloud-first application security. We will...PlatformSenior$184k - $287.5k
NVIDIA’s accelerated computing platform is foundational to modern HPC... ...of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software. We are hiring a Senior Software Engineer to develop the Rust experience...PlatformSeniorFull time$184k - $287.5k
NVIDIA’s accelerated computing platform is foundational to modern HPC... ...of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software.We are hiring a Senior Software Engineer to advance the Python experience...PlatformSeniorFull time$184k - $287.5k
NVIDIA’s accelerated computing platform is foundational to modern... ...of this platform are CUDA Core Libraries that provide the... ...capabilities needed to build fast, reliable, and scalable GPU-... ...accelerated software.We are hiring a Senior Software Engineer to advance the C++...PlatformSeniorFull time$174k - $252k
...accessible technologies.Google's software engineers develop the next-generation technologies... ..., and enhance software solutions.The Core team builds the technical foundation behind... ...underlying design elements, developer platforms, product components, and infrastructure...PlatformSenior$90k - $180k
...than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our... ...of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists... ...demands, including distributed systems and cloud deployments in Azure. Work closely...PlatformSeniorRemote work$160k - $240k
...a day - quickly, reliably, and securely. Any... ...Fiserv.Job TitleSenior Site Reliability... ...Site Reliability Engineer do at Fiserv?You will... ...operate financial platforms at scale. You will... ...improvement across our cloud-native... ...DevOps at a mid-to-senior level.Strong shell...PlatformSeniorFull time$184k - $287.5k
We are looking for an experienced Senior System Software Engineer to help build NVIDIA Halos. Halos is our full-stack safety platform for physical AI: autonomous vehicles, humanoids, industrial robots, and intelligent machines that sense, decide, and act in the real world...PlatformSeniorFull time$168k - $270.25k
...deploy and run an AI data center. We take great pride in providing excellent, comprehensive support to our customers! Sr Site Reliability Engineer in this role will significantly impact and contribute to the overall success of both external customers running their clusters...SeniorFull timeWorldwide$184k - $287.5k
Our Autonomous Vehicles Platform team is searching for engineers to develop and bring NVIDIA's automotive platform out to the world. You will participate in a focused effort to develop and productize ground-breaking solutions that will revolutionize the world of transportation...PlatformSeniorFull time$272k - $431.25k
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team... ...like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers...PlatformFull timeWork experience placementWorldwide$152k - $241.5k
The Autonomous Vehicles Platform team is seeking a Senior System Software Engineer to help bring NVIDIA's autonomous vehicle platform to new markets! This role involves developing and productizing innovative solutions that will transform transportation and the field of...PlatformSeniorFull time- DDN is seeking a Senior Software Engineering Manager to lead the engineering organization... ...for our KV Cache Platform—a distributed memory and storage... ..., ensuring scalability, reliability, security, and operational... ...distributed systems, cloud infrastructure, storage platforms...PlatformSenior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Core Cloud Platform. Be the first to apply!
- site reliability engineer San Jose, CA
- site reliability engineer sre San Jose, CA
- senior cloud security engineer San Jose, CA
- senior cloud solutions architect San Jose, CA
- senior cloud data engineer San Jose, CA
- cloud engineering manager San Jose, CA
- informatica cloud developer San Jose, CA
- senior cloud engineer San Jose, CA
- cloud architect San Jose, CA
- aws cloud architect San Jose, CA

