Senior Site Reliability Engineer - Managed Kubernetes
$240k - $356kLambda
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.
What You’ll Do
Operate and maintain bare-metal Kubernetes clusters , scaling up to thousands of nodes
Handle cluster degradation, recovery, resizing, and incident response using fleet management tools
Participate in a well-managed on-call rotation for critical incidents
Assist customers with Kubernetes questions, workload integration, storage, and authentication
Work closely with our HPC Ops and Datacenter Ops teams for low-level or cross-functional issues
Use Python and Golang to create tooling and automate the validation of platform quality.
Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes
Develop automation for cluster lifecycle management : provisioning, upgrades, patching, and deletion.
Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.
You
6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and systems
Strong programming skills in Go and Python ; experience with GitOps (e.g., ArgoCD), Helm, and Kubernetes operators
Proven experience operating Kubernetes clusters in production environments (on-prem, EKS, GKE, or similar)
Can work either independently with limited direction or as part of a team
Can work with customers during incidents either via tickets, live messaging, or as part of a larger call.
Familiarity with observability tools like Prometheus, Grafana, FluentBit , and CI/CD pipelines
Proven experience provisioning Kubernetes using tools such as kubeadm, Cluster API, or similar
Nice To Have
Deep Kubernetes expertise: CRDs, CSI, CNI, Kubernetes Operator Coding experience
Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters
Hybrid or multi-cloud Kubernetes environment experience
Contributions to CNCF projects or Kubernetes SIGs
Why Join Us
Work on cutting-edge Managed Kubernetes platforms for AI/ML workloads
Influence the platform roadmap and help shape operations and reliability best practices
Collaborate with a highly skilled engineer
Opportunity to mentor and grow within a fast-growing, technology-driven environment
Salary Range Information
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda
Founded in 2012, with 500+ employees, and growing fast
Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
Our values are publicly available:
We offer generous cash & equity compensation
Health, dental, and vision coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible paid time off plan that we all actually use
Equal Opportunity Employer
Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible... ...our multi-cloud-provider Kubernetes infrastructure, networking,... ...alerting systems.The Fleet Management team provides the core... ...components that ensure cluster reliability and security (e.g., CoreDNS,...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours- ...integrated solutions to manage everything from... ...next.About the teamThe Engineering team at Airwallex is... ...together to build scalable, reliable, and secure products... ....What you’ll doAs a Senior Site Reliability Engineer,... ...platforms (AWS/GCP), Kubernetes, observability, and...SeniorTemporary workLocal areaWorldwide
$180k - $210k
...actively seeking an exceptional Senior Software Engineer for our cloud software... ...model, as well as managing critical hardware, software... ...considering their impact on reliability, scalability, operational... ...in advancing our managed Kubernetes and AI training clusters,...SeniorTemporary work- Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services. You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The...Senior
$148.5k - $223.9k
...Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco.... ...efficiency.Incident Management: Lead the coordinated response... ...containerized architectures (Docker, Kubernetes) and orchestration platforms.Strong...SeniorFull timeWorldwideWeekend work$220k - $235k
...for Contract Lifecycle Management, a Fortune Great... ...strategic, high-output Staff/Senior Staff SRE to define... ...and champion engineering excellence across Ironclad... ...strategic direction for the Site Reliability Engineering team and... ...knowledge of Kubernetes and Google Cloud Platform...SeniorFull timeContract workWork at office$300k
...full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance,... ...ensuring seamless orchestration across environments managed by Slurm, Kubernetes, or direct SSH access. As well as supporting their...SeniorPermanent employment- ...services, operators, and custom controllers for Kubernetes. Develop automation for cluster lifecycle management (provisioning, upgrades, patching, and deletion... ...’s degree or foreign equivalent in Computer Engineering, Computer Science, Electrical Engineering, or related...Full timeWork experience placementWork at officeLocal areaRemote workWork from homeFlexible hours
$194k - $267k
...automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$266k - $395k
...day is currently Tuesday. About the Role We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial role in shaping the architecture, reliability, and automation of our Kubernetes-based infrastructure, which...SeniorFull timeWork at officeLocal areaWork from homeFlexible hours- ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with... ...(Docker), and orchestration (Kubernetes)Strong knowledge of Linux systems, networking... ...-to-HaveExperience with cloud and managed services (e.g. AWS)Experience supporting...Senior
$152.5k - $205k
...What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design... ...and automation, developing reliable Kubernetes platforms, and using Terraform to... ...authoring reusable modules, managing state and environments, and delivering...SeniorFlexible hours$15k
...frontier of applying AI/ML to investment management. We have become a multibillion-... ...catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our... ...frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow,...SeniorWork at officeLocal areaRemote work$117k - $209.33k
...to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build... ...SLIs, production readiness, incident management, observability, resilience testing,... ...requirementsExperience with containers, Kubernetes, cloud-native architectures, APIs,...SeniorFull timeFor contractors$165k - $225.6k
...across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this... ...container orchestration environments (Kubernetes) and utilizing monitoring and...SeniorPermanent employmentLocal areaWorldwideFlexible hours$167.7k - $245.2k
...2 days per week on-site at Cisco offices in... ...intended, improving reliability and reducing risks.... ...confidently deploy and manage AI-powered applications... ...and control.As a Senior Site Reliability Engineer (SRE), you will build... ...Operate and improve Kubernetes-based production...SeniorFull timeTemporary workLocal areaFlexible hours2 days per week$165k - $241.4k
...performance, efficiency, change management, monitoring, emergency... ...looking for talented engineers with a software or... ...teams to ensure the reliability, performance and... ...Terraform, Puppet, and Kubernetes.Preferred QualificationsGood... ...see the Cisco careers site to discover more...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week- ...automagically help companies find and manage tools. Whether it is... .... We're looking for a senior software engineer to not only amplify our... ...with a focus on performance, reliability, and security.... ...like Google Cloud Run or Kubernetes. Solution-seeking. Strong...SeniorFull timeLocal areaRemote work
- ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have... ...strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves defining SLIs/SLOs, optimizing...Senior
- ...perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the... ...network device configuration management end to end, ensuring consistency... ...network peering.Familiarity with Kubernetes networking (CNI plugins,...Senior
$215k - $275k
...the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale... ...plane, which orchestrates cluster management, scheduling, and user access, and... ..., along with expertise in Kubernetes, container orchestration, and cloud...SeniorWork at office$170k - $230k
...one card issuer processor and program management platform. We give digital-first organizations... ...be learned and care more about your engineering skill over frameworks ~ Excitement... ...GCP or AWS SpringBoot, Docker, and Kubernetes Experience with big data technologies...SeniorFull timeWork at officeLocal areaHome officeFlexible hours- ...Staff Engineer Lambda is building the AI Cloud of the future. We are seeking a Staff... ...to help our development of our Managed Kubernetes platform. Think GKE, but purpose-built... ...infrastructure to build systems that are reliable, performant, and elegantly simple for...
$13 per hour
...Agentforce Operations — managing complex, infinitely configurable... ..., and enterprise-grade reliability at scale.What You'll... ...seeking a highly experienced Senior, Lead, and Principal Engineers to serve as a key... ...technologies (e.g., Docker, Kubernetes)Experience designing or working...SeniorFull time$165k - $241.4k
...Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with... ...and operations. You will design and manage large-scale, highly available distributed... ...our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus,...SeniorFull timeTemporary workWork at officeLocal areaFlexible hours1 day per week$200k - $260k
...started.Role OverviewAs a Software Engineer on the Site Reliability team at Harvey, you will ensure the... ...What You'll DoDesign, implement, and manage monitoring, alerting, and infrastructure... ...understanding of CI/CD, Kubernetes, containerization, networking, databases...SeniorRelocation package- ...drivers, and networking, managed as code (Ansible,... ...working closely with on-site deployment teams Participate... ...back to other engineering teams on... ...of experience in Site Reliability Engineering, HPC Engineering... ...technologies (e.g., Docker, Kubernetes) Experience building...SeniorFull timeWork at officeLocal areaRemote workWork from homeFlexible hours
$212.5k - $250k
...infrastructure that founders use to manage equity, fund managers use to... ...lifecycle so Carta's engineers can do the best work of their... ...The Problems You’ll SolveAs a Senior Software Engineer II on DevEx... ...cloud-native (AWS, Terraform, Kubernetes); familiarity with secrets management...SeniorFull time- ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public... ...maintain, monitor and improve our Kubernetes clusters. Work with development... ...hand experience with configuration management and infrastructure as code (Ansible...Senior
- ...Cassandra, Kafka, Microservices, Spring Boot, Spark Streaming, Kubernetes, Docker Containers, and Splunk Logging &... ...issues with performance, networking, kernel drivers, package management, etc. or have a good idea where to startYou have used at least...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Managed Kubernetes. Be the first to apply!
- site reliability engineer remote San Francisco, CA
- site reliability engineer sre San Francisco, CA
- site reliability engineer San Francisco, CA
- senior operations associate San Francisco, CA
- senior safety specialist San Francisco, CA
- senior technology project manager San Francisco, CA
- remote senior business analyst San Francisco, CA
- senior director fp&a San Francisco, CA
- senior manager clinical operations San Francisco, CA
- senior supervisor San Francisco, CA



