Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer - Managed Kubernetes

$240k - $356k
Full-time

Lambda

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.

If you'd like to build the world's best AI cloud, join us.

*Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.


Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.


What You’ll Do

  • Operate and maintain bare-metal Kubernetes clusters , scaling up to thousands of nodes

  • Handle cluster degradation, recovery, resizing, and incident response using fleet management tools

  • Participate in a well-managed on-call rotation for critical incidents

  • Assist customers with Kubernetes questions, workload integration, storage, and authentication

  • Work closely with our HPC Ops and Datacenter Ops teams for low-level or cross-functional issues

  • Use Python and Golang to create tooling and automate the validation of platform quality.

  • Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes

  • Develop automation for cluster lifecycle management : provisioning, upgrades, patching, and deletion.

  • Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.

You

  • 6+ years of experience in a SRE, operations engineer, or similar role, with a deep knowledge of running Linux clusters and systems

  • Strong programming skills in Go and Python ; experience with GitOps (e.g., ArgoCD), Helm, and Kubernetes operators

  • Proven experience operating Kubernetes clusters in production environments (on-prem, EKS, GKE, or similar)

  • Can work either independently with limited direction or as part of a team

  • Can work with customers during incidents either via tickets, live messaging, or as part of a larger call.

  • Familiarity with observability tools like Prometheus, Grafana, FluentBit , and CI/CD pipelines

  • Proven experience provisioning Kubernetes using tools such as kubeadm, Cluster API, or similar

Nice To Have

  • Deep Kubernetes expertise: CRDs, CSI, CNI, Kubernetes Operator Coding experience

  • Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters

  • Hybrid or multi-cloud Kubernetes environment experience

  • Contributions to CNCF projects or Kubernetes SIGs

Why Join Us

  • Work on cutting-edge Managed Kubernetes platforms for AI/ML workloads

  • Influence the platform roadmap and help shape operations and reliability best practices

  • Collaborate with a highly skilled engineer

  • Opportunity to mentor and grow within a fast-growing, technology-driven environment

Salary Range Information

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, with 500+ employees, and growing fast

  • Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove

  • We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Our values are publicly available:

  • We offer generous cash & equity compensation

  • Health, dental, and vision coverage for you and your dependents

  • Wellness and commuter stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible paid time off plan that we all actually use

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer - Managed Kubernetes in San Francisco, CA vacancy
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible...  ...our multi-cloud-provider Kubernetes infrastructure, networking,...  ...alerting systems.The Fleet Management team provides the core...  ...components that ensure cluster reliability and security (e.g., CoreDNS,... 
    Senior
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    San Francisco, CA
    3 days ago
  •  ...integrated solutions to manage everything from...  ...next.About the teamThe Engineering team at Airwallex is...  ...together to build scalable, reliable, and secure products...  ....What you’ll doAs a Senior Site Reliability Engineer,...  ...platforms (AWS/GCP), Kubernetes, observability, and... 
    Senior
    Temporary work
    Local area
    Worldwide

    Airwallex

    San Francisco, CA
    1 day ago
  • $180k - $210k

     ...actively seeking an exceptional Senior Software Engineer for our cloud software...  ...model, as well as managing critical hardware, software...  ...considering their impact on reliability, scalability, operational...  ...in advancing our managed Kubernetes and AI training clusters,... 
    Senior
    Temporary work

    Crusoe

    San Francisco, CA
    1 day ago
  • Apple Service Engineering (ASE) seeks a senior SRE software engineer to own the architectural direction of Kubernetes internals powering Apple services. You will define controllers and namespace management, raise reliability, and contribute to upstream Kubernetes. The... 
    Senior

    Socket.dev

    San Francisco, CA
    3 days ago
  • $148.5k - $223.9k

     ...Salesforce.Salesforce is seeking a senior engineering candidate to join the Site Reliability organization in San Francisco....  ...efficiency.Incident Management: Lead the coordinated response...  ...containerized architectures (Docker, Kubernetes) and orchestration platforms.Strong... 
    Senior
    Full time
    Worldwide
    Weekend work

    Salesforce

    San Francisco, CA
    9 hours ago
  • $220k - $235k

     ...for Contract Lifecycle Management, a Fortune Great...  ...strategic, high-output Staff/Senior Staff SRE to define...  ...and champion engineering excellence across Ironclad...  ...strategic direction for the Site Reliability Engineering team and...  ...knowledge of Kubernetes and Google Cloud Platform... 
    Senior
    Full time
    Contract work
    Work at office

    Ironclad

    San Francisco, CA
    9 hours ago
  • $300k

     ...full-scale model training, or inference. As a Platform Engineer/Senior Site Reliability Engineer, you’ll own the reliability, performance,...  ...ensuring seamless orchestration across environments managed by Slurm, Kubernetes, or direct SSH access. As well as supporting their... 
    Senior
    Permanent employment
    San Francisco, CA
    more than 2 months ago
  •  ...services, operators, and custom controllers for Kubernetes. Develop automation for cluster lifecycle management (provisioning, upgrades, patching, and deletion...  ...’s degree or foreign equivalent in Computer Engineering, Computer Science, Electrical Engineering, or related... 
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    9 hours ago
  • $194k - $267k

     ...automate it” and who can rapidly self-educate on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing Kubernetes platforms that support cloud-native applications and services. This position focuses on... 
    Permanent employment
    Work at office
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    3 days ago
  • $266k - $395k

     ...day is currently Tuesday. About the Role We are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial role in shaping the architecture, reliability, and automation of our Kubernetes-based infrastructure, which... 
    Senior
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  •  ...RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with...  ...(Docker), and orchestration (Kubernetes)Strong knowledge of Linux systems, networking...  ...-to-HaveExperience with cloud and managed services (e.g. AWS)Experience supporting... 
    Senior

    Alembic

    San Francisco, CA
    1 day ago
  • $152.5k - $205k

     ...What you’ll be responsible for:As a Senior Site Reliability Engineer on Circle’s platform team, you’ll design...  ...and automation, developing reliable Kubernetes platforms, and using Terraform to...  ...authoring reusable modules, managing state and environments, and delivering... 
    Senior
    Flexible hours

    Circle

    San Francisco, CA
    9 hours ago
  • $15k

     ...frontier of applying AI/ML to investment management. We have become a multibillion-...  ...catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our...  ...frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow,... 
    Senior
    Work at office
    Local area
    Remote work

    The Voleon Group

    Berkeley, CA
    4 days ago
  • $117k - $209.33k

     ...to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build...  ...SLIs, production readiness, incident management, observability, resilience testing,...  ...requirementsExperience with containers, Kubernetes, cloud-native architectures, APIs,... 
    Senior
    Full time
    For contractors

    Autodesk

    San Francisco, CA
    2 days ago
  • $165k - $225.6k

     ...across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this...  ...container orchestration environments (Kubernetes) and utilizing monitoring and... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    9 hours ago
  • $167.7k - $245.2k

     ...2 days per week on-site at Cisco offices in...  ...intended, improving reliability and reducing risks....  ...confidently deploy and manage AI-powered applications...  ...and control.As a Senior Site Reliability Engineer (SRE), you will build...  ...Operate and improve Kubernetes-based production... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours
    2 days per week

    CISCO Systems

    San Francisco, CA
    9 hours ago
  • $165k - $241.4k

     ...performance, efficiency, change management, monitoring, emergency...  ...looking for talented engineers with a software or...  ...teams to ensure the reliability, performance and...  ...Terraform, Puppet, and Kubernetes.Preferred QualificationsGood...  ...see the Cisco careers site to discover more... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    9 hours ago
  •  ...automagically help companies find and manage tools. Whether it is...  .... We're looking for a senior software engineer to not only amplify our...  ...with a focus on performance, reliability, and security....  ...like Google Cloud Run or Kubernetes. Solution-seeking. Strong... 
    Senior
    Full time
    Local area
    Remote work

    Brm

    San Francisco, CA
    9 hours ago
  •  ...A tech startup in San Francisco is looking for Site Reliability Engineers to enhance system reliability and performance. Ideal candidates have...  ...strong expertise in cloud infrastructure, including AWS and Kubernetes. The role involves defining SLIs/SLOs, optimizing... 
    Senior

    Breakout Tools

    San Francisco, CA
    1 day ago
  •  ...perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the...  ...network device configuration management end to end, ensuring consistency...  ...network peering.Familiarity with Kubernetes networking (CNI plugins,... 
    Senior

    Alembic

    San Francisco, CA
    3 days ago
  • $215k - $275k

     ...the role:Anyscale is looking for a Senior Site Reliability Engineer to join the Infrastructure team. Anyscale...  ...plane, which orchestrates cluster management, scheduling, and user access, and...  ..., along with expertise in Kubernetes, container orchestration, and cloud... 
    Senior
    Work at office

    Anyscale

    San Francisco, CA
    3 days ago
  • $170k - $230k

     ...one card issuer processor and program management platform. We give digital-first organizations...  ...be learned and care more about your engineering skill over frameworks  ~ Excitement...  ...GCP or AWS SpringBoot, Docker, and Kubernetes Experience with big data technologies... 
    Senior
    Full time
    Work at office
    Local area
    Home office
    Flexible hours

    Highnote

    San Francisco, CA
    9 hours ago
  •  ...Staff Engineer Lambda is building the AI Cloud of the future. We are seeking a Staff...  ...to help our development of our Managed Kubernetes platform. Think GKE, but purpose-built...  ...infrastructure to build systems that are reliable, performant, and elegantly simple for... 

    Lambda

    San Francisco, CA
    3 days ago
  • $13 per hour

     ...Agentforce Operations — managing complex, infinitely configurable...  ..., and enterprise-grade reliability at scale.What You'll...  ...seeking a highly experienced Senior, Lead, and Principal Engineers to serve as a key...  ...technologies (e.g., Docker, Kubernetes)Experience designing or working... 
    Senior
    Full time

    Salesforce

    San Francisco, CA
    2 days ago
  • $165k - $241.4k

     ...Your ImpactWe are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with...  ...and operations. You will design and manage large-scale, highly available distributed...  ...our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus,... 
    Senior
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours
    1 day per week

    CISCO Systems

    San Francisco, CA
    2 days ago
  • $200k - $260k

     ...started.Role OverviewAs a Software Engineer on the Site Reliability team at Harvey, you will ensure the...  ...What You'll DoDesign, implement, and manage monitoring, alerting, and infrastructure...  ...understanding of CI/CD, Kubernetes, containerization, networking, databases... 
    Senior
    Relocation package

    Harvey

    San Francisco, CA
    2 days ago
  •  ...drivers, and networking, managed as code (Ansible,...  ...working closely with on-site deployment teams Participate...  ...back to other engineering teams on...  ...of experience in Site Reliability Engineering, HPC Engineering...  ...technologies (e.g., Docker, Kubernetes) Experience building... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Flexible hours

    Lambda

    San Francisco, CA
    1 day ago
  • $212.5k - $250k

     ...infrastructure that founders use to manage equity, fund managers use to...  ...lifecycle so Carta's engineers can do the best work of their...  ...The Problems You’ll SolveAs a Senior Software Engineer II on DevEx...  ...cloud-native (AWS, Terraform, Kubernetes); familiarity with secrets management... 
    Senior
    Full time

    Carta

    San Francisco, CA
    2 days ago
  •  ...About the job Senior Site Reliability Engineer About the Company Stellar is a decentralized, public...  ...maintain, monitor and improve our Kubernetes clusters. Work with development...  ...hand experience with configuration management and infrastructure as code (Ansible... 
    Senior

    TechChain Talent

    San Francisco, CA
    5 days ago
  •  ...Cassandra, Kafka, Microservices, Spring Boot, Spark Streaming, Kubernetes, Docker Containers, and Splunk Logging &...  ...issues with performance, networking, kernel drivers, package management, etc. or have a good idea where to startYou have used at least... 
    Senior

    Phenom People

    San Francisco, CA
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer - Managed Kubernetes. Be the first to apply!