Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer Kubernetes Platform

Full-time

SFE

Job Title: Senior Site Reliability Engineer Kubernetes Platform

Location: San Jose, CA

Full-Time




Job Description

Must Have Technical/Functional Skills:

10+ years of experience in SRE, DevOps, or infrastructure engineering

Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)

Hands-on experience working in FedRAMP High and/or DoD IL5 environments

Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals

Experience with Infrastructure as Code (Terraform preferred)

Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)

Proficiency in scripting or programming (Python, Go)

Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)

Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)

Roles & Responsibilities:

Design, build, and operate production-grade Kubernetes platforms in regulated environments

Improve system reliability through automation, thoughtful design, and continuous iteration

Define and drive SLOs, SLIs, and error budgets to guide reliability decisions

Build and evolve CI/CD pipelines that are secure, scalable, and easy to use

Implement robust observability (metrics, logs, traces) to make systems understandable and actionable

Reduce operational toil by automating repetitive processes and improving workflows

Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity

Support ATO processes, including documentation, controls implementation, and audit readiness Confidential

Participate in on-call rotations supporting customer requests and paging alerts

Participate in incident response, blameless postmortems, and continuous improvement efforts

Help shape a platform that engineers enjoy using

Nice to Have

Experience with service mesh technologies (Istio, Linkerd)

Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)

Experience with GitOps workflows

Exposure to multi-cluster or hybrid cloud architectures

Knowledge of FIPS-compliant systems or DoD Cloud SRG

Relevant certifications (CKA, CKS, cloud provider certs, Security+)

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer Kubernetes Platform in San Jose, CA vacancy
  •  ...Job Title: Mid-Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must Have Technical/Functional Skills: 8+ years of experience in SRE, DevOps, or platform engineering Hands-on experience... 
    Senior
    Full time

    SFE

    San Jose, CA
    2 days ago
  •  ...home day is currently Tuesday.Engineering at Lambda is responsible for...  ...and maintain bare-metal Kubernetes clusters, scaling up to thousands...  ...automate the validation of platform quality.Design, build, and...  ...services, workloads, and platform reliability.You6+ years of experience in... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  •  ...security, delivering an AI-powered platform that governs and secures...  .... As a Staff Platform Engineer, you will play a critical role...  ...leadership role. You will own reliability for major platform domains,...  ...design scalable solutions on Kubernetes and AWS, and drive... 
    Senior

    Saviynt

    Milpitas, CA
    a month ago
  •  ...is currently Tuesday.Engineering at Lambda is responsible...  ...cloud networking platform and SDN infrastructureOperate and improve Kubernetes-based control plane services...  ...to improve service reliability and deployment workflowsDeploy...  ...of experience in Site Reliability... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $168k - $270.25k

     ...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing...  ...technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).NVIDIA is widely considered to be one of the technology... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    21 hours ago
  • $267k - $356k

     ...Tuesday.Lambda's Storage Engineering team is the backbone...  ...of Lambda's data platform services—from low-level...  ...industry, which means reliability and performance aren't...  ...across new and existing sites using tools such as Ansible...  ...experience with Kubernetes (GitOps tooling such as... 
    Senior
    Work experience placement
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    21 hours ago
  • $148k - $235.75k

     ...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation...  ...+ resource efficiency on track.Own Kubernetes deployments end-to-end (runbooks,... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $101k - $161k

     ...awards, such as Best Engineering Team, Best Company for...  ...’re looking for Site Reliability Engineers to join our...  ...stack is built entirely Kubernetes-native. Familiarity with...  ...GCP (Google Cloud Platform) and GKE (Google Kubernetes...  ...level: Mid-Senior LevelIndustry: Computer... 
    Senior

    Arista Networks

    Santa Clara, CA
    21 hours ago
  • $152k - $241.5k

     ...intelligence.We’re looking for a Senior SRE to join our...  ...our global services platform. At NVIDIA, you’ll...  ...clusters using Slurm, LSF or Kubernetes clusters, including...  ...management, fleet reliability/auto-healing, E2E observability...  ...Ruby.Mentored other engineers and influenced... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $184k - $287.5k

     ...worldwide. Our team of skilled engineers is committed to addressing...  ...the world!We are looking for a Senior Systems Software Engineer with strong experience in Kubernetes node engineering, OS image...  ...depth needed to maintain cluster reliability at frontier AI scale. In this... 
    Senior
    Full time
    Worldwide

    Nvidia

    Santa Clara, CA
    1 day ago
  • $184k - $287.5k

     ...looking for a hardworking Sr. Systems Software Engineer to work on platform software based on open-source container runtimes and Kubernetes technologies. We expect you to have...  ...Understanding of performance, security and reliability in complex distributed systems.Ways to... 
    Senior
    Full time
    Work experience placement
    Remote work

    Nvidia

    Santa Clara, CA
    2 days ago
  • $184k - $287.5k

     ...group of forward‑thinking engineers tackling some of the...  ...We’re searching for a Senior Systems Software...  ...distributed systems, Kubernetes, containers, and systems...  ...serving, and major cloud platforms. You’ll own hard technical...  ...projects to enable reliable operation at... 
    Senior
    Full time
    Remote work

    Nvidia

    Santa Clara, CA
    1 day ago
  •  ....About the RoleWe are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial...  ...in shaping the architecture, reliability, and automation of our...  ...critical workloads across our global platform.Lambda is building the AI Cloud... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    1 day ago
  • $102.5k - $187.9k

     ...actionable audit technology. Your key responsibilities As a Kubernetes DevOps Engineer, you are responsible to design, deploy, and manage containerized applications and orchestration platforms (Kubernetes) to ensure scalable, secure, and automated infrastructure... 
    Senior
    Summer holiday
    Work at office
    Flexible hours

    Ernst & Young

    San Jose, CA
    more than 2 months ago
  •  ...Company Description Mirantis is the Kubernetes-native AI infrastructure company,...  ...Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable,...  ...Overall system architecture, scalability, reliability, and performance experience... 
    Senior

    Mirantis

    San Jose, CA
    21 days ago
  • $222k - $300.5k

     ...financial technology platform that powers prosperity...  ...s Infrastructure and Site Reliability organization owns the...  ...Platform Systems Engineering team builds and operates...  ...OpportunityWe're hiring a Senior Manager, Site...  ...containerization/orchestration (Kubernetes), and infrastructure-... 
    Senior
    Worldwide
    Shift work

    Intuit

    Mountain View, CA
    21 hours ago
  • $168k - $270.25k

    Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain...  ...cloud enabling technologies like Kubernetes and Public Cloud. SRE at NVIDIA ensures...  ...AI Agents, AI Skills to accelerate platform operationsDrive automation and... 
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $207.4k - $259.2k

     ...an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft...  ...experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In...  ...infrastructure on Amazon Elastic Kubernetes Service (EKS).Develop and implement comprehensive... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Visa sponsorship

    Archer Aviation

    San Jose, CA
    21 hours ago
  • $124k - $271.2k

     ...Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical...  ...leads for our DevOps Platforms organization. This group is...  ...to improve our datacenter kubernetes infrastructure, our cloud infrastructure...  ...including security teams, senior leadership, and external... 
    Full time
    Work at office
    Remote work

    Zoom

    San Jose, CA
    1 day ago
  • $160k - $240k

     ...a day - quickly, reliably, and securely. Any...  ...Fiserv.Job TitleSenior Site Reliability...  ...Site Reliability Engineer do at Fiserv?You will...  ...operate financial platforms at scale. You will...  ...DevOps at a mid-to-senior level.Strong shell...  ...and orchestration (Kubernetes).Working knowledge... 
    Senior
    Full time

    Fiserv

    Sunnyvale, CA
    21 hours ago
  • $192.4k - $275.8k

     ...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service...  ...automation and frameworks that make the whole platform more resilient. If you are the kind...  ...you Your ImpactYou will be the most senior technical individual contributor on... 
    Senior
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    San Jose, CA
    3 days ago
  • $152k - $241.5k

    Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA is an engineering role to compose, build, and maintain...  ...enabling technologies like Kubernetes and OpenStack. Senior Systems...  ...cloud services run maximum reliability and uptime as promised to the... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    1 day ago
  • $224k - $356.5k

     ...join our group of highly skilled and motivated engineers who bring GeForce NOW to life! As a member of the GeForce Now Platform Engineering team, you will help design,...  ...platformsPrior experience with tools like GPUView, Kubernetes, Nvidia Nsight, Kernel shark, etc.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  •  ...seeking a hands-on Infrastructure Engineer responsible for building, operating, and scaling Kubernetes clusters on physical (bare-...  ..., automation, and cluster reliability. The ideal candidate has...  ...issues Collaborate with platform, DevOps, and application teams... 

    Purple Drive

    Milpitas, CA
    1 day ago
  • $174k - $252k

     ...networking features to enable AI/ML workloads in GDC platforms. alization solutions for Container/VM workloads running on Kubernetes platforms.Design and develop network...  ...Windows (optional) and Linux OSs.Google's software engineers develop the next-generation technologies that... 
    Senior

    Google

    Sunnyvale, CA
    3 days ago
  • $152k - $241.5k

     ...coordinated access control platform. Our UAM platform is a cutting...  ..., enhance scalability and reliability, and boost the operational efficiency...  ...Science, Electrical Engineering, Computer Engineering, or a...  ...level knowledge of AWS Cloud, Kubernetes (k8s), and the GitOps model.... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    2 days ago
  • $168k - $270.25k

     ...of artificial intelligence.Join our team at NVIDIA as a Senior Storage Platform Engineer responsible for designing, deploying, and operating the...  ...containerization technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).You've contributed to or maintained internal... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $280k - $380k

     ...the #1 TV streaming platform in the U.S., Canada,...  ...internet scale. We focus on reliability and automation, engineering systems that perform...  ...DevOps/SRE (Site Reliability Engineering) Senior Software Engineer to...  ...number of the following: Kubernetes, Docker, Service Mesh... 
    Senior
    Work at office
    Local area
    Remote work
    Monday to Thursday
    Flexible hours

    Roku

    San Jose, CA
    2 days ago
  • $132.4k - $217.6k

     ...intelligence directly into our platforms to improve usability,...  ...insight. This Senior Software Engineer role sits at the...  ...of our San Jose, CA site. Responsibilities:Lead...  ...that POC work can be reliably promoted to production...  ...Azure Familiarity with Kubernetes and multi-service... 
    Senior
    Full time
    Temporary work
    Internship
    Worldwide
    Flexible hours

    Bio-Techne

    San Jose, CA
    2 days ago
  • DDN is seeking a Senior Software Engineering Manager to lead the engineering organization...  ...for our KV Cache Platform—a distributed memory and storage...  ..., ensuring scalability, reliability, security, and operational...  ...Storage, BlueField DPUs, Kubernetes, or related AI... 
    Senior

    DataDirect Networks

    Santa Clara, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer Kubernetes Platform. Be the first to apply!