Senior Site Reliability Engineer Kubernetes Platform
SFE
Job Title: Senior Site Reliability Engineer Kubernetes Platform
Location: San Jose, CA
Full-Time
Job Description
Must Have Technical/Functional Skills:
10+ years of experience in SRE, DevOps, or infrastructure engineering
Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)
Hands-on experience working in FedRAMP High and/or DoD IL5 environments
Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals
Experience with Infrastructure as Code (Terraform preferred)
Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
Proficiency in scripting or programming (Python, Go)
Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)
Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)
Roles & Responsibilities:
Design, build, and operate production-grade Kubernetes platforms in regulated environments
Improve system reliability through automation, thoughtful design, and continuous iteration
Define and drive SLOs, SLIs, and error budgets to guide reliability decisions
Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
Implement robust observability (metrics, logs, traces) to make systems understandable and actionable
Reduce operational toil by automating repetitive processes and improving workflows
Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity
Support ATO processes, including documentation, controls implementation, and audit readiness Confidential
Participate in on-call rotations supporting customer requests and paging alerts
Participate in incident response, blameless postmortems, and continuous improvement efforts
Help shape a platform that engineers enjoy using
Nice to Have
Experience with service mesh technologies (Istio, Linkerd)
Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
Experience with GitOps workflows
Exposure to multi-cluster or hybrid cloud architectures
Knowledge of FIPS-compliant systems or DoD Cloud SRG
Relevant certifications (CKA, CKS, cloud provider certs, Security+)
- ...Job Title: Mid-Senior Site Reliability Engineer Kubernetes Platform Location: San Jose, CA Full-Time Job Description Must Have Technical/Functional Skills: 8+ years of experience in SRE, DevOps, or platform engineering Hands-on experience...SeniorFull time
- ...home day is currently Tuesday.Engineering at Lambda is responsible for... ...and maintain bare-metal Kubernetes clusters, scaling up to thousands... ...automate the validation of platform quality.Design, build, and... ...services, workloads, and platform reliability.You6+ years of experience in...SeniorWork at officeLocal areaWork from homeFlexible hours
- ...security, delivering an AI-powered platform that governs and secures... .... As a Staff Platform Engineer, you will play a critical role... ...leadership role. You will own reliability for major platform domains,... ...design scalable solutions on Kubernetes and AWS, and drive...Senior
- ...is currently Tuesday.Engineering at Lambda is responsible... ...cloud networking platform and SDN infrastructureOperate and improve Kubernetes-based control plane services... ...to improve service reliability and deployment workflowsDeploy... ...of experience in Site Reliability...SeniorWork at officeLocal areaWork from homeFlexible hours
$168k - $270.25k
...of artificial intelligence.Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing... ...technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).NVIDIA is widely considered to be one of the technology...SeniorFull time$267k - $356k
...Tuesday.Lambda's Storage Engineering team is the backbone... ...of Lambda's data platform services—from low-level... ...industry, which means reliability and performance aren't... ...across new and existing sites using tools such as Ansible... ...experience with Kubernetes (GitOps tooling such as...SeniorWork experience placementWork at officeLocal areaWork from homeFlexible hours$148k - $235.75k
...the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation... ...+ resource efficiency on track.Own Kubernetes deployments end-to-end (runbooks,...SeniorFull time$101k - $161k
...awards, such as Best Engineering Team, Best Company for... ...’re looking for Site Reliability Engineers to join our... ...stack is built entirely Kubernetes-native. Familiarity with... ...GCP (Google Cloud Platform) and GKE (Google Kubernetes... ...level: Mid-Senior LevelIndustry: Computer...Senior$152k - $241.5k
...intelligence.We’re looking for a Senior SRE to join our... ...our global services platform. At NVIDIA, you’ll... ...clusters using Slurm, LSF or Kubernetes clusters, including... ...management, fleet reliability/auto-healing, E2E observability... ...Ruby.Mentored other engineers and influenced...SeniorFull time$184k - $287.5k
...worldwide. Our team of skilled engineers is committed to addressing... ...the world!We are looking for a Senior Systems Software Engineer with strong experience in Kubernetes node engineering, OS image... ...depth needed to maintain cluster reliability at frontier AI scale. In this...SeniorFull timeWorldwide$184k - $287.5k
...looking for a hardworking Sr. Systems Software Engineer to work on platform software based on open-source container runtimes and Kubernetes technologies. We expect you to have... ...Understanding of performance, security and reliability in complex distributed systems.Ways to...SeniorFull timeWork experience placementRemote work$184k - $287.5k
...group of forward‑thinking engineers tackling some of the... ...We’re searching for a Senior Systems Software... ...distributed systems, Kubernetes, containers, and systems... ...serving, and major cloud platforms. You’ll own hard technical... ...projects to enable reliable operation at...SeniorFull timeRemote work- ....About the RoleWe are seeking a Senior Software Engineer to join our Managed Kubernetes (Mk8s) team. You will play a crucial... ...in shaping the architecture, reliability, and automation of our... ...critical workloads across our global platform.Lambda is building the AI Cloud...SeniorWork at officeLocal areaWork from homeFlexible hours
$102.5k - $187.9k
...actionable audit technology. Your key responsibilities As a Kubernetes DevOps Engineer, you are responsible to design, deploy, and manage containerized applications and orchestration platforms (Kubernetes) to ensure scalable, secure, and automated infrastructure...SeniorSummer holidayWork at officeFlexible hours- ...Company Description Mirantis is the Kubernetes-native AI infrastructure company,... ...Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable,... ...Overall system architecture, scalability, reliability, and performance experience...Senior
$222k - $300.5k
...financial technology platform that powers prosperity... ...s Infrastructure and Site Reliability organization owns the... ...Platform Systems Engineering team builds and operates... ...OpportunityWe're hiring a Senior Manager, Site... ...containerization/orchestration (Kubernetes), and infrastructure-...SeniorWorldwideShift work$168k - $270.25k
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain... ...cloud enabling technologies like Kubernetes and Public Cloud. SRE at NVIDIA ensures... ...AI Agents, AI Skills to accelerate platform operationsDrive automation and...Full time$207.4k - $259.2k
...an end-to-end advanced air mobility platform that delivers air taxis, unmanned aircraft... ...experienced and passionate Sr. Staff Site Reliability Engineer (SRE) to join our growing team. In... ...infrastructure on Amazon Elastic Kubernetes Service (EKS).Develop and implement comprehensive...SeniorPermanent employmentLocal areaWorldwideVisa sponsorship$124k - $271.2k
...Can ExpectAs a Lead Staff Site Reliability Engineer, you will be one of the technical... ...leads for our DevOps Platforms organization. This group is... ...to improve our datacenter kubernetes infrastructure, our cloud infrastructure... ...including security teams, senior leadership, and external...Full timeWork at officeRemote work$160k - $240k
...a day - quickly, reliably, and securely. Any... ...Fiserv.Job TitleSenior Site Reliability... ...Site Reliability Engineer do at Fiserv?You will... ...operate financial platforms at scale. You will... ...DevOps at a mid-to-senior level.Strong shell... ...and orchestration (Kubernetes).Working knowledge...SeniorFull time$192.4k - $275.8k
...demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service... ...automation and frameworks that make the whole platform more resilient. If you are the kind... ...you Your ImpactYou will be the most senior technical individual contributor on...SeniorFull timeTemporary workLocal areaFlexible hours$152k - $241.5k
Senior Systems Software Engineer, Observability and Telemetry Platform at NVIDIA is an engineering role to compose, build, and maintain... ...enabling technologies like Kubernetes and OpenStack. Senior Systems... ...cloud services run maximum reliability and uptime as promised to the...SeniorFull time$224k - $356.5k
...join our group of highly skilled and motivated engineers who bring GeForce NOW to life! As a member of the GeForce Now Platform Engineering team, you will help design,... ...platformsPrior experience with tools like GPUView, Kubernetes, Nvidia Nsight, Kernel shark, etc....SeniorFull time- ...seeking a hands-on Infrastructure Engineer responsible for building, operating, and scaling Kubernetes clusters on physical (bare-... ..., automation, and cluster reliability. The ideal candidate has... ...issues Collaborate with platform, DevOps, and application teams...
$174k - $252k
...networking features to enable AI/ML workloads in GDC platforms. alization solutions for Container/VM workloads running on Kubernetes platforms.Design and develop network... ...Windows (optional) and Linux OSs.Google's software engineers develop the next-generation technologies that...Senior$152k - $241.5k
...coordinated access control platform. Our UAM platform is a cutting... ..., enhance scalability and reliability, and boost the operational efficiency... ...Science, Electrical Engineering, Computer Engineering, or a... ...level knowledge of AWS Cloud, Kubernetes (k8s), and the GitOps model....SeniorFull time$168k - $270.25k
...of artificial intelligence.Join our team at NVIDIA as a Senior Storage Platform Engineer responsible for designing, deploying, and operating the... ...containerization technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).You've contributed to or maintained internal...SeniorFull time$280k - $380k
...the #1 TV streaming platform in the U.S., Canada,... ...internet scale. We focus on reliability and automation, engineering systems that perform... ...DevOps/SRE (Site Reliability Engineering) Senior Software Engineer to... ...number of the following: Kubernetes, Docker, Service Mesh...SeniorWork at officeLocal areaRemote workMonday to ThursdayFlexible hours$132.4k - $217.6k
...intelligence directly into our platforms to improve usability,... ...insight. This Senior Software Engineer role sits at the... ...of our San Jose, CA site. Responsibilities:Lead... ...that POC work can be reliably promoted to production... ...Azure Familiarity with Kubernetes and multi-service...SeniorFull timeTemporary workInternshipWorldwideFlexible hours- DDN is seeking a Senior Software Engineering Manager to lead the engineering organization... ...for our KV Cache Platform—a distributed memory and storage... ..., ensuring scalability, reliability, security, and operational... ...Storage, BlueField DPUs, Kubernetes, or related AI...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer Kubernetes Platform. Be the first to apply!
- site reliability engineer sre San Jose, CA
- site reliability engineer San Jose, CA
- platform engineer San Jose, CA
- senior platform engineer San Jose, CA
- platform developer San Jose, CA
- data platform engineer San Jose, CA
- senior operations technician San Jose, CA
- senior cloud service delivery manager San Jose, CA
- senior it service manager San Jose, CA
- senior project engineer San Jose, CA


