Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer II

Apex Systems

Site Reliability Engineer II

Location: Chandler, Arizona (Hybrid)

Duration: 12 months

Role Overview

This position is for a Site Reliability Engineer responsible for the reliability and support of on-premise and external cloud Container Platforms, including Azure, AWS, and Google. The role involves monitoring, troubleshooting, and enhancing the performance and security of container environments such as OpenShift, Rancher (RKE), and Azure (AKS). The ideal candidate will be a key stakeholder in the design of cloud services and will work to improve automation and operational excellence.

Key Responsibilities

  • Provide reliability and support for Container Platforms on-premise and in external clouds (Azure, AWS, Google).
  • Monitor and troubleshoot performance, connectivity, and security issues for OpenShift, Rancher (RKE), and Azure (AKS) environments.
  • Conduct deep dives into systemic and latent reliability issues, managing incidents and problems.
  • Identify, analyze, and resolve infrastructure vulnerabilities and application deployment issues.
  • Perform blameless Root Cause Analysis (RCA) and partner with engineering and operation teams to implement fixes.
  • Manage application onboarding and provide troubleshooting support throughout the application lifecycle.
  • Identify and implement automation opportunities to reduce operational tasks and improve efficiency.
  • Collaborate with risk and compliance teams to implement controls and remediate vulnerabilities.
  • Ensure resiliency during implementation and work with engineering teams to resolve resiliency problems.
  • Participate in a 24x7 on-call rotation following a follow-the-sun model.

Required Qualifications

Education: B.S./M.S. degree in Computer Science or a related technical field, or equivalent practical experience.

Experience: Minimum of 5+ years of hands-on experience supporting Kubernetes, OpenShift, RKE, or EKS Container platforms.

Technical Skills:

  • Experience with Python, Ansible, Golang, and shell scripting.
  • Experience with major services related to Compute, Storage, Network, and Security.
  • Experience with monitoring tools like Prometheus and Dynatrace, and cloud-native tools like Azure Monitor and Log Analytics.
  • Strong understanding of complex IAM infrastructure, including Active Directory, Azure AD, and SSO solutions like Ping Identity.
  • Advanced knowledge of Linux OS, DNS, DHCP, Kerberos, and Windows Authentication.
  • Experience with CI/CD tools such as Git and Jenkins, and GitOps models.
  • Excellent understanding of Linux/Windows operating systems administration.
  • Experience in container security and vulnerability remediation.

Preferred Qualifications

  • Experience in OpenShift, RKE, and CSP Kubernetes services such as AKS and EKS.
  • Experience in Terraform, ArgoCD, Tekton, and K-native technologies.
  • Experience in agile deployment methodologies (GitOps).
  • Knowledge of various container runtimes.
  • Familiarity with the operator deployment pattern.
  • Experience working in a highly available multi-datacenter environment.
  • Experience with monitoring tools such as Prometheus, Splunk, Dynatrace, or Sysdig.
  • Understanding of cost management, inventory management, and the FinOps model.
Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer II. Be the first to apply!