Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America ATM

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.

Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.

We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.

Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.

At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Position Summary:

The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:

Reliability & Operations

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Kubernetes & OpenShift Platform Engineering

  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption. 

Observability & Automation

  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.

Capacity & Performance Engineering

  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.

Security & Compliance

  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.

Required Qualifications:

Education / Experience

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Technical Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.

Desired Qualifications

  • BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications.
  • Experience running large-scale enterprise container platforms.
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
  • Experience implementing cloud-native security controls and platform governance.
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
  • Experience with vulnerability management and container security scanning solutions.
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicate effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.


This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices.

Responsibilities:

  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Participates regularly in architecture community of practice meetings and communication via other channels
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations

Skills:

  • Automation
  • Collaboration
  • Influence
  • Production Support
  • Result Orientation
  • Analytical Thinking
  • Application Development
  • Architecture
  • Solution Design
  • Stakeholder Management
  • Adaptability
  • DevOps Practices
  • Project Management
  • Risk Management
  • Solution Delivery Process

Shift:

1st shift (United States of America)

Hours Per Week: 

40
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP) in Chandler, AZ vacancy
  • Job DescriptionWe are seeking an experienced Site Reliability Engineer SRE Lead Messaging Services to drive platform reliability observability and operational excellence...  ...for messaging workloadsFamiliarity witho Kubernetes containerized messaging platformsExperience witho... 
    Platform

    LTM

    Chandler, AZ
    1 day ago
  • $152.6k - $191.5k

     ...leaders across engineering and technology...  ...objective reliability goals for services...  ...Senior Azure Site Reliability Engineer...  ...Azure platform. The role focuses...  ....Mentor SRE engineers and...  ...with AKS, ACR, Kubernetes, container networking, CI...  ...provide industry-leading benefits, access... 
    Platform
    Full time
    Work at office
    Day shift

    Bank of America

    Chandler, AZ
    2 days ago
  •  ...skilled Software Engineer with expert-level...  ...design capabilities, Kubernetes development...  ...enterprise-ready platform capabilities.This...  ...design and build reliable distributed systems...  ...platform engineering, SRE, security, DevOps,...  ...or contributed to internal developer platforms... 
    Platform
    Full time
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America

    Chandler, AZ
    3 days ago
  •  ...and financial platform for global businesses...  ...by world-leading investors...  ...responsible for the reliability, performance,...  ...: product engineers should be able...  ...the Database SRE team, you will...  ...of AI-powered internal tooling while...  ...(Terraform), Kubernetes, and Docker. Demonstrated... 
    Platform
    Worldwide

    GrabJobs

    Gilbert, AZ
    4 days ago
  •  ...below: Collaborate with internal and client teams to...  ..., TKG and managed Kubernetes services such as AKS,...  ...specifically as it relates to containers.• Experience in agile...  ...effectively across engineering, infrastructure,...  ...Technology|Container Platform|Kubernetes Technical/... 
    Platform
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Chandler, AZ
    3 days ago
  • $160k - $240k

     ...work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering organization and help us scale...  ...Ansible. CDK a plus. Good understanding of containers, Fargate, Kubernetes, and overall distributed microservice architectures... 
    Platform
    Permanent employment
    Full time
    Remote work
    Work from home
    Relocation
    Flexible hours

    GrabJobs

    Chandler, AZ
    1 day ago
  • $122k - $200k

     ...The Senior Software Engineer - Kubernetes Platform Engineering, Service...  ...the bank's enterprise container ecosystem. This role...  ...The individual will lead the architecture and...  ..., networking, SRE, DevOps, and application...  ...automationObservability, reliability, and platform... 
    Platform
    Full time
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America

    Chandler, AZ
    1 day ago
  • $70.24 per hour

     ...Description* We are seeking an experienced Site Reliability Engineer (SRE) to join a high-performing cloud...  ...portfolio of enterprise cloud platforms. This individual will serve as a true...  ...TEKsystems Global Services We're a leading provider of business and technology... 
    Platform
    Permanent employment
    Contract work
    Temporary work

    TEKsystems

    Chandler, AZ
    1 day ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type...  ...focus on ensuring the platform and services customers...  ...You will interact with internal staff, managers, and customers...  ...on experience of Azure Kubernetes Services (AKS) with container-based deployment skills... 
    Platform
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Mesa, AZ
    4 days ago
  •  ...Chandler, AZ seeks a senior Software Engineer with expert-level Go...  ...build scalable backend services, platform APIs, and Kubernetes‑native automation. You will craft reliable distributed systems, integrate...  ...collaborating with platform engineering, SRE, #J-18808-Ljbffr Bank of... 
    Platform

    Bank of America

    Chandler, AZ
    2 days ago
  •  ...Fargo is seeking a Senior Lead Technology Control...  ...most critical technology platforms, infrastructure, and...  ...appetite.Partner with internal audit, compliance, and...  ..., Architecture & Engineering PartnershipCollaborate...  ...databases, OpenShift Container Platform) and financial... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Chandler, AZ
    4 days ago
  •  ...SRE/DevOps Location Options: Chandler, AZ (highly preferred...  ...the ongoing investment in the reliability and modernization of these...  ...Banking Technology division, the Platform team provides management of...  ...reliability. Integration Engineering, ensuring that client is consuming... 
    Platform
    Work experience placement
    3 days per week

    Syntricate Technologies

    Chandler, AZ
    3 days ago
  •  ...seeking a Principal Engineer to support the...  ...goal is to improve reliability, reduce time to...  ...automation, and Site Reliability Engineering (SRE) practices across...  ...data engineering, platform teams, and vendors...  ...with cloud and container platforms (Kubernetes, OpenShift)Desired... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Chandler, AZ
    6 hours ago
  •  ...Fargo is seeking a Lead Systems Operations Engineer in technology as...  ...Modernization Platform Engineer / SRE to join Payments...  ...management, and reliability. You will be embedded...  ...This role embeds Site Reliability Engineering...  ...experience in Kubernetes / container... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Chandler, AZ
    4 days ago
  • $102.5k - $210.6k

     ...Work you’ll do As a Lead Cloud Security Analyst...  ...cybersecurity expertise and cloud engineering execution. This highly...  ...deploy cutting-edge internal and go-to-market...  ..., and/or Google Cloud Platform, including...  ...working knowledge of Docker containers and microservices architecture... 
    Platform
    Full time
    Flexible hours
    Shift work

    Deloitte

    Gilbert, AZ
    3 days ago
  • $170k - $240k

     ...The Role Are you a Platform Engineer looking to make a...  ...plans, internal tooling, architecture...  ...applications in Kubernetes and AWS, and solving...  ...the security and reliability of our applications...  ...experience in DevOps, Site Reliability...  ...understanding of networking, containers, security and... 
    Platform
    Remote work

    GrabJobs

    Gilbert, AZ
    1 day ago
  •  ...Oracle/Java Based applications to .Net/SQL Server and integrate platform applications.Responsibilities- Communicate technical...  ...version control in SVN or TFS- Knowledge of integration to other internal applications with middleware (web) services: including, but not... 
    Platform
    For contractors
    Night shift
    Weekend work

    NTT DATA

    Chandler, AZ
    2 days ago
  • $180.5k - $236.91k

     ...a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering...  ...stack technology platform and a relentless...  ..., Terraform, and Kubernetes with developer...  ...domains such as DevOps, site reliability, and cloud best...  ...deliverablesExperience leading technical... 
    Platform
    Full time
    Work at office
    Flexible hours

    Oscar Health Insurance

    Tempe, AZ
    3 days ago
  • $105.75k - $141k

     ...Purpose Senior Software Engineer is responsible for leading software...  ...data across various platforms. They also mentor junior...  ...of consistency and reliability. Ensure our systems...  ...troubleshooting to internal teams or external...  ...such as Docker or Kubernetes. Proficiency in Git... 
    Platform
    Flexible hours

    CNH

    Tempe, AZ
    3 days ago
  •  ...this role:Wells Fargo is seeking a Senior Lead Analytics Consultant to join the Global...  ...and cross‑functional risk teams to assess internal fraud events, identify emerging risk patterns...  ...with additional analytics tools and platforms is desired to accelerate analytical development... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Chandler, AZ
    4 days ago
  •  ...responsible for delivering the engineering, procurement, and...  ...to complete and finish-out a leading-edge chipmaking facility — a...  ...Automation Lead liaises with internal and external stakeholders to...  ...Bulk Material Utilities), Power Platform solutions, and Bechtel Data Mesh... 
    Platform
    Full time
    Part time
    Work experience placement
    Work at office
    Local area
    Remote work
    Relocation

    Bechtel

    Chandler, AZ
    10 days ago
  •  ...responsible for leading the most...  ...improve the reliability and efficiency...  ...and platform owner, driving...  ...standards, and engineering practices across...  ...for internal consumersFully...  ...region/multi-site high availability...  ...bare metal, container...  ...patterns, enable Kubernetes/operator-based... 
    Platform
    Full time
    Work at office
    Day shift

    Bank of America

    Chandler, AZ
    11 hours ago
  •  ...training, and annual training programs in coordination with site leadership. Maintain accurate training records, track...  ...to-date, accessible training records for regulatory and internal reviews. Utilize LMS platforms to assign, track, and report on training completion and... 
    Platform
    Work at office
    Local area
    Shift work

    hims & hers

    Gilbert, AZ
    5 days ago
  •  ...The Construction Automation Lead is responsible for the day...  ...serves as the primary on-site delivery lead for digital...  ...Requirements Bachelor's degree (or international equivalent) in Engineering, Construction Management,...  ...construction automation platforms and digital workflows in... 
    Platform
    Full time
    For contractors
    Work experience placement
    Work at office
    Local area
    Remote work
    Relocation

    Bechtel Global Corporation

    Chandler, AZ
    2 days ago
  • $56.49 - $66.49 per hour

     ...applications and services are constructed using the Azure and AWS platforms, and helping drive automation & integration aspects around the...  ...experience 8-10 years infrastructure or software engineering / development experience Desired Qualifications Experience... 
    Platform
    Hourly pay
    Contract work
    Temporary work
    Work experience placement
    Local area
    Remote work
    Shift work
    Weekend work
    Chandler, AZ
    28 days ago
  •  ...Our Deloitte AI & Engineering team to transform technology platforms, drive innovation, and help...  ...solutions for client and internal environmentsConfiguring...  ...relationshipsAbility to lead projects or workstreamsAbility...  ...(Docker, Kubernetes, OpenShift, AWS EKS, AWS... 
    Platform
    Local area

    Deloitte

    Gilbert, AZ
    3 days ago
  •  ...Holdings seeks a Director of Cloud & Platform Engineering to lead our cloud platform, owning the Azure...  ...strategy, architecture, operations, and reliability. You will evolve the team from...  ...infrastructure into a modern, service-oriented, SRE-aligned function. This role... 
    Platform

    Jobleads-US

    Tempe, AZ
    4 days ago
  •  ...Pinecone is the leading vector database for...  ...senior software engineer to help design and...  ..., and system reliability. Responsibilities...  ...and build scalable platform components leveraging...  ...Design APIs for internal and external user...  ...tools, such as Kubernetes, cloud-native architectures... 
    Platform
    Local area
    Remote work
    Work from home
    Flexible hours

    GrabJobs

    Chandler, AZ
    3 days ago
  • $16 per hour

    Slot Lead (Internal & GRIC Members Only) Job Category : Slots Requisition Number : SLOTL005109 Posted : August 13, 2026 Full-Time On-site Locations Showing 1 location Description Closing Date: August 21st, 2026 at 4:00 PM Pay Rate: $16.00 per hour This position... 
    Hourly pay
    Full time
    Shift work

    Playatgila

    Chandler, AZ
    4 days ago
  • $119.46k - $155.3k

     ...relationships with clients and internal teams, securing the...  ..., cloud infrastructure, or platform engineering roles Experience in Terraform...  ...containerization using Docker and container orchestration using AWS ECS...  ...scaling, Helm charts, Kubernetes RBAC, and node group management... 
    Platform
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Tempe, AZ
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP). Be the first to apply!