Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Bank of America ATM
At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.
Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits. We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve. Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs. At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!Position Summary:
The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.
The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.
Key Responsibilities:
Reliability & Operations
- Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
- Lead critical incident response, root cause analysis, and problem management activities.
- Serve as a senior escalation point for L3 platform support and on-call operations.
- Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
- Drive production readiness reviews for new platform capabilities and services.
- Ensure platforms meet enterprise resiliency and availability objectives.
- Conduct resilience exercises and continuous improvement activities following recovery testing.
Kubernetes & OpenShift Platform Engineering
- Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
- Improve platform scalability, performance, and resource utilization across production and non-production environments.
- Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
- Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption.
Observability & Automation
- Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
- Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
- Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
- Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.
Capacity & Performance Engineering
- Perform platform capacity planning and trend analysis.
- Forecast infrastructure growth requirements and optimize platform resource consumption.
- Conduct performance tuning for clusters, workloads, networking, and storage services.
- Support enterprise-scale growth while maintaining platform stability and customer experience.
Security & Compliance
- Partner with security teams to implement platform security controls and governance requirements.
- Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
- Implement and maintain RBAC, Network Policies, and container security controls.
- Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.
Required Qualifications:
Education / Experience
- 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
- 5+ years managing Kubernetes and/or OpenShift production environments.
- Experience operating large-scale mission-critical distributed systems.
- Experience supporting enterprise production environments with 24x7 operational responsibilities.
Technical Skills
- Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
- Linux administration and troubleshooting.
- Terraform, Ansible, GitOps, ArgoCD, Helm
- CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
- Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
- Infrastructure as Code and automation frameworks.
- Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
- Storage platforms, backup technologies, and disaster recovery solutions.
- Scripting in Python, Go, Bash, or similar languages.
Desired Qualifications
- BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
- OpenShift Administration or Kubernetes certifications.
- Experience running large-scale enterprise container platforms.
- Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
- Experience implementing cloud-native security controls and platform governance.
- Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
- Experience with vulnerability management and container security scanning solutions.
- Drives operational excellence and continuous improvement.
- Demonstrates strong ownership and accountability.
- Influences cross-functional teams without direct authority.
- Communicate effectively with senior technical and business leaders.
- Champions automation-first and reliability-first engineering culture.
This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices.
Responsibilities:
- Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
- Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
- Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
- Participates regularly in architecture community of practice meetings and communication via other channels
- Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
- Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations
Skills:
- Automation
- Collaboration
- Influence
- Production Support
- Result Orientation
- Analytical Thinking
- Application Development
- Architecture
- Solution Design
- Stakeholder Management
- Adaptability
- DevOps Practices
- Project Management
- Risk Management
- Solution Delivery Process
Shift:
1st shift (United States of America)Hours Per Week:
40- Job DescriptionWe are seeking an experienced Site Reliability Engineer SRE Lead Messaging Services to drive platform reliability observability and operational excellence... ...for messaging workloadsFamiliarity witho Kubernetes containerized messaging platformsExperience witho...Platform
$152.6k - $191.5k
...leaders across engineering and technology... ...objective reliability goals for services... ...Senior Azure Site Reliability Engineer... ...Azure platform. The role focuses... ....Mentor SRE engineers and... ...with AKS, ACR, Kubernetes, container networking, CI... ...provide industry-leading benefits, access...PlatformFull timeWork at officeDay shift- ...skilled Software Engineer with expert-level... ...design capabilities, Kubernetes development... ...enterprise-ready platform capabilities.This... ...design and build reliable distributed systems... ...platform engineering, SRE, security, DevOps,... ...or contributed to internal developer platforms...PlatformFull timeWork at officeFlexible hoursShift workDay shift
- ...and financial platform for global businesses... ...by world-leading investors... ...responsible for the reliability, performance,... ...: product engineers should be able... ...the Database SRE team, you will... ...of AI-powered internal tooling while... ...(Terraform), Kubernetes, and Docker. Demonstrated...PlatformWorldwide
- ...below: Collaborate with internal and client teams to... ..., TKG and managed Kubernetes services such as AKS,... ...specifically as it relates to containers.• Experience in agile... ...effectively across engineering, infrastructure,... ...Technology|Container Platform|Kubernetes Technical/...PlatformFull timeTemporary workRelocation
$160k - $240k
...work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering organization and help us scale... ...Ansible. CDK a plus. Good understanding of containers, Fargate, Kubernetes, and overall distributed microservice architectures...PlatformPermanent employmentFull timeRemote workWork from homeRelocationFlexible hours$122k - $200k
...The Senior Software Engineer - Kubernetes Platform Engineering, Service... ...the bank's enterprise container ecosystem. This role... ...The individual will lead the architecture and... ..., networking, SRE, DevOps, and application... ...automationObservability, reliability, and platform...PlatformFull timeWork at officeFlexible hoursShift workDay shift$70.24 per hour
...Description* We are seeking an experienced Site Reliability Engineer (SRE) to join a high-performing cloud... ...portfolio of enterprise cloud platforms. This individual will serve as a true... ...TEKsystems Global Services We're a leading provider of business and technology...PlatformPermanent employmentContract workTemporary work$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type... ...focus on ensuring the platform and services customers... ...You will interact with internal staff, managers, and customers... ...on experience of Azure Kubernetes Services (AKS) with container-based deployment skills...PlatformFull timeTemporary workWork experience placementRemote work- ...Chandler, AZ seeks a senior Software Engineer with expert-level Go... ...build scalable backend services, platform APIs, and Kubernetes‑native automation. You will craft reliable distributed systems, integrate... ...collaborating with platform engineering, SRE, #J-18808-Ljbffr Bank of...Platform
- ...Fargo is seeking a Senior Lead Technology Control... ...most critical technology platforms, infrastructure, and... ...appetite.Partner with internal audit, compliance, and... ..., Architecture & Engineering PartnershipCollaborate... ...databases, OpenShift Container Platform) and financial...PlatformFull timeWork experience placement
- ...SRE/DevOps Location Options: Chandler, AZ (highly preferred... ...the ongoing investment in the reliability and modernization of these... ...Banking Technology division, the Platform team provides management of... ...reliability. Integration Engineering, ensuring that client is consuming...PlatformWork experience placement3 days per week
- ...seeking a Principal Engineer to support the... ...goal is to improve reliability, reduce time to... ...automation, and Site Reliability Engineering (SRE) practices across... ...data engineering, platform teams, and vendors... ...with cloud and container platforms (Kubernetes, OpenShift)Desired...PlatformFull timeWork experience placement
- ...Fargo is seeking a Lead Systems Operations Engineer in technology as... ...Modernization Platform Engineer / SRE to join Payments... ...management, and reliability. You will be embedded... ...This role embeds Site Reliability Engineering... ...experience in Kubernetes / container...PlatformFull timeWork experience placement
$102.5k - $210.6k
...Work you’ll do As a Lead Cloud Security Analyst... ...cybersecurity expertise and cloud engineering execution. This highly... ...deploy cutting-edge internal and go-to-market... ..., and/or Google Cloud Platform, including... ...working knowledge of Docker containers and microservices architecture...PlatformFull timeFlexible hoursShift work$170k - $240k
...The Role Are you a Platform Engineer looking to make a... ...plans, internal tooling, architecture... ...applications in Kubernetes and AWS, and solving... ...the security and reliability of our applications... ...experience in DevOps, Site Reliability... ...understanding of networking, containers, security and...PlatformRemote work- ...Oracle/Java Based applications to .Net/SQL Server and integrate platform applications.Responsibilities- Communicate technical... ...version control in SVN or TFS- Knowledge of integration to other internal applications with middleware (web) services: including, but not...PlatformFor contractorsNight shiftWeekend work
$180.5k - $236.91k
...a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering... ...stack technology platform and a relentless... ..., Terraform, and Kubernetes with developer... ...domains such as DevOps, site reliability, and cloud best... ...deliverablesExperience leading technical...PlatformFull timeWork at officeFlexible hours$105.75k - $141k
...Purpose Senior Software Engineer is responsible for leading software... ...data across various platforms. They also mentor junior... ...of consistency and reliability. Ensure our systems... ...troubleshooting to internal teams or external... ...such as Docker or Kubernetes. Proficiency in Git...PlatformFlexible hours- ...this role:Wells Fargo is seeking a Senior Lead Analytics Consultant to join the Global... ...and cross‑functional risk teams to assess internal fraud events, identify emerging risk patterns... ...with additional analytics tools and platforms is desired to accelerate analytical development...PlatformFull timeWork experience placement
- ...responsible for delivering the engineering, procurement, and... ...to complete and finish-out a leading-edge chipmaking facility — a... ...Automation Lead liaises with internal and external stakeholders to... ...Bulk Material Utilities), Power Platform solutions, and Bechtel Data Mesh...PlatformFull timePart timeWork experience placementWork at officeLocal areaRemote workRelocation
- ...responsible for leading the most... ...improve the reliability and efficiency... ...and platform owner, driving... ...standards, and engineering practices across... ...for internal consumersFully... ...region/multi-site high availability... ...bare metal, container... ...patterns, enable Kubernetes/operator-based...PlatformFull timeWork at officeDay shift
- ...training, and annual training programs in coordination with site leadership. Maintain accurate training records, track... ...to-date, accessible training records for regulatory and internal reviews. Utilize LMS platforms to assign, track, and report on training completion and...PlatformWork at officeLocal areaShift work
- ...The Construction Automation Lead is responsible for the day... ...serves as the primary on-site delivery lead for digital... ...Requirements Bachelor's degree (or international equivalent) in Engineering, Construction Management,... ...construction automation platforms and digital workflows in...PlatformFull timeFor contractorsWork experience placementWork at officeLocal areaRemote workRelocation
$56.49 - $66.49 per hour
...applications and services are constructed using the Azure and AWS platforms, and helping drive automation & integration aspects around the... ...experience 8-10 years infrastructure or software engineering / development experience Desired Qualifications Experience...PlatformHourly payContract workTemporary workWork experience placementLocal areaRemote workShift workWeekend work- ...Our Deloitte AI & Engineering team to transform technology platforms, drive innovation, and help... ...solutions for client and internal environmentsConfiguring... ...relationshipsAbility to lead projects or workstreamsAbility... ...(Docker, Kubernetes, OpenShift, AWS EKS, AWS...PlatformLocal area
- ...Holdings seeks a Director of Cloud & Platform Engineering to lead our cloud platform, owning the Azure... ...strategy, architecture, operations, and reliability. You will evolve the team from... ...infrastructure into a modern, service-oriented, SRE-aligned function. This role...Platform
- ...Pinecone is the leading vector database for... ...senior software engineer to help design and... ..., and system reliability. Responsibilities... ...and build scalable platform components leveraging... ...Design APIs for internal and external user... ...tools, such as Kubernetes, cloud-native architectures...PlatformLocal areaRemote workWork from homeFlexible hours
$16 per hour
Slot Lead (Internal & GRIC Members Only) Job Category : Slots Requisition Number : SLOTL005109 Posted : August 13, 2026 Full-Time On-site Locations Showing 1 location Description Closing Date: August 21st, 2026 at 4:00 PM Pay Rate: $16.00 per hour This position...Hourly payFull timeShift work$119.46k - $155.3k
...relationships with clients and internal teams, securing the... ..., cloud infrastructure, or platform engineering roles Experience in Terraform... ...containerization using Docker and container orchestration using AWS ECS... ...scaling, Helm charts, Kubernetes RBAC, and node group management...PlatformFull timeTemporary workRelocation
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP). Be the first to apply!
- site leader Chandler, AZ
- on-site clinical research associate (traveling/remote) Chandler, AZ
- on site coordinator Chandler, AZ
- official site Chandler, AZ
- site recruiter Chandler, AZ
- historic site Chandler, AZ
- IT site lead Chandler, AZ
- junior website developer Chandler, AZ
- site safety Chandler, AZ
- website coordinator Chandler, AZ


