Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America Corporation

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.

Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.

We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.

Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.

At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Position Summary:

The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:

Reliability & Operations

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Kubernetes & OpenShift Platform Engineering

  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption. 

Observability & Automation

  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.

Capacity & Performance Engineering

  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.

Security & Compliance

  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.

Required Qualifications:

Education / Experience

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Technical Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.

Desired Qualifications

  • BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications.
  • Experience running large-scale enterprise container platforms.
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
  • Experience implementing cloud-native security controls and platform governance.
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
  • Experience with vulnerability management and container security scanning solutions.
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicate effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.


This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices.

Responsibilities:

  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Participates regularly in architecture community of practice meetings and communication via other channels
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations

Skills:

  • Automation
  • Collaboration
  • Influence
  • Production Support
  • Result Orientation
  • Analytical Thinking
  • Application Development
  • Architecture
  • Solution Design
  • Stakeholder Management
  • Adaptability
  • DevOps Practices
  • Project Management
  • Risk Management
  • Solution Delivery Process

Shift:

1st shift (United States of America)

Hours Per Week: 

40

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP) in Charlotte, NC vacancy
  • $125.3k - $167.9k

     ...Join us! Position Summary: The IKCP Site Reliability Engineer Lead is responsible for ensuring the...  ...excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a...  ...cloud, platform engineering, or SRE experience. ~5+ years managing... 
    Platform
    Full time
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America

    Charlotte, NC
    1 day ago
  • $55 - $60 per hour

     ...hiring for a Senior Site Reliability Engineer (SRE) Position Type:...  ...you will: Lead infrastructure provisioning...  ...workloads using Kubernetes (GKE) and Docker,...  ...microservices architectures. Platform Reliability & Cloud...  ...CI/CD pipelines and container platforms using IAM,... 
    Platform
    Hourly pay
    Full time
    Contract work
    Temporary work
    Work experience placement
    Immediate start
    Worldwide
    Flexible hours

    Innova Solutions

    Charlotte, NC
    3 days ago
  • $84.24k - $142.48k

     ...and passionate engineers to deliver...  ...with a team of SRE engineers to operate...  ...managing Kubernetes (EKS), logging...  ...Prometheus), and container technologies (...  ...computing platforms (AWS)Strong knowledge...  ...industry-leading health and welfare...  ...national and international standards,... 
    Platform
    Worldwide
    Flexible hours

    ESRI

    Charlotte, NC
    5 hours ago
  •  ...are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS, container platforms, observability,...  ...will focus on platform reliability, incident management...  ...Python and Bash. Lead incident response, root...  .../SLI governance and site reliability... 
    Platform

    PB consulting

    Matthews, NC
    3 days ago
  •  ...Site Reliability Engineer Our Financial Services client...  ...Reliability Engineer (SRE) to help develop our client's platform operations across...  ...Platform (GCP), container orchestration,...  ...workloads using Kubernetes (GKE) and Docker,...  ...architectures. Lead infrastructure provisioning... 
    Platform
    Weekend work

    Experis

    Charlotte, NC
    17 hours ago
  •  ...Observability engineering (metrics/logs/traces...  ...monitoring). SRE principles (incident...  ...RCA, automation, reliability). Hybrid/SaaS/...  ...systems. SAAS Platform support....  ...managing SAAS-to-SAAS internal to external, and...  ...engineering. Openshift, Kubernetes and containerized... 
    Platform

    Cloud Analytics Technologies LLC

    Charlotte, NC
    6 days ago
  • $122k - $200k

     ...The Senior Software Engineer - Kubernetes Platform Engineering, Service...  ...the bank's enterprise container ecosystem. This role...  ...The individual will lead the architecture and...  ..., networking, SRE, DevOps, and application...  ...automationObservability, reliability, and platform... 
    Platform
    Full time
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America

    Charlotte, NC
    3 days ago
  • $153.84k - $246.15k

     ...believe that belonging leads to better outcomes and...  ...“I can succeed as a Platform Adoption & Enablement...  ...capabilities that thousands of engineers depend on to deliver...  ...-as-code, CI/CD, and container/Kubernetes ecosystems—enough...  ...credibly and set standards; SRE experience is a plus... 
    Platform
    Full time
    Temporary work
    Interim role
    Work at office
    Local area
    Flexible hours

    Capital Group

    Charlotte, NC
    3 days ago
  • $153.84k - $246.15k

     ...We believe that belonging leads to better outcomes and a...  ...ones “I can succeed as a Platform Engineer Lead at Capital Group.”...  ...engineering mindset, treating internal capabilities as software...  ..., load balancing, container networking), Kubernetes, IAM design, TLS/KMS, and... 
    Platform
    Full time
    Temporary work
    Local area
    Flexible hours

    Capital Group

    Charlotte, NC
    3 days ago
  • $167k - $260k

     ...channels, including leadership forums, digital platforms, and global broadcasts at scale. The...  .... The Wells Fargo job profile is Senior Lead Communications Consultant.In this role, you...  ...Qualifications:Experience developing internal and executive communication plans within... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Charlotte, NC
    23 hours ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract...  ...with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally-visible... 
    Platform
    Contract work
    Local area

    InterSources

    Charlotte, NC
    4 days ago
  •  ...Job Title: Lead AI Architect Role Overview...  ...partner closely with our internal architecture team, engineering leads, and business...  .... • Cloud & Platform - hands-on experience...  ...AWS, Azure, or GCP); containers and orchestration (Docker, Kubernetes); serverless services... 
    Platform
    Contract work

    Diverse Lynx

    Charlotte, NC
    2 days ago
  •  ...Senior Site Reliability Engineer Jersey City, New Jersey;Charlotte...  ...maturity, platform resiliency, and operational...  ...workload onboarding Lead complex platform reliability...  ...adoption Mentor SRE engineers and raise...  ...to the provisions contained in the plan documents... 
    Platform
    Work at office
    Shift work
    Day shift

    Bank of America

    Charlotte, NC
    18 hours ago
  • $55 - $62 per hour

     ...Position: Infrastructure Engineer 3 - Contingent...  ...and Deposit platforms, including digital...  ...workloads to the OpenShift Container Platform (OCP)...  ...will bring strong Site Reliability Engineering (SRE) and production support...  ...patterns. ~ Experience leading incident resolution... 
    Platform
    Full time
    Contract work

    Pinnacle Group

    Charlotte, NC
    5 days ago
  •  ...industry-standard oracle platform bringing the capital...  ..., Fidelity International, UBS, S&P Dow Jones...  ...the Role As a Senior Site Reliability Engineer on the CCIP Platform...  ...operated production Kubernetes environments supporting...  ...launches. Experience leading on-call operations,... 
    Platform
    Full time
    Remote work

    Chainlink Labs

    Charlotte, NC
    11 days ago
  •  ...Lead AI Architect Location: Charlotte,...  ...partner closely with the internal architecture team, engineering leads, and business...  .... Cloud & Platform - hands-on experience...  ...AWS, Azure, or GCP); containers and orchestration (Docker, Kubernetes); serverless... 
    Platform
    Contract work

    ReqRoute,Inc

    Charlotte, NC
    5 days ago
  • $153.84k - $246.15k

     ...Platform Engineering And Developer Portal ManagerI...  ...belonging leads to better outcomes...  ...evolve our Internal Developer Portal...  ...platforms such as Kubernetes.You are...  ...across Security, SRE, Architecture,...  ...ability to build reliable, scalable...  ...platforms and container orchestration... 
    Platform
    Temporary work
    Local area
    Flexible hours

    Capital Group

    Charlotte, NC
    4 days ago
  •  ...Observability Lead-Sr. Infrastructure Engineer with a strong...  ...engineering, SRE, platform teams,...  ...improving system reliability, reducing mean...  ...or reusable internal tools that support...  ...in Kubernetes, cloud platforms, containers, microservices...  ...our Benefits site. Depending on... 
    Platform
    Permanent employment
    Full time
    Part time
    Work experience placement
    H1b
    Work at office
    Work visa
    Shift work
    Day shift

    Truist

    Charlotte, NC
    33 minutes ago
  • $101k - $203k

     ...We are the leading provider of professional...  ..., and site reliability engineering (SRE) best practices...  ...on client cloud platforms to include, to...  ...Development of internal solution summaries...  ...Experience with CSP container registries and...  ...services (Kubernetes, Docker, ECR, ECS... 
    Platform
    Full time
    Work experience placement
    Internship
    Local area

    RSM US LLP

    Charlotte, NC
    3 days ago
  •  ...Solutions & Adoption Lead to join our IT team. As...  ...Nichol is Ranked #1 in Engineering News-Record for Marine...  ...Top 50 Designers in International Markets. Moffatt & Nichol...  ..., and automation platforms to support the right solution...  ...designed to cover or contain a comprehensive... 
    Platform
    For contractors
    Work at office
    Worldwide

    Moffatt & Nichol

    Charlotte, NC
    3 days ago
  •  ...review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)...  ...across cloud and application platforms. This is a full-time leadership role...  ...engineering teams toward modern SRE practices.This individual will... 
    Platform
    Full time
    Temporary work
    Shift work
    Day shift

    TIH

    Charlotte, NC
    5 hours ago
  •  ...Fargo is seeking a Senior Lead Technology Control...  ...most critical technology platforms, infrastructure, and...  ...appetite.Partner with internal audit, compliance, and...  ..., Architecture & Engineering PartnershipCollaborate...  ...databases, OpenShift Container Platform) and financial... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Charlotte, NC
    6 days ago
  • $141.2k - $278.3k

     ...Summary Google AI Lead Architect/AI & Engineering:Join our AI &...  ...transforming technology platforms, driving innovation, and...  ...optimize for scalability, reliability, security, and cost.Design...  ...and services, including containers (Docker, Kubernetes), serverless functions,... 
    Platform
    Local area
    Visa sponsorship
    Flexible hours

    Deloitte

    Charlotte, NC
    2 days ago
  • $83.52k - $125.28k

     ...currently seeking a Lead ML Platform Engineer (SRE / FTE / Onsite) to...  ...models efficiently and reliably. The successful...  ...Design and operate Kubernetes-based ML platforms...  ...OpenShift, and associated container, workload...  ...DATA offices or client sites. This ensures we can... 
    Platform
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Charlotte, NC
    13 hours ago
  •  ...role:Wells Fargo is seeking a Lead Systems Architect -...  ...applications, APIs, cloud-native platforms, and AI-enabled...  ...cloud-native architectures, containers, Kubernetes, microservices, and APIsKnowledge...  ...platform teamsPartner with engineering, architecture, risk, and cybersecurity... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Charlotte, NC
    5 hours ago
  • $47.52 - $55.52 per hour

     ...Genesis10 is currently seeking a SRE Lead for a contract position with a Global Financial...  ...for designing and leading the Site Reliability Engineering (SRE) strategy across banking and payments...  ..., on average.) ~ Behavioral Health Platform ~ Medical, Dental, Vision ~... 
    Platform
    Hourly pay
    Permanent employment
    Contract work
    Flexible hours

    Genesis10

    Charlotte, NC
    4 days ago
  •  ...experienced professional to join our team as a Site Reliability Engineer III. This role involves collaborating...  ...DynaTrace. Familiarity with cloud platforms like AWS, Azure, or Google Cloud....  ...technologies (Docker, Kubernetes) and orchestration tools. Knowledge... 
    Platform
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Artech

    Charlotte, NC
    5 days ago
  • $85.1k - $161.7k

     ...We are the leading provider of...  ...serving the cloud engineering needs of our...  ..., and site reliability engineering (SRE) best practices...  ...cloud platforms to include,...  ...Development of internal solution summaries...  ...with CSP container registries and...  ...services (Kubernetes, Docker, ECR... 
    Platform
    Full time
    Work experience placement
    Internship
    Local area

    RSM US LLP

    Charlotte, NC
    3 days ago
  •  ...business applications using leading technologies such as TIBCO,...  ...Experience• Hands-on experience with Kubernetes, Docker, and containerized...  ...workloads on Kubernetes platforms• Experience working in Agile...  ...• Experience with Docker, container image creation, image optimization... 
    Platform
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Charlotte, NC
    3 days ago
  •  ...for performance, reliability, and scalability,...  ...benchmarking & tuning• Kubernetes, GKE, KServe / ML...  ...)Observability & SRE• Prometheus/...  ..., data engineering, and software engineering...  ...Kubernetes-based serving platforms: o KServe,...  ...Provide reusable internal libraries, templates... 
    Platform
    Full time
    Temporary work
    Relocation

    Infosys Technologies

    Charlotte, NC
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP). Be the first to apply!