Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America ATM

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.

Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates’ physical, emotional, and financial wellness through affordable, competitive and flexible benefits.

We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.

Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role-specific responsibilities and business needs.

At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Position Summary:

The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:

Reliability & Operations

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Kubernetes & OpenShift Platform Engineering

  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption. 

Observability & Automation

  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.

Capacity & Performance Engineering

  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.

Security & Compliance

  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.

Required Qualifications:

Education / Experience

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Technical Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.

Desired Qualifications

  • BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications.
  • Experience running large-scale enterprise container platforms.
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization.
  • Experience implementing cloud-native security controls and platform governance.
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models.
  • Experience with vulnerability management and container security scanning solutions.
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicate effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.


This job is responsible for partnering with engineering and technology teams to implement measures prescribed by the Site Reliability Engineer teams it leads. Key responsibilities include ensuring appropriate instrumentation, tooling, ticketing, alerting and on call routines are in place for key services, demonstrating technical expertise within domains, and decomposing objectives into work units. Job expectations include advancing efficient solution delivery practices and promoting exceptional design, engineering, and organizational practices.

Responsibilities:

  • Collaborates with Development and Infrastructure teams to understand technical solutions and implement monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)
  • Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and operational needs, and when mentoring SRE resources on reliability practices and established tools/capabilities
  • Partners to implement code changes to make use of common reliability libraries and tools and helps Application Production Services and Application Development teammates understand how to use them
  • Participates regularly in architecture community of practice meetings and communication via other channels
  • Identifies vulnerabilities and opportunities for reliability improvement, such as investigating low level error rates and 'noise' in monitoring, and defines solutions to reduce manual support effort and/or improve system reliability
  • Engages as a subject matter expert in major incident triage efforts and failure scenario modelling and diagnosis with Problem Manager root causes for major incident/problem management investigations

Skills:

  • Automation
  • Collaboration
  • Influence
  • Production Support
  • Result Orientation
  • Analytical Thinking
  • Application Development
  • Architecture
  • Solution Design
  • Stakeholder Management
  • Adaptability
  • DevOps Practices
  • Project Management
  • Risk Management
  • Solution Delivery Process

Shift:

1st shift (United States of America)

Hours Per Week: 

40

Vacancy posted 29 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP) in Charlotte, NC vacancy
  • $80k - $90k

     ...Role: Site Reliability Engineer Location: Charlotte...  ...deliver industry-leading digital...  ...Reliability Engineer (SRE) to improve...  ...technology platforms. The SRE will...  ..., Kubernetes, infrastructure...  ...Systems: Linux Containers: Docker, Kubernetes...  ..., mentoring, internal mobility,... 
    Platform
    Full time
    Temporary work
    Flexible hours

    Synechron

    Charlotte, NC
    2 days ago
  •  ...description:The Site Reliability Engineer role focuses...  ...enterprise platforms across hybrid...  ...Responsibilities include leading major...  ..., mentoring SRE team members,...  ...systems, Kubernetes, automation scripting...  ..., container orchestration...  ..., Linux/Unix internals, and modern infrastructure... 
    Platform
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Charlotte, NC
    2 days ago
  •  ...technology projects, from international trading platforms to critical applications for leading airlines. We recruit...  ...Responsibilities: Platform & Reliability Engineering Embed SRE and production...  ...springboot, mongodb, kakfa, Kubernetes/ CI/CD pipelines Hands... 
    Platform
    Full time
    Worldwide
    Visa sponsorship
    Work visa

    Mthree

    Charlotte, NC
    1 day ago
  •  ...experienced Senior Azure Kubernetes Service (AKS) Platform Engineer to manage, scale, and...  ...Platform Operations & Reliability: Participate in production...  ..., Azure Key Vault, Azure Container Registry (ACR). Scripting...  ...deployment strategies, and SRE/observability best... 
    Platform
    Full time
    Flexible hours

    TMT IT Solutions

    Charlotte, NC
    3 days ago
  • $91.2k - $136.8k

    Reliability Engineer - IE08GEWe’re determined to make a difference...  ...a crucial role to lead infrastructure resilience...  ...Infrastructure Engineering, Site Reliability Engineering (SRE), or DevOps.Hands-on...  ....Expertise in cloud platforms (AWS) and Kubernetes-based microservices... 
    Platform
    Full time
    Temporary work
    Work at office
    3 days per week

    The Hartford Financial Services Group

    Charlotte, NC
    1 day ago
  • $125k - $130k

     ...deliver industry-leading digital solutions....  ...Data, and Software Engineering, servicing an array...  .... Cloud & Platform — hands-on experience...  ..., Azure, or GCP); containers and orchestration (Docker, Kubernetes); serverless services...  ..., mentoring, internal mobility, learning... 
    Platform
    Full time
    Temporary work
    Flexible hours

    Synechron

    Charlotte, NC
    3 days ago
  • $100k - $120k

     ...Details Job Title SRE Operations Lead - AWS, GitLab & AIOps Job...  ...Work Arrangement On site Duration Full-Time...  ...Skills SRE / Application Reliability Engineering (ARE) and 24x7 production...  ...initiatives. Oversee AWS platform operations, batch processing... 
    Platform
    Full time

    SFE

    Charlotte, NC
    1 day ago
  •  ...Job Title: Senior Site Reliability Engineer Duration: 18 months...  ...Top Skills: SRE exp is non-negotiable...  ...critical applications and platforms. This role blends...  ...and reduce toil - lead and complete production...  ...supporting applications on Kubernetes / OpenShift and GCP... 
    Platform
    Shift work
    3 days per week

    Veracity

    Charlotte, NC
    1 day ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract...  ...with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally-visible... 
    Platform
    Contract work
    Local area

    InterSources

    Charlotte, NC
    15 hours ago
  •  ...Senior Site Reliability Engineer Jersey City, New Jersey;Charlotte...  ...maturity, platform resiliency, and operational...  ...Champions modern SRE, observability, and AIOps...  ...platform services. Leads intelligent operations...  ...according to the provisions contained in the plan documents... 
    Platform
    Work at office
    Shift work
    Day shift

    Bank of America

    Charlotte, NC
    1 day ago
  • $127.6k - $191.4k

    Staff Reliability Engineer - IE07KEWe’re determined to...  ...dedicated Staff Data Site Reliability...  ...key partner to the SRE team, concentrating...  ..., and running on platforms like Amazon EMR/...  ...Management (Data Focus): Lead the response and...  ...(compute, containers, databases, APIs... 
    Platform
    Full time
    Temporary work
    Work at office
    3 days per week

    The Hartford Financial Services Group

    Charlotte, NC
    15 hours ago
  •  ...Platform Engineering And Developer Portal ManagerYou...  ...belonging leads to better outcomes...  ...evolve our Internal Developer Portal...  ...such as Kubernetes.You are familiar...  ...across Security, SRE, Architecture,...  ...to build reliable, scalable platforms...  ...platforms and container orchestration... 
    Platform
    Local area
    Flexible hours

    Capital Group

    Charlotte, NC
    15 hours ago
  •  ...Observability Lead-Sr. Infrastructure Engineer with a strong...  ...engineering, SRE, platform teams,...  ...improving system reliability, reducing mean...  ...or reusable internal tools that support...  ...in Kubernetes, cloud platforms, containers, microservices...  ...our Benefits site. Depending on... 
    Platform
    Permanent employment
    Full time
    Part time
    Work experience placement
    H1b
    Work at office
    Work visa
    Shift work
    Day shift

    Truist

    Charlotte, NC
    2 days ago
  • $5,250 per month

     ...Senior Platform Engineer (Multiple Positions)...  ...Engineering, SRE, Product, and...  ...effectiveness and conduct internal studies to...  ...Pillars of Reliability, Security,...  ...with Terraform; containers and container...  ...orchestration (Azure Kubernetes Service [AKS])...  ...is a leading provider of accounts... 
    Platform
    16 hours
    Full time
    Temporary work
    Work at office
    Local area
    Remote work
    Relocation
    Shift work

    Jobleads-US

    Charlotte, NC
    5 days ago
  •  ..., collaborating with leading tech professionals to...  ...an experienced Lead Site Reliability Engineer to join our engineering...  ...* Build and maintain internal tools and services to...  ...* Mentor junior SRE team members and promote...  ...) and orchestration (Kubernetes/ECS) Proficiency with... 
    Full time
    Part time

    Federal Reserve Bank of San Francisco

    Charlotte, NC
    1 day ago
  •  ...are looking for a seasoned SRE who combines technical expertise...  ...operations and approach reliability as an engineering discipline, using automation...  ...~4-8 years of experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or related disciplines... 
    Platform
    Full time

    Synechron

    Charlotte, NC
    2 days ago
  •  ...Position Summary We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our...  ...strategy across cloud and application platforms. This is a full-time leadership...  ...engineering teams toward modern SRE practices. This individual will... 
    Platform
    Full time
    Shift work

    CRC Group

    Charlotte, NC
    15 hours ago
  •  ...Solutions & Adoption Lead to join our IT team. As...  ...Lead works directly with internal clients across the...  ...including individual engineers and scientists, offices...  ...improve productivity.Platform Integration and Ecosystem...  ...designed to cover or contain a comprehensive listing... 
    Platform
    For contractors
    Work at office
    Worldwide

    Moffatt & Nichol

    Charlotte, NC
    4 days ago
  •  ...Site Reliability Engineer (SRE, Terraform, AWS, Dynatrace) Optomi, in partnership with a Fortune 500 digital platform leader, is seeking a Site Reliability Engineer to join their Digital SRE...  ...applications ~ Proven ability to lead incident response efforts and create... 
    Platform

    LinkedIn

    Charlotte, NC
    1 day ago
  • $120k - $140k

     ...Role: AI Lead Architect – RAG & Python Location...  ...DevOps, Data, and Software Engineering, servicing an array of...  ..., Python, LLMs, cloud platforms, and enterprise...  ...Implement Docker, Kubernetes, CI/CD, MLOps, and LLMOps...  ...arrangements, mentoring, internal mobility, learning and... 
    Platform
    Full time
    Temporary work
    Flexible hours

    Synechron

    Charlotte, NC
    2 days ago
  •  ...with system analysts, engineers, and programmers to design...  ...and OpenSearch, and container orchestration using...  ...applications and BPM platforms. Preferred SkillsWorking...  ...role, base salary of internal peers, prior performance...  ...today weʼre a market-leading retirement company fueled... 
    Platform
    Full time
    Work at office

    Teachers Insurance and Annuity Association

    Charlotte, NC
    15 hours ago
  •  ...operating VPI's cloud platform and engineering infrastructure....  ...to enable highly reliable and scalable...  ...infrastructure automation, Kubernetes, DevOps practices...  ...platform.Lead AWS platform architecture...  ...Kubernetes and container platforms...  ...monitoring solutions, and SRE practices.... 
    Platform
    Full time

    Vanguard

    Charlotte, NC
    2 days ago
  • $141.2k - $278.3k

    Position Summary Lead Architect with...  ...AI Experience, AI & Engineering/Engineering as a ServiceAgentic...  ...technology platforms, driving innovation,...  ...services including containers (Docker, Kubernetes), serverless functions...  ...the performance, reliability and usability of AI... 
    Platform
    Work at office
    Local area
    Visa sponsorship
    Flexible hours
    Shift work

    Deloitte

    Charlotte, NC
    2 days ago
  • $160k - $225k

    DescriptionThe Lead Architect is a hands...  ...client and internal initiatives. The...  ...paths, establishing engineering guardrails, and leading...  ...on Google Cloud Platform (GCP).· 5+ years...  ...(Docker, Kubernetes)· Strong understanding...  ...long-term security, reliability, and cost-... 
    Platform
    Permanent employment
    Full time
    Temporary work
    Work experience placement
    Remote work

    TEKsystems

    Charlotte, NC
    1 day ago
  •  ...experienced professional to join our team as a Site Reliability Engineer III. This role involves collaborating...  .... Familiarity with cloud platforms like AWS, Azure, or Google Cloud....  ...containerization technologies (Docker, Kubernetes) and orchestration tools.... 
    Platform
    Work experience placement
    Immediate start
    Remote work
    Flexible hours

    Artech

    Charlotte, NC
    1 day ago
  • $119.77k - $140.9k

     ...API ecosystem, leading engineering efforts across...  ...GraphQL platforms.Design and deliver...  ...GitLab CI/CD, Kubernetes, Helm, Ansible...  ...Improve platform reliability through...  ...Kubernetes, Helm, containers, and Azure services...  ..., and Site Reliability Engineering (SRE).Proven ability... 
    Platform
    Full time
    Work experience placement
    Local area
    3 days per week

    US Bank

    Charlotte, NC
    4 days ago
  •  ...seeking a Principal Engineer, Platform, Cloud...  ...infrastructure SRE, Kubernetes/OpenShift, observability...  ..., and platform reliability. This role drives...  ...security, and container platforms. As a...  ...years of experience leading enterprise-scale...  ...Engineering, Site Reliability Engineering... 
    Platform
    Full time
    Work experience placement

    Wells Fargo

    Charlotte, NC
    4 days ago
  •  ...professional services company with leading capabilities in digital,...  ...at . A successful Network Engineer Architect combines deep...  ...Familiarity with DevOps, NetOps, SRE, and platform engineering operating models...  ...of virtualization, container networking, edge computing,... 
    Platform
    Full time
    Work experience placement
    Live in
    Work at office
    Local area
    Remote work

    Accenture

    Charlotte, NC
    3 days ago
  • $108.8k - $136k

     ...single, secure platform. By...  ...and proudly leading the next generation...  ...software engineering department....  ...faster and more reliably. You will...  ...and finetune container orchestration...  ...Develop internal tooling and...  ...Engineer, or Site Reliability...  ...orchestration ( Kubernetes/EKS ). ~ Solid... 
    Platform
    Local area
    Flexible hours

    Capital Rx

    Charlotte, NC
    4 days ago
  •  ...seeking a OpenShift DevOps Engineer to join our team in...  ..., OCP(Openshift Container platform), Kubernetes, Helm, Infrastructure...  ...GCP)- Function as an SRE to enhance platform stability...  ...one of the world's leading AI and digital...  ...DATA offices or client sites. This ensures we can... 
    Platform
    Work experience placement
    Work at office
    Remote work
    Flexible hours
    Weekend work

    NTT DATA

    Charlotte, NC
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP). Be the first to apply!