Lead Platform Reliability Engineer
Wells Fargo Bank
Wells Fargo is seeking a Lead Platform Reliability Engineer to join the CTO Platform organization. This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database, or Storage) and have demonstrated experience collaborating across at least one additional infrastructure domains (Enterprise Tools, Cloud, Observability Tools). The expectation is that this engineer will elevate themselves in looking for trends and patterns that are not limited to these streams and investigate systemic issues that span multiple streams or need a deeper troubleshooting.As part of our Platform Reliability Engineering (PRE) team, you will apply modern Site Reliability Engineering (SRE) practices to improve the availability, resiliency, observability, scalability, and operational excellence of critical enterprise platforms. You will leverage your domain expertise to identify systemic issues, drive automation, and deliver engineering solutions that strengthen platform stability at scale.In This Role You WillServe as the reliability engineering expert for your primary domain (Network, Middleware, Database, or Storage) while partnering across adjacent technology disciplinesLead the investigation and resolution of complex production incidents, identifying root causes and implementing long-term corrective actionsApply SRE principles including service level indicators (SLIs), service level objectives (SLOs), error budgets, and reliability engineering practices to improve platform healthLead capacity analysis, forecasting, and utilization reviews to identify future scaling risks and prevent service degradation before customer impact occursPerform deep performance analysis across infrastructure layers, identifying bottlenecks, contention points, latency drivers, and resource inefficienciesIdentify and remediate configuration drift, operational debt, and platform hygiene issues that impact long-term reliabilityDrive proactive reliability improvements through observability, automation, performance optimization, and resiliency engineeringDesign and implement automation solutions that eliminate operational toil, reduce manual intervention, and improve recovery capabilitiesDefine and enhance enterprise observability standards through metrics, logging, tracing, alerting, and service health monitoringPartner closely with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availabilityLead blameless post-incident reviews and convert recurring operational issues into measurable engineering improvementsIdentify reliability risks and communicate technical recommendations to engineering leaders and senior stakeholdersMentor engineers and technical teams on reliability engineering, operational excellence, automation, and platform best practicesRequired Qualifications5+ years of Systems Engineering, Infrastructure Engineering, Platform Engineering, Technology Architecture, or equivalent experience demonstrated through work experience, military experience, training, or education5+ years supporting and engineering enterprise-scale production environments5+ years of experience with hands-on expertise in one of the following technology domains:Network Engineering (routing, switching, load balancing, DNS, network observability, performance analysis)Middleware Engineering (WebSphere, Tomcat, JBoss, Kafka, MQ, application platforms, integration technologies)Database Engineering (Oracle, SQL Server, PostgreSQL, MongoDB, database performance, replication, HA/DR)Storage Engineering (SAN/NAS technologies, storage virtualization, backup/recovery, performance and capacity management)Desired QualificationsStrong experience applying SRE principles, including SLI/SLO development, error budgets, incident analysis, and reliability measurementExperience supporting highly available, mission-critical production environmentsProven success troubleshooting complex issues spanning multiple technology domains in large-scale distributed environmentsExperience with capacity planning, resiliency engineering, fault tolerance, disaster recovery, and performance optimizationHands-on experience with observability and monitoring platforms such as Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, or similar technologiesExperience building dashboards, alerts, service health indicators, and operational reportingStrong automation and scripting experience using Python, Bash, PowerShell, or similar technologiesExperience developing operational tooling, API integrations, self-healing capabilities, and automated remediation solutionsFamiliarity with Git-based development practices, CI/CD pipelines, infrastructure automation, and Infrastructure as Code tools such as Ansible or TerraformExperience diagnosing and resolving issues that span multiple infrastructure layersAbility to influence technical direction across infrastructure and engineering organizationsExperience leading major incident reviews and driving sustainable operational improvementsDemonstrated success mentoring engineers and promoting reliability engineering best practicesStrong communication skills with the ability to translate technical concepts into business-focused outcomesJob Expectations:This position offers a hybrid scheduleThis position does not offer Visa sponsorshipPay RangeReflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities.$119,000.00 - $224,000.00Benefits Wells Fargo provides eligible employees with a comprehensive set of benefits, many of which are listed below. Visit Benefits - Wells Fargo Jobs for an overview of the following benefit plans and programs offered to employees.Health benefits401(k) PlanPaid time offDisability benefitsLife insurance, critical illness insurance, and accident insuranceParental leaveCritical caregiving leaveDiscounts and savingsCommuter benefitsTuition reimbursementScholarships for dependent childrenAdoption reimbursementPosting End Date:1 Sep 2026*Job posting may come down early due to volume of applicants.We Value Equal OpportunityWells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit’s risk appetite and all risk and compliance program requirements.Applicants with DisabilitiesTo request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo.Drug and Alcohol PolicyWells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.Wells Fargo Recruitment and Hiring Requirements:a. Third-Party recordings are prohibited unless authorized by Wells Fargo.b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.SummaryLocation: IRVING, TX; WEST DES MOINES, IA; CHARLOTTE, NC; MINNEAPOLIS, MN; ISELIN, NJType: Full time
- ...an impact. Join us! Position Summary: The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability,... ...excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform...SuggestedWork at officeFlexible hoursShift workDay shift
- ...Hiring Alert | DevSecOps & SRE AI Platform Engineering Location: Charlotte, NC (Onsite) Employment Type: Full-Time Experience... ...Architecture & Roadmap Planning Production On-Call & Reliability Engineering Terraform Module Design, Drift Detection &...SuggestedFull time
$153.84k - $246.15k
...designed to celebrate differences. We believe that belonging leads to better outcomes and a stronger community of associates united... ...existing skills and learn new ones “I can succeed as a Platform Engineer Lead at Capital Group.” As a Platform Engineer Lead on our...SuggestedFull timeTemporary workLocal areaFlexible hours$153.84k - $246.15k
...differences. We believe that belonging leads to better outcomes and a stronger community... ...learn new ones “I can succeed as a Platform Adoption & Enablement Lead at Capital Group... ...and the capabilities that thousands of engineers depend on to deliver business-critical...SuggestedFull timeTemporary workInterim roleWork at officeLocal areaFlexible hours- ...We are seeking an experienced Lead Alteryx Administrator / BI Platform Engineer with 10+ years of experience administering, deploying, and supporting enterprise Business Intelligence and Analytics platforms. The ideal candidate will have deep expertise in Alteryx Server...SuggestedTemporary work
- ...Hiring Alert | DevSecOps & SRE AI/ML Engineer AI Platform Location: Charlotte, NC (Onsite) Employment Type: Full-Time Experience Required: 10 13 Years Visa Type: USC / GC / GC EAD Only Must-Have Skills: Python FastAPI, Pydantic & boto3...Full time
- ...Responsibilities:Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and... ...and reliability standards.Lead blameless post‑incident reviews... ...participate in reliability engineering and resilience communities of practice...Full time
$152.6k - $191.5k
...partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities... ..., observability maturity, platform resiliency, and operational... ...benefits eligible. We provide industry-leading benefits, access to paid time...Full timeWork at officeDay shift$84.24k - $142.48k
...talented team of dynamic and passionate engineers to deliver capabilities that enable our... ...understanding and experience with cloud computing platforms (AWS)Strong knowledge of Linux Operating... ...rewards strategy includes industry-leading health and welfare benefits: medical,...WorldwideFlexible hours- ...United States of America)Please review the following job description:The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical...Permanent employmentFull timePart timeH1bWork at officeLocal areaImmediate startWork visaMonday to FridayShift workDay shift
$55 - $62 per hour
...POSTINGPosition: Infrastructure Engineer 3 - ContingentLocation:... ...Business Lending and Deposit platforms, including digital banking applications... ...You will bring strong Site Reliability Engineering (SRE) and... ...patterns.Experience leading incident resolution, root cause...Full time- ...America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a... ...and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for...Full timeTemporary workShift workDay shift
- ...around the world.At Electrolux Group, a leading global appliance company, we strive... ...Charlotte, NC.About the Role:The Site Reliability Engineer plays a critical role in designing, building... ...best practices to ensure optimal platform performance and resilience.Working closely...Full timeFlexible hours
$91.2k - $136.8k
Reliability Engineer - IE08GEWe’re determined to make a difference and are proud to be an insurance... ...This position will play a crucial role to lead infrastructure resilience in ensuring... ...best practices.Expertise in cloud platforms (AWS) and Kubernetes-based microservices...Full timeTemporary workWork at office3 days per week$152.8k - $229.2k
Principal Reliability Engineering - IE06JEWe’re determined to make a difference and are proud to be... ...availability, and performance of all data platforms, cloud infrastructure, data products,... ...Engineering within EDS and leads the definition, implementation, and continuous...Full timeTemporary workWork at office3 days per week- Job Description : Set up CI/CD pipelines and Helm charts for deployments. Implement GitOps for declarative infrastructure. Integrate observability tools for logs, metrics, and traces. Skills: Infrastructure: OpenShift/Kubernetes, Helm/GitOps, Secure pipelines...
- ...seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS, container platforms, observability, infrastructure... ...role will focus on platform reliability, incident management, SLO/SLI... ...scripts using Python and Bash. Lead incident response, root cause...
- ...Financial Services client is seeking a Site Reliability Engineer to join their team. Title: Site... ...(SRE) to help develop our client's platform operations across Windows, Linux, and cloud... ...microservices architectures. Lead infrastructure provisioning and configuration...Weekend work
- Job TitleJoin us to work collaboratively with our talented team of dynamic and passionate engineers to deliver capabilities that enable our customers to make a difference. You'll deploy and operate ArcGIS Velocity and ArcGIS Workflow Manager SaaS solution.
- ...Job Title: Senior Site Reliability Engineer Duration: 18 months (possibility to extend or convert... ...for mission-critical applications and platforms. This role blends software engineering... ...opportunities and reduce toil - lead and complete production readiness activities...Shift work3 days per week
$124.36k - $146.3k
...Bank, Elavon is committed to building the platforms and ecosystems that help over 1.5... ...DescriptionResponsibilitiesAs a senior-level Reliability Engineer specializing in observability, this... ...and improve overall service reliability.Lead the definition, documentation, implementation...Full timeWork experience placementLocal area3 days per week- ...Reliability Test Engineer (Manufacturing)Location: Greensboro / Raleigh / Charlotte, NCDuration: FulltimeJob Description:Skills Desired:4+ years... ...with relevant test standards.FCC/EMC test setups, environmental stress screening, and accelerated life test platforms....
- ...PFB the JD must have skills Hands on SRE Engineer with good analytical skills Good Exposure to both incident and Problem Management... ...(Servers,Load Balancer/Trace Logs) and have the acumen to lead any troubleshooting efforts from end to end until permanent...Permanent employmentFull time
- ...Mandatory Skills: Observability engineering (metrics/logs/traces, tooling like. Grafana... ...principles (incident response, RCA, automation, reliability). Hybrid/SaaS/on-prem integration... ...flows across multiple systems. SAAS Platform support. Experience managing SAAS-to...
- ...Description SERC OVERVIEWSERC Reliability Corporation (SERC) is one of... ...(RAPA) Department. The team leads the development of steady-state... ....The Senior Reliability Engineer is a senior-level technical and... ...PSSE, TARA, SERVM, or similar platforms preferred.Proven ability to work...Temporary workWork at office2 days per week
$90 per hour
Site Reliability Engineering (SRE) Consultant This range is provided by TekWissen ®. Your actual pay... ...services and consulting company and is a leading provider of information technology,... ...leadership. Deep expertise in cloud platforms (AWS, Azure), observability tools (Dynatrace...Contract workTemporary work$127.6k - $191.4k
Staff Reliability Engineer - IE07KEWe’re determined to make a difference and are proud to be an insurance... ..., Python/Pyspark, and running on platforms like Amazon EMR/Hadoop, Informatica and... ...Incident and Problem Management (Data Focus): Lead the response and resolution for data-...Full timeTemporary workWork at office3 days per week$17 - $27.75 per hour
...deliver an exceptional customer experience * Serves as a Brand Ambassador embodying of Coach values and increasing brand awareness * Leads implementation of Company initiatives and support full operation of the business * Maintain a growth mindset for business and...Minimum wageShift work$56 - $66 per hour
DescriptionA client with Kforce is seeking a Remote Appian Platform Engineer II to join their team.Summary:As a Platform Engineer within the client's Technology team, this person will provide critical support for Appian Applications that support Finance and Accounting function...Remote work$125k - $142k
Job Description:AssetMark is a leading strategic provider of innovative investment and... ..., and purpose.We are seeking a Platform Engineer to join our platform engineering team,... ...while ensuring security, performance, and reliability across all AssetMark systems.We can only...Full timeWork at officeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Platform Reliability Engineer. Be the first to apply!
- lead infrastructure engineer Charlotte, NC
- lead operating engineer Charlotte, NC
- lead engineer Charlotte, NC
- lead network engineer Charlotte, NC
- platform engineer Charlotte, NC
- client platform engineer Charlotte, NC
- platform developer Charlotte, NC
- senior platform engineer Charlotte, NC
- data platform engineer Charlotte, NC
- reliability maintenance engineering technician Charlotte, NC


