Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Principal Reliability Engineer - EDS

$152.8k - $229.2k

The Hartford Financial Services Group

Principal Reliability Engineering - IE06JEWe’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals – and to help others accomplish theirs, too. Join our team as we help shape the future. The Enterprise Data Services (EDS) organization is seeking a Principal Reliability Engineer (Principal RE) to serve as the senior technical authority responsible for the reliability, resilience, availability, and performance of all data platforms, cloud infrastructure, data products, and data pipelines across the enterprise data organization. This role sets the strategic vision for Reliability Engineering within EDS and leads the definition, implementation, and continuous evolution of RE practices, tooling, automation, observability frameworks, and AIOps/AI‑driven operations.As the Principal RE, you will influence architectural direction, lead large‑scale, cross‑organizational technical initiatives, and drive a culture of engineering excellence, automation‑first operations, and proactive reliability improvement. You will partner closely with platform engineering, data engineering, security, architecture, and product teams to embed RE principles into every stage of the data product lifecycle.This role will have a Hybrid work schedule, with the expectation of working in an office (Columbus, OH, Chicago, IL, Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).​Key ResponsibilitiesEnterprise Reliability Strategy & LeadershipWork closely with the AVP, RE & Production Support, EDS defining the Reliability Engineering strategy for data platforms, data cloud environments, and data products.Establish long‑term RE roadmaps, target operating models, and architectural patterns that scale with organizational growth.Serve as the highest‑level technical escalation point for systemic reliability issues, influencing executive stakeholders and engineering leaders.Platform & Cloud Reliability (AWS, GCP, Snowflake, EMR, Hadoop, ETL/ELT)Leverage Enterprise provided standards and building blocks to Architect and evolve highly reliable, performant, and cost‑efficient cloud‑based platforms across AWS and GCP for all EDS services.Influence and work directly with Platform Solution Architecture on new product enablement, hyper automation (end to end blueprint automation).Oversee reliability controls and fail‑safe patterns for Snowflake, EMR, Hadoop/Spark clusters, container platforms (e.g., Kubernetes), and mission‑critical data systems.Lead the creation and enforcement of SLO/SLI frameworks that span the entire data lifecycle.AI‑Enabled Operations, AIOps & Intelligent AutomationDevelop and implement AI‑driven automation for anomaly detection, alert correlation, autonomous remediation, and predictive capacity management.Leverage LLMs, prompt engineering, and cloud‑native AI services (AWS Bedrock, SageMaker, Vertex AI) to build intelligent runbooks, advanced troubleshooting agents, and generative‑AI‑enabled operational tooling.Champion the adoption of machine learning–based observability and reliability analytics.End‑to‑End Observability & Operational ExcellenceAdopt and architect enterprise‑wide data observability frameworks—including logging, metrics, tracing, distributed profiling, and event pipelines—for all data platforms and pipelines.Establish gold‑standard incident response patterns, post‑incident reviews, and continuous improvement processes.Drive elimination of toil across EDS, focusing on self‑healing systems, proactive detection, and autonomous operations.Data Pipeline & Data Product ReliabilityDefine RE best practices for modern data products, governed data pipelines, real‑time/streaming systems, and operational analytics platforms.Ensure data quality, data timeliness, and SLAs for data products through automated checks, lineage-informed alerting, and pipeline reliability tooling.Partner with Data Engineering to embed resilience patterns (idempotency, checkpointing, replayability, disaster recovery) into pipeline architectures.Engineering Standards, Governance & Cross‑Org InfluenceSet and enforce standards for IaC, CI/CD, platform automation, reliability frameworks, operational readiness, and runbook quality across EDS.Provide technical leadership and mentorship to Staff/Senior Engineers in the RE team and Production Support teams, influencing engineering culture and helping grow RE capabilities across the organization.Represent Reliability Engineering in architectural reviews, enterprise governance forums, and executive‑level discussions.Technical Experience10+ years in one or more of the following areas: data, cloud, platform engineering, site/reliability engineering, or large‑scale distributed systems, with experience in leadership or technology leader roles.Proficiency with data or cloud platforms, including architectural patterns for resilience, networking, security, and distributed data infrastructure.Deep experience supporting or engineering platforms such as Snowflake, EMR, Hadoop/Spark, Data Integration, and cloud‑native data ecosystems.Scripting and programming (preferably Python) for large‑scale automation, platform tooling, and reliability frameworks.Experience with Infrastructure‑as‑Code (Terraform, CloudFormation) and enterprise CI/CD.Preferred QualificationsExperience in regulated or highly complex enterprise environments (financial services, insurance, healthcare).Prior experience as a Senior Staff Engineer, Engineering or Architecture leader with hands on experience, or similar senior technical role.Knowledge of data governance, metadata, lineage systems, and data quality engineering practices.Certifications in AWS, GCP, Kubernetes, or SRE/DevOps frameworks.AI & AIOpsBackground applying machine learning to operations—anomaly detection, event correlation, predictive modeling, and automated remediation.Understand of AI‑enabled developer/operations tools using LLMs, prompt engineering, or cloud AI services for reliability improvements.Observability & Platform OperationsExpertise with enterprise observability stacks (Prometheus, Grafana, Datadog, Splunk, Dynatrace, OpenTelemetry).Ability to design and enforce advanced SLI/SLO frameworks across complex data ecosystems.Leadership & Cross‑Functional InfluenceDemonstrated ability to lead technical strategy at scale, influence senior engineering leaders, and set enterprise‑wide standards.Strong capability in mentoring engineers, providing architectural guidance, and fostering engineering excellence.Exceptional communication skills for interacting with executives, senior architects, product leaders, and engineering teams.Candidate must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.CompensationThe listed annualized base pay range is primarily based on analysis of similar positions in the external market. Actual base pay could vary and may be above or below the listed range based on factors including but not limited to performance, proficiency and demonstration of competencies required for the role. The base pay is just one component of The Hartford’s total compensation package for employees. Other rewards may include short-term or annual bonuses, long-term incentives, and on-the-spot recognition. The annualized base pay range for this role is:$152,800 - $229,200Equal Opportunity Employer/Sex/Race/Color/Veterans/Disability/Sexual Orientation/Gender Identity or Expression/Religion/AgeAbout Us | Our Culture | What It’s Like to Work Here | Perks & BenefitsSummaryLocation: Hartford, CT; Charlotte, NC; United States - RemoteType: Full time

Vacancy posted 9 hours ago
Similar jobs that could be interesting for youBased on the Principal Reliability Engineer - EDS in Charlotte, NC vacancy
  • $91.2k - $136.8k

    Reliability Engineer - IE08GEWe’re determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals - and to help others accomplish theirs, too. Join... 
    Suggested
    Full time
    Temporary work
    Work at office
    3 days per week

    The Hartford Financial Services Group

    Charlotte, NC
    1 day ago
  •  ...is the world's leading company delivering sustainable design, engineering, and consultancy solutions for natural and built assets.We are...  ...Role description:Arcadis is currently seeking a highly motivated Principal Substation Electrical Engineer to join our Power Delivery &... 
    Suggested
    Full time
    Part time
    For contractors

    Arcadis

    Charlotte, NC
    3 days ago
  • $165k - $210k

     ...Principal Mechanical EngineerLocation: Charlotte, NCType: Direct HireCompensation: $165,000.00 - $210,000.00Work Model: Onsite – onsiteResponsibilities...  ...vendor documentation• Develop scopes of work, job notes, and engineering deliverables• Conduct field walkdowns and support... 
    Suggested

    System One Holdings, LLC

    Charlotte, NC
    2 days ago
  •  ...Electrical Engineering ServicesYou will provide electrical engineering services to Worley and its customers, including technical support and supervision within the electrical engineering team.Responsibilities:Deliver electrical engineering services that meet Worley, its... 
    Suggested
    Local area

    Worley

    Charlotte, NC
    2 days ago
  •  ...Principal Electrical EngineerSystem One is seeking a Principal Electrical Engineer for a Direct Hire opportunity located in Charlotte, NC.ResponsibilitiesLead electrical engineering efforts for combined cycle power generation projects from conceptual design through commissioning... 
    Suggested

    System One Holdings, LLC

    Charlotte, NC
    2 days ago
  •  ...Senior Principal Electrical Engineer Aspen Technical Staffing is seeking an experienced Senior Principal Electrical Engineer for a contract-to-hire opportunity with our client. This is a leadership position responsible for providing technical expertise, project leadership... 
    Contract work
    Work at office

    Aspen Technical Staffing

    Charlotte, NC
    2 days ago
  • $156.5k - $230k

     ...with opportunities to learn, grow, and make an impact. Join us!Job Description:This job is responsible for defining and leading the engineering approach for solutions at the program or portfolio level, to deliver significant business outcomes. Key responsibilities include... 
    Full time
    Work experience placement
    Work at office
    Flexible hours
    Shift work
    Day shift

    Bank of America

    Charlotte, NC
    14 hours ago
  • This position will be fully remote and can be hired anywhere in the continental U.S. Optiv is seeking a Principal SailPoint Engineer to join Optiv Security’s 24x7x365 Security Operations Center as a member of the Advanced Fusion Center (AFC) team. The Optiv AFC line of... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    Optiv

    Charlotte, NC
    3 days ago
  • $147k - $237.5k

     ...your work truly matters.Job SummaryThe TeamEngineering - Our engineering team is at the core of our products and connected directly to...  ...only enabled by a secure digital environment.Job SummaryAs a Principal Escalation Engineer, you will hold a highly visible, senior-level... 
    Full time
    Remote work
    Visa sponsorship
    Work visa

    Palo Alto Networks

    Charlotte, NC
    4 days ago
  • $95.2k - $176.8k

     ...Component Prep & Compounding, Liquid PFS Filling, Automated Inspection, Autoinjector Assembly, Packaging/Finished Products). The Principal MES Engineer is being hired to participate in Greenfield Project execution and subsequently support the facility after going live. Focus... 
    Full time
    Work at office
    Local area

    Genentech

    Charlotte, NC
    1 day ago
  • $156.5k - $230k

     ...with opportunities to learn, grow, and make an impact. Join us!Job Description:This job is responsible for defining and leading the engineering approach for solutions at the program or portfolio level, to deliver significant business outcomes. Key responsibilities include... 
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, NC
    14 hours ago
  • $117.5k - $234.5k

     ...Carrier social media at @Carrier. About This Role As a Principal Systems Engineer you will lead the design, modeling, and architectural...  ...architecture, ensuring all designs meet performance and reliability requirements through rigorous testing and DOE... 
    Temporary work
    Local area

    Carrier

    Charlotte, NC
    4 days ago
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Full time

    Vanguard

    Charlotte, NC
    4 days ago
  • $152.6k - $191.5k

     ...opportunities to learn, grow, and make an impact. Join us!This job is responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include composing observability designs through instrumentation... 
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, NC
    4 days ago
  • $84.24k - $142.48k

    OverviewJoin us to work collaboratively with our talented team of dynamic and passionate engineers to deliver capabilities that enable our customers to make a difference. You'll deploy and operate ArcGIS Velocity and ArcGIS Workflow Manager SaaS solutions. You will also... 
    Worldwide
    Flexible hours

    ESRI

    Charlotte, NC
    1 day ago
  •  ...(Required)Work Shift:1st Shift (United States of America)Please review the following job description:Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish... 
    Full time
    Temporary work
    Shift work
    Day shift

    TIH

    Charlotte, NC
    1 day ago
  •  ...related.Applicants must have 2 years’ experience with:1. Software Reliability Engineering2. Supporting Business Applications - Dynamics AX 4...  ...business production processes4. Leading a team of skilled engineers remotely,5. IT Software/Systems architectures6. Microsoft Ecosystem... 
    Full time
    Work at office
    Remote work
    Worldwide
    Monday to Friday
    2 days per week
    3 days per week

    Asahi Kasei

    Charlotte, NC
    14 hours ago
  •  ...!Where you’ll be:This position will be based at our Corporate Headquarters located in Charlotte, NC.About the Role:The Site Reliability Engineer plays a critical role in designing, building, and maintaining scalable, secure, and highly available cloud infrastructure that... 
    Full time
    Flexible hours

    Electrolux

    Charlotte, NC
    3 days ago
  •  ...Principal EngineerIRALOGIX is a high-growth, institutional technology platform focused on providing uniquely capable solutions to IRA...  ...competitiveness, far beyond industry expectations.As a PRINCIPAL ENGINEER, reporting to senior engineering leadership, you will be one... 
    Permanent employment

    iraLogix

    Charlotte, NC
    2 days ago
  •  ...Agentic Ai Framework & Harness Principal EngineerAt Bank of America, we are guided by a...  ...responsible for defining and leading the engineering approach for solutions at the program...  ...contracts, and execution patterns that enable reliable integration with enterprise systems,... 
    Temporary work
    Work at office
    Local area
    Flexible hours

    Bank of America

    Charlotte, NC
    2 days ago
  •  ...Principal Engineerin TechnologyWells Fargo is seeking a principal engineer responsible for the design, development, and implementation of reusable Terraform modules that enable observability capabilities across Azure, AWS, and GCP. Serves as the technical lead for cloud... 
    Work experience placement

    Wells Fargo

    Charlotte, NC
    2 days ago
  •  ...Fluency: English (Required)Work Shift:1st shift (United States of America)Please review the following job description:The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and... 
    Permanent employment
    Full time
    Part time
    H1b
    Work at office
    Local area
    Immediate start
    Work visa
    Monday to Friday
    Shift work
    Day shift

    Truist

    Charlotte, NC
    2 days ago
  •  ...dashboards, and status reports for transparency and accountability. The role requires 7+ years in tech project management, strong Agile/Jira skills, exceptional written communication for execs, and experience with Site Reliability Engineering/DevOps. #J-18808-Ljbffr... 

    ManpowerGroup Global, Inc.

    Charlotte, NC
    3 days ago
  •  ...Matlen Silver is seeking a Senior Platform Engineer to ensure the reliability, performance, and long-term sustainability of the ServiceNow platform that underpins the Service-Aware Operating Model. You will treat the platform as a product, improving standards, runtime... 

    Matlen Silver

    Charlotte, NC
    1 day ago
  •  ...Details Job Description: Mandatory Skills: Observability engineering (metrics/logs/traces, tooling like. Grafana/Splunk/OTel/...  ...monitoring). SRE principles (incident response, RCA, automation, reliability). Hybrid/SaaS/on-prem integration monitoring. Vendor-... 

    Cloud Analytics Technologies LLC

    Charlotte, NC
    2 days ago
  • About this role:Wells Fargo is seeking a Principal Engineer to lead the strategy, architecture and delivery of secure network solutions in...  ...engineers while driving measurable improvements in network reliability, service resilience, operational maturity, and the overall... 
    Full time
    Work experience placement
    Free visa

    Wells Fargo

    Charlotte, NC
    3 days ago
  •  ...Site Reliability Engineer Job Location: San Francisco, CA or Charlotte, NC. Job Type: Contract Work with local API development squads, platform teams, product owners, scrum masters, and architects. The SRE ensures that both our internally critical and our externally... 
    Contract work
    Local area

    InterSources

    Charlotte, NC
    14 hours ago
  •  ...Site Reliability Engineer III Our client, a leading organization in the technology and infrastructure sector, is seeking a Site Reliability Engineer III to join their team. As a Site Reliability Engineer III, you will be part of the Infrastructure Support team supporting... 
    Flexible hours

    Experis

    Charlotte, NC
    2 days ago
  • Job Description :  Set up CI/CD pipelines and Helm charts for deployments. Implement GitOps for declarative infrastructure. Integrate observability tools for logs, metrics, and traces. Skills: Infrastructure: OpenShift/Kubernetes, Helm/GitOps, Secure pipelines...

    Cloud Analytics Technologies LLC

    Charlotte, NC
    more than 2 months ago
  • Wells Fargo is seeking a Lead Platform Reliability Engineer to join the CTO Platform organization. This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database,... 
    Full time
    Work experience placement
    Free visa

    Wells Fargo

    Charlotte, NC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Principal Reliability Engineer - EDS. Be the first to apply!