Lead Data Engineer
Zodiac Solutions
Role: Lead Data Engineer
Location: Cary, NC (Onsite)
Job Type: Fulltime
About The Engagement:
Client is building a centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform ingests 150+ inbound data feeds, distributes to 35+ downstream systems, and is organized as a medallion architecture (Bronze / Silver / Gold). AI is embedded in ingestion, canonical mapping, data quality, reconciliation, and business user access from day one — not bolted on at the edges.
This is a senior hands-on leadership role. You will own the end-to-end technical design of the data and AI layers, build the reference implementations your engineers work from, and ship production-grade Python, Scala, and PySpark code every week. If you have not written or reviewed production code in the past year, this is not the right fit.
WHAT YOU WILL OWN:
Data Platform Architecture & Engineering
- Own the end-to-end lakehouse architecture: Bronze / Silver / Gold layer contracts, zone layout on ADLS Gen2, Delta Lake table design, partitioning, schema evolution, and retention strategy.
- Design and build metadata-driven, parameterized ingestion frameworks that onboard new data sources without bespoke pipeline code for every feed.
- Write canonical PySpark and Scala Spark transformation jobs that serve as the team reference; set coding and testing standards, review pull requests, and debug production incidents.
- Design for scale and cost: tune Spark clusters and pools, apply partition pruning and caching strategies, and set cost guardrails as data volumes grow.
- Build and automate CI/CD for Databricks and ADF pipelines in Azure DevOps using Databricks Asset Bundles and Terraform; maintain platform observability with Azure Monitor and Log Analytics.
- Design repeatable patterns for batch files, database extracts, CDC feeds, and streaming ingestion using Azure Event Hubs / Kafka and Spark Structured Streaming.
AI-Augmented Ingestion & Canonical Mapping:
- Design and build the AI-augmented metadata ingestion framework that auto-generates bridge documents, DML statements, canonical table definitions, and control metadata from source schemas.
- Build AI-assisted source-to-canonical attribute mapping: schema reasoning, data profiling, confidence scoring, and a human review / approval gate before any mapping reaches production.
- Generate file-level and record-level validation rules from historical data and metadata analysis, feeding a configurable, rules-engine-backed data quality framework.
AI-Driven Data Quality, Anomaly Detection & Testing
- Build AI-assisted data quality that analyses patterns across Bronze, Silver, and Gold layers to propose DQ rules beyond predefined checks.
- Deliver anomaly detection covering outliers, data drift, schema drift, volume shifts, and reconciliation breaks — with actionable alerting, not noise.
- Build AI-assisted automated reconciliation and test-data generation to feed the platform's automated testing framework.
- Produce synthetic, privacy-preserving datasets for lower environments using differential privacy, format-preserving masking, and referential-integrity-safe generation.
Semantic Layer & Conversational Data Access
- Design and build an ontology-driven semantic layer and knowledge graph modelling relationships across finance data entities, powering data discovery, semantic integration, and AI/BI tooling.
- Build a GPT-powered conversational interface for natural-language querying of financial data: text-to-SQL or semantic-layer-mediated retrieval grounded in the knowledge graph, with row-level and column-level security enforced and every answer traceable to source.
Governance, Architecture Reviews & Team Leadership:
- Own end-to-end AI architecture decisions: model selection, RAG and retrieval design, prompt strategy, evaluation harnesses, guardrails, cost and latency budgets, and observability.
- Implement Unity Catalog for cataloguing, lineage, and fine-grained access control; define PII classification, masking, tokenization, and encryption standards across every layer.
- Take designs through Architecture Review Boards and AI governance forums, covering responsible AI, data residency, model approval, auditability, and human-in-the-loop controls.
- Mentor data engineers, run design reviews, and set the engineering patterns the team builds on — without becoming a bottleneck.
- Produce documentation and reusable components good enough for the client's team to operate the platform independently at engagement end.
MUST-HAVE SKILLS & EXPERIENCE:
Programming & Data Engineering
- Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large-scale batch and streaming data pipelines using Delta Lake and Spark technologies.
- Strong SQL and data modelling — dimensional and normalised; schema design and data contract definition.
- Databricks expertise — Delta Lake, Unity Catalog, Jobs & Workflows, cluster and pool management, performance tuning, Model Serving.
- Azure data stack — ADLS Gen2 (zone design, ACLs, lifecycle), Azure Data Factory (parameterized / metadata-driven frameworks, error handling), Azure Event Hubs.
AI & Machine Learning
- 3+ years designing and shipping LLM-based systems in production: RAG pipelines, agentic / tool-calling workflows, structured output, chunking and embedding strategy, vector and hybrid retrieval, and prompt engineering.
- Evaluation discipline — golden datasets, regression suites, accuracy and hallucination tracking, human-in-the-loop feedback loops; you measure AI quality, not assert it.
- Hands-on experience with LangChain, LlamaIndex, or LangGraph, plus at least one provider stack (Azure OpenAI, OpenAI, or Databricks Model Serving).
- Metadata-driven thinking — schema inference, data profiling, lineage, catalogs, and configuration-driven frameworks that onboard the next source without new code.
Architecture & Governance:
- 12–18 years of total experience in data engineering, data platform delivery, or related disciplines.
- Proven delivery of a medallion / lakehouse architecture at enterprise scale — not just familiarity with the concept.
- Azure security and governance — Entra ID, managed identities, RBAC, POSIX ACLs on ADLS Gen2, Key Vault, private endpoints, and PII handling.
- CI/CD and infrastructure as code — Azure DevOps, Terraform, Databricks Asset Bundles, and automated testing of data pipelines.
- Clear technical writing and the ability to present and defend a design to both engineers and non-technical stakeholders.
STRONGLY PREFERRED:
- Knowledge graphs and ontologies: RDF/SPARQL, property graphs (Neo4j), or graph modelling over a lakehouse.
- Text-to-SQL or semantic-layer-backed natural-language query systems at enterprise scale, including access control and ambiguity handling.
- ML-based anomaly detection on time-series or transactional financial data.
- Financial services or insurance domain knowledge: finance close, general ledger, subledger, reconciliation, or actuarial data.
- LLMOps and MLOps: model versioning, prompt versioning, cost governance, and observability tooling.
- Databricks Data Engineer Professional, Azure DP-203 / DP-700, or AZ-305 certification.
- dbt, Great Expectations, or similar data-quality and transformation tooling.
- Workday, Prism, or Accounting Center exposure.
WHAT MAKES SOMEONE SUCCESSFUL HERE
- You prototype in days, not sprints — and the prototype is production-close enough to survive an architecture review.
- You know where AI genuinely helps and where a deterministic rule is the better engineering answer. On finance data, that judgement matters more than enthusiasm.
- You design for human review by default. Every AI-generated mapping, rule, and artefact lands in front of a reviewer with the reasoning attached.
- You are comfortable working with US-based client stakeholders and can explain a technical trade-off to a finance business owner without jargon.
- You leave behind documentation and patterns the client's own team can operate without you.
- ...READ CAREFULLY BEFORE APPLYING | URGENT HIRING ROLE: Lead Data Engineer – Hands-On Client: MetLife Location: Cary, NC — On-site / Hybrid Experience: 12–18 Years Employment Type: Full-Time Eligibility: US CITIZENS ONLY ABOUT THE ENGAGEMENT...SuggestedFull time
- ...Programming & Data Engineering Expert - level proficiency in Python, Scala, and PySpark, with a strong track record of designing and delivering production-ready, modular, and well-tested solutions; developing and troubleshooting Spark workloads; and optimizing large...SuggestedFull timeContract work
- ...ABOUT THE ENGAGEMENT A centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks... ...data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code...Suggested
- ...A centralized, AI-first enterprise Data Hub for a global insurance and financial services client on Azure Databricks. The platform... ...the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code...SuggestedShift work
$128.94k - $168k
...Lead Data EngineerThis is a remote role that may be hired in several markets across the United States. Priority given to those candidates... ..., Texas and Florida.We are seeking an experienced Lead Data Engineer to join our Data Platform Engineering Organization and lead...SuggestedRemote work- We partner with global enterprises to design and scale data platforms that power products used by millions of users.... ...mentoring while remaining hands-on. Your Role as a Tech Lead As a Tech Lead, Data Engineering, you will own the technical direction of data initiatives...Full timeRemote workMonday to FridayFlexible hours
- ...First Citizens Bank is seeking a Lead Data Engineer to design, build, and operate production-grade data pipelines in a cloud environment. You will own end-to-end delivery, set engineering standards, and mentor junior engineers to raise the team’s capabilities. The...
- ...This is a remote role that may only be hired in the following location(s): North Carolina, Arizona & Texas. We are seeking a Lead Data Engineer with deep, hands-on experience designing, building, and operating production-grade data platforms. This role is targeted at...Remote work
- ...About the Engagement Build a centralized, AI-first enterprise data hub for a global insurance and financial services client on... ...the data and AI layers, build reference implementations for the engineering team, and ship production-grade Python, Scala, and PySpark code...Full timeShift work
- ...Summary: The engineer in this position should have a wealth of experience engineering/installing... ...include the complete engineering of data centers including, but not limited to,... ...field engineers, installation supervisors, lead installers, and installation techs....For contractorsWork at officeLocal area
- ...SQL Data EngineerThis is for BECU Business Intelligence Data Engineer. For the SQL Data Engineer role: experience with Snowflake, SQL Server migration, ADF, Power BI semantic models, Clienture DevOps/GitHub, Cortex, AIM, and Apache Iceberg. Must have skills: experience...
- ...Sr. Data EngineerCGI is seeking an experienced Senior Data Engineer with a strong database development background to design, develop, and modernize enterprise data solutions. The ideal candidate will bring expertise in Oracle, ETL technologies, cloud platforms, and modern...Full time
- ...Palantir Data EngineerMandatory Skills: Data Engineer with experience on Palantir Palantir Data Engineer We are seeking a highly skilled Palantir Data Engineer to design, develop, and optimize data solutions using Palantir Foundry, ensuring seamless data integration,...
$95k - $154k
...helped thousands of candidates secure full-time roles with leading companies and recognizable brands—think Google, Apple,... ...programmer, Java full stack developer, Python/Java developer, data analyst, data engineer, data scientist, and machine learning/AI engineer. In...Full timeH1bRemote workNight shift$95k - $125k
...future of technology forward.The OpportunityThe Senior Engineer - Infrastructure Operations, Data Security is responsible for engineering, automating,... ...its subsidiaries and affiliates, is one of the world’s leading financial services companies; providing insurance, annuities...Full timeTemporary workWork at officeLocal areaWorldwideRelocation package3 days per week- ...Data EngineerCGI is seeking an experienced Data Engineer with a strong database development background to design, develop, and modernize enterprise data solutions. The ideal candidate will bring expertise in Oracle, ETL technologies, cloud platforms, and modern data engineering...Full time
- ...building award-winning games or crafting engine technology that enables others to make visually... ...game development.ANALYTICSWhat We DoOur Data & Analytics teams build powerful stories... ...sabbatical.ABOUT USEpic Games is a leading interactive entertainment company. For over...
- ...are about our mission.Why Join Q2?Q2 is a leading provider of digital banking and lending... ...grow stronger relationships. As a Senior Data Scientist on the team, you will work closely... ...and user value.A Typical Day:Explore and engineer features from large, complex commercial...Full timeBank staffWork visa
- ...The Role The Electrical Shift Engineer provides support for the multiple critical... ...processes, procedures, and notifies/advises Lead engineers, Engineering Manager and Director... ...Universal EPA certification is required. Data Center Engineering and supervision...Full timeShift workAfternoon shift
- ...services and energy industries , and we're looking for a Senior Data Scientist to implement cutting-edge data science solutions,... ...Collaborate with cross-functional teams, including product management, engineering, and business stakeholders, to translate business requirements...Contract work
$90k - $115k
...At MetLife, data isn’t just a tool – it is a catalyst for growth. As part of our Data... ...data – one that’s governed responsibly, engineered for scalability, and designed for growth.... ...subsidiaries and affiliates, is one of theworld’s leading financial services companies; providing...Temporary workInternshipWork at officeLocal area3 days per week- ...push boundaries, elevate standards, and deliver with purpose. As a senior subject matter expert within our Data Center team, you'll lead structural engineering efforts across a portfolio of data center projects, including new builds and retrofits. As part of the vision...Contract workFor contractorsRemote work
$145k - $165k
...Job Description We're looking for a hands-on data engineer to own, maintain and expand the data infrastructure behind our Siding & Accessories... ...problem-solving and operational ownership. ~ Experience leading or contributing to a data platform build or migration. ~...Full time$113k - $140k
...Company Description We are Olsson. We engineer and design solutions that improve the world... ...to work for. We design large hyperscale data center campuses and colocation data... ...as a project manager on some projects and lead design engineer on others. Prepare planning...Full timeWork at officeRemote workFlexible hours- ...Role Descriptions: Design and build batch and streaming data pipelines using Azure, Databricks, and Kafka. Develop ETL/ELT solutions... ...Skills & Experience: ~8+ years of experience in data engineering in azure using Azure Databricks. ~ Strong experience with Azure...Full time
- ...Work Location: Hartford, CT, Raleigh, NC, Richardson, TX Job Title: Data Engineer Type-Full Time Your contribution to the team: ⦁ A collaborative spirit and excellent communication skills. ⦁ The ability to handle end to end SDLC phases from requirement gathering...Full time
- ...The Infosys Data and Analytics (DNA) unit is at the forefront of transforming data into actionable insights, driving business growth and operational efficiency. We specialize in leveraging advanced AI and analytics to create innovative solutions that address complex business...Full timeTemporary workRelocation
- ...179Location: Raleigh, NC, United StatesJob Title: Data EngineerLocation: Raleigh, NCJob Summary - Data Engineer:We are seeking a skilled mid-level+ Data Engineer... ...that do not meet ingestion requirements, which can lead to load failures or data backouts.Collaboration: Work...
- ...specialize in leveraging advanced technologies such as AI, cloud, and data-led innovation to help our clients accelerate growth and unlock... ...Carolina, TexasCompanyITL USA Interest GroupInfosys Limited Job RoleTechnology Consultant 2Career RoleTechnology Lead - US: 153693BRFull timeTemporary workRelocation
$128k - $249k
...to join us. Perhaps our search for talented visionaries and your search for important and impactful work lead to the same place.We are seeking a Senior Data Engineer to join the firm. The Senior Data Engineer contributes to the Enterprise Applications, Data, and AI...Full timeTemporary workLocal areaRemote workFlexible hoursAfternoon shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Data Engineer. Be the first to apply!





