Lead Data Engineer
Full-time
The Phoenix Group
Role Overview
The Lead Data Engineer will own the design, development, and maintenance of robust data architecture and pipelines leveraging cloud-based technologies, supporting strategic analytics, AI/ML initiatives, and data governance initiatives across the organization’s enterprise-scale environment. Collaborating with cross-functional teams, this individual will ensure high-quality, secure, and performant data solutions that enable decision-making and operational excellence.
Key Responsibilities
- Design, build, test, deploy, and support scalable data pipelines utilizing orchestration tools (Airflow, Astronomer), data transformation frameworks (dbt), cloud data platforms (Snowflake), SQL, and Python.
- Develop and optimize cloud-based data architecture, including database schemas, data warehouses, security roles, workload isolation, cost management, and performance tuning within Snowflake.
- Implement medallion data architecture principles—raw/bronze, silver, and gold layers—for structured and semi-structured data management.
- Lead migration and modernization of legacy workflows to cloud orchestration platforms, establishing reusable patterns for dependency management, retries, backfills, alerting, recovery, and ensuring idempotency.
- Troubleshoot and maintain the full runtime environment of orchestration platforms, including schedulers, workers, executors, metadata stores, queues, and autoscaling components.
- Collaborate with cloud infrastructure teams to manage AWS services such as S3, IAM, EKS, ECS, Fargate, EC2, Lambda, Secrets Manager, KMS, CloudWatch, and networking components.
- Build secure, resilient integrations across APIs, SFTP, databases, cloud storage, vendor APIs, flat files, Excel files, and event-driven systems.
- Develop ingestion and processing pipelines for structured, semi-structured, and unstructured data, including OCR/image-based document extraction.
- Embed data quality testing, monitoring, lineage, reconciliation, and observability into data workflows to ensure reliability and compliance.
- Partner with Data Governance, Security, and Business teams to enforce data ownership, classification, access policies, retention, and auditing standards.
- Build and maintain CI/CD pipelines for orchestration workflows, transformations, and infrastructure code, ensuring automated testing, environment promotion, deployment, and rollback capabilities.
- Support production environments through on-call and incident management processes, conducting root cause analysis, stakeholder communication, and corrective actions.
- Establish reusable frameworks, engineering standards, and best practices to improve platform reliability, consistency, and delivery velocity.
- Lead technical workstreams, mentoring engineering team members through design reviews, code reviews, documentation, and troubleshooting.
Core Qualifications & Requirements
- Minimum of 5+ years in data engineering or software engineering roles with progressive responsibilities.
- Proven experience with cloud-based data platforms, including Snowflake, AWS cloud infrastructure, and data orchestration tools such as Airflow/Astronomer.
- Strong expertise in designing and implementing data architectures, including dimensional and medallion models.
- Hands-on experience migrating and maintaining production workflows in orchestration platforms.
- Deep understanding of cloud compute environments, including Kubernetes, containers, schedulers, workers, and autoscaling architectures.
- Proficiency in AWS services such as S3, IAM, EKS/ECS, Lambda, Secrets Manager, KMS, and CloudWatch for security, storage, and monitoring.
- Skilled in integrating complex data sources and systems, including APIs, SFTP, relational and NoSQL databases, cloud storage, and file-based workflows.
- Experience processing semi-structured and unstructured data, including OCR, document extraction, and image processing.
- Strong background in implementing data quality controls, automated testing, data lineage, and observability tools.
- Knowledge of data governance principles, access controls, data classification, retention policies, and audit requirements.
- Experience with CI/CD pipelines, Git workflows, and automated deployment practices for data solutions and infrastructure.
- Excellent troubleshooting skills across cloud infrastructure, data pipelines, application code, and security permissions.
- Ability to lead technical initiatives and communicate effectively with both technical teams and business stakeholders.
Vacancy posted more than 2 months ago
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Data Engineer. Be the first to apply!
Related searches
