Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Data Engineer - Kafka / PySpark / Hadoop

Full-time

Acestack

  • Job Title: Data Engineer Kafka / PySpark / Hadoop
  • Location: Toronto, ON
  • Work Model: Onsite
  • Job Type: Full Time (FTE)
  • Job Description

    We are seeking an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop to design, develop, and support scalable batch and real-time data pipelines. The ideal candidate will have strong experience working with large-scale distributed data processing environments and enterprise data integration solutions.

    Key Responsibilities
    • Design, develop, and maintain scalable batch and real-time data pipelines .
    • Develop data processing applications using Python and PySpark/Apache Spark .
    • Build and support Kafka-based data ingestion and streaming pipelines .
    • Work with Hadoop and related technologies to process large volumes of data.
    • Develop and maintain ETL/ELT pipelines for data ingestion, transformation, cleansing, and integration.
    • Perform data validation, reconciliation, and quality checks.
    • Troubleshoot pipeline failures, data discrepancies, and performance issues.
    • Optimize Spark/PySpark jobs and SQL queries for performance and scalability.
    • Monitor data pipelines and resolve production issues.
    • Collaborate with data architects, developers, analysts, and business teams.
    • Participate in Agile development, testing, deployment, and production support activities.
    Required Skills
    • Strong hands-on experience with Python for data engineering and automation.
    • Strong expertise in PySpark / Apache Spark .
    • Hands-on experience with Apache Kafka for real-time data ingestion and streaming.
    • Strong experience with the Hadoop ecosystem and distributed data processing.
    • Strong SQL skills and experience working with large datasets.
    • Experience developing and maintaining ETL/ELT data pipelines .
    • Strong understanding of distributed computing and data processing concepts.
    • Experience with data ingestion, transformation, cleansing, and integration.
    • Strong troubleshooting and performance optimization skills.
    Good to Have
    • Hive
    • Databricks
    • AWS, Azure, or GCP
    • Git and CI/CD
    • Unix/Linux
    • Airflow or Autosys
    • Relational and NoSQL databases
Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Data Engineer - Kafka / PySpark / Hadoop. Be the first to apply!