13 Sep
|
Realign
|
Toronto
Toronto, Ontario M5V 3L9 Posted September 11th, 2026
Job Type: Full Time
Job Category: IT
- Job Title: Data Engineer – Kafka / PySpark / Hadoop
- Location: Toronto, ON
- Work Model: Onsite
- Job Type: Full Time (FTE)
- - We are seeking an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop to design, develop, and support scalable batch and real-time data pipelines. The ideal candidate will have strong experience working with large-scale distributed data processing environments and enterprise data integration solutions.
- Key Responsibilities
- - Design, develop, and maintain scalable batch and real-time data pipelines.
- Develop data processing applications using Python and PySpark/Apache Spark.
- Build and support Kafka-based data ingestion and streaming pipelines.
- Work with Hadoop and related technologies to process large volumes of data.
- Develop and maintain ETL/ELT pipelines for data ingestion, transformation, cleansing, and integration.
- Perform data validation, reconciliation, and quality checks.
- Troubleshoot pipeline failures, data discrepancies, and performance issues.
- Optimize Spark/PySpark jobs and SQL queries for performance and scalability.
- Monitor data pipelines and resolve production issues.
- Collaborate with data architects, developers, analysts, and business teams.
- Participate in Agile development, testing, deployment, and production support activities.
- Required Skills
- - Strong hands-on experience with Python for data engineering and automation.
- Strong expertise in PySpark / Apache Spark.
- Hands-on experience with Apache Kafka for real-time data ingestion and streaming.
- Strong experience with the Hadoop ecosystem and distributed data processing.
- Strong SQL skills and experience working with large datasets.
- Experience developing and maintaining ETL/ELT data pipelines.
- Solid understanding of distributed computing and data processing concepts.
- Experience with data ingestion, transformation, cleansing, and integration.
- Strong troubleshooting and performance optimization skills.
- Good to Have
- - Hive
- Databricks
- AWS, Azure, or GCP
- Git and CI/CD
- Unix/Linux
- Airflow or Autosys
- Relational and NoSQL databases
Required Skills
Cloud Developer Data / Python Engineer DevOps Engineer IT Business Continuity Analyst Net Back Engineer Python / DevOps Engineer
📌 Data Engineer (Kafka / PySpark / Hadoop) (Toronto)
🏢 Realign
📍 Toronto