Job Title: Data Engineer – Kafka / PySpark / Hadoop
Location: Toronto, ON
Work Model: Onsite
Employment Type: Full-Time (FTE)
Experience: 10+ Years
Job Summary
"We are looking for an experienced Data Engineer with strong hands-on expertise in Kafka, PySpark, Python, and Hadoop. The ideal candidate will have experience designing, developing, and supporting scalable batch and real-time data pipelines and working with large-scale distributed data processing environments.
Must-Have Skills:
- Strong hands-on experience with Python for data engineering and automation
- Strong experience with PySpark / Apache Spark
- Hands-on experience with Apache Kafka for real-time data ingestion and streaming
- Robust experience with Hadoop ecosystem and distributed data processing
- Strong SQL skills and experience working with large datasets
- Experience developing and maintaining ETL/ELT data pipelines
- Experience with data ingestion, transformation, cleansing, and integration
- Strong understanding of distributed computing and data processing concepts
Key Responsibilities
- Design, develop, and maintain scalable batch and real-time data pipelines
- Develop data processing applications using Python and PySpark
- Build and support Kafka-based data ingestion and streaming pipelines
- Work with Hadoop and related technologies to process large volumes of data
- Perform data transformation, validation, and integration
- Troubleshoot data pipeline failures, data discrepancies, and performance issues
- Analyze and optimize Spark/PySpark jobs and SQL queries
- Monitor data pipelines and resolve production issues
- Collaborate with data architects, developers, analysts, and business teams
- Participate in Agile development, testing, deployment, and production support