05 Oct
|
Capgemini
|
Mississauga
05 Oct
Capgemini
Mississauga
Job Title: Data Engineer – PySpark & SQL
Location: Mississauga, ON
Fulltime
We are looking for an experienced 3+ years of Data Engineer with strong hands-on expertise in PySpark, Python, SQL, and data engineering . The ideal candidate should have experience building scalable data pipelines and working with large datasets in cloud or distributed data environments.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using PySpark and Python .
- Develop complex and optimized SQL queries for data processing and transformation.
- Build ETL/ELT workflows to ingest, transform, and process large datasets.
- Perform data cleansing, validation, transformation, and performance optimization.
- Work with structured and semi-structured datasets.
- Troubleshoot data quality and pipeline performance issues.
- Collaborate with business, analytics, and engineering teams to understand data requirements.
- Follow data engineering best practices, coding standards, and documentation processes.
Required Skills
- Strong hands-on experience with PySpark / Apache Spark .
- Excellent proficiency in SQL , including complex queries, joins, CTEs, window functions, and query optimization.
- Good programming experience with Python .
- Experience developing ETL/ELT data pipelines .
- Valuable understanding of data warehousing and data modeling concepts.
- Experience working with large-scale/distributed data processing.
- Exposure to at least one major cloud platform such as AWS, Azure, or GCP .
- Strong problem-solving and debugging skills.
Certification Requirement Candidate must hold at least one relevant, active Data Engineering or Cloud certification , preferably from:
- Databricks Certified Generative AI Engineer Associate
- Google Professional Machine Learning Engineer
- NVIDIA Certified Professional - Agentic AI (NCP-AAI)
- AWS Certified AI Practitioner
📌 Data Engineer – PySpark & SQL (Mississauga)
🏢 Capgemini
📍 Mississauga