Develop and maintain scalable data pipelines using PySpark and Spark SQL for processing large datasets efficiently. Write clean, reusable, and optimized code in Python for data manipulation, analysis, and automation tasks. Design and implement ETL workflows to extract, transform, and load data from various structured and unstructured sources. Collaborate with data engineers, analysts, and stakeholders to understand data requirements and deliver solutions.
Optimize
Spark jobs for performance tuning, resource utilization, and minimizing execution time.
Work with distributed computing frameworks to process and analyze big data in a cloud or on-premises setting.
Utilize
Spark SQL for querying and managing large datasets stored in distributed systems like Hadoop or cloud storage. Monitor and troubleshoot data pipeline issues, ensuring reliability and timely delivery of data. Stay updated with the latest advancements in PySpark, Spark SQL, and big data technologies to improve existing systems.
EEOC Compliance
We are an equal chance employer, and all qualified applicants will receive consideration for employment.
DISCLAIMER
AI Usage Policy: Pacer Group uses AI to assist in screening applications. Final hiring decisions are made by human recruiters based on qualifications and experience.
📌 Big Data Lead Mississauga
🏢 Pacer Group
📍 Mississauga
Reply to this offer
Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.