Title: Senior Data Engineer Location: Toronto, ON, CA (Onsite) Duration: Months Job Description We are seeking a highly skilled Senior Data Engineer with deep expertise in Google Cloud Platform (GCP), distributed data processing, and cloud-native data architectures.
This role involves designing, building, optimizing, and maintaining scalable data pipelines and analytical platforms supporting enterprise grade workloads.
The ideal candidate brings strong hands on experience in Big
Query, Dataflow, Data
Proc, Dataform, Cloud Composer (Airflow), PySpark, and end to end ELT/ETL frameworks, along with robust knowledge of metadata, lineage, data quality, and CI/CD automation.
Key Responsibilities .
Data Engineering & Architecture Design and implement end to end data architectures on GCP, including data lakes, data marts, and warehouse models.
Build scalable batch and streaming pipelines using Dataflow, Data
Proc (Spark), Dataform, and Pub/Sub.
Architect low latency, high throughput processing solutions supporting advanced analytics and ML workloads.
Develop pre aggregated models, materialized views, and optimized analytical structures in Big
Query. . ETL/ELT Pipeline Development Design, develop, test, and optimize ELT/ETL pipelines for structured and unstructured data.
Use Dataform and Cloud Composer (Airflow) for orchestration, dependency management, and metadata logging.
Implement best practices for ingestion, transformation, storage, and data access patterns. .
Data Quality, Metadata & Governance Implement enterprise?grade data quality checks using Outstanding Expectations or custom Python frameworks.
Manage metadata, lineage tracking, data cataloging, and compliance with governance standards.
Ensure data integrity, schema enforcement, and security by design principles across all data pipelines. .
Cloud Infrastructure & Dev
Ops Build and automate cloud infrastructure using Terraform, Jenkins, Git
Lab CI, and IaC best practices.
Develop CI/CD workflows for pipeline deployments, testing gates, and operational automation.
Monitor pipelines using Cloud Monitoring & Logging, optimizing for performance and cost. .
Cross Functional Collaboration Work closely with data scientists, analysts, platform engineering, and product owners to translate complex business needs into scalable data solutions.
Support legacy-to-GCP migration initiatives, including Hadoop and on premise workloads.
Enable advanced analytics and ML workloads through ML ready data pipelines. .
Advanced Analytics & ML Support Support feature engineering and ML data preparation for Vertex AI, Gemini, Hugging
Face, or other ML platforms.
Enable vector database workflows and generative AI data pipelines.
Required Technical Skills Cloud & Big Data Google Cloud Platform: Big
Query, Data
Proc, Dataflow, Cloud Composer (Airflow), GCS, Cloud Run, Event
Arc Distributed Computing: Apache Spark, PySpark, Kafka Data Lake & Lakehouse Architectures Programming & Tools Python, SQL, Java Git, Bitbucket, Jenkins, Git
Lab CI Terraform (IaC) REST APIs, FastAPI Airflow DAG development Data Engineering Competencies Data modeling (OLTP/OLAP) Data Warehousing ELT/ETL pipelines Streaming & real time processing Data profiling and validation Metadata, lineage, quality management Experience Required: years (per request) Skills: Cloud & Big Data Engineering GCP Distributed Data Processing ELT/ETL Data Architecture
📌 Senior Data Engineer (Toronto)
🏢 eTeam
📍 Toronto