17 Aug
|
KData AI
|
Ontario
We are looking for a Senior, Super Hands-On Databricks Data Engineer who lives and breathes code, query optimization, and modern data architecture. In this role, you won't just design architectures on whiteboards—you will write production PySpark/SQL, optimize Databricks clusters, build streaming and batch pipelines, and enforce data governance.
You will own end-to-end pipeline execution from raw ingestion to curated Gold layer models, playing a lead role in modernizing our Lakehouse platform.
Key Responsibilities
1. Hands-On Pipeline Development & Lakehouse Architecture
Design, build, and maintain enterprise-scale batch and real-time streaming pipelines using PySpark, SQL, Delta Live Tables (DLT), and Auto Loader .
Implement and refine Medallion Architecture (Bronze Silver Gold) to support downstream BI, reporting, and Machine Learning workloads.
Enforce schema evolution, ACID transactions, and data compaction using Delta Lake core constructs .
2. Performance Tuning & Optimization (Deep Tech)
Diagnose and resolve Spark performance bottlenecks: data skew, OOM errors, excessive shufflings, and memory spills .
Optimize queries using Liquid Clustering, Z-Ordering, Data Partitioning, AQE (Adaptive Query Execution), and Photon engine tuning .
Benchmark and optimize Databricks compute workloads to minimize DBU (Databricks Unit) consumption and cloud costs (FinOps) .
3. Governance, Security & Quality
Implement end-to-end data governance, fine-grained access control (row/column-level security), and lineage tracking using Unity Catalog .
Automate automated data quality validation checks and alert mechanisms across the pipeline life cycle.
4. Operations, CI/CD & DevOps
Automate pipeline orchestration using Databricks Asset Bundles (DABs) or Databricks Workflows / Apache Airflow .
Build CI/CD pipelines (GitHub Actions, Azure DevOps, or GitLab) for automated testing, deployment, and code promotions.
Required Skills & Qualifications
Must-Haves
Experience: 8+ years in Data Engineering , with 4+ years of intensive, hands-on production experience on Databricks .
Programming Mastery: Fluent in PySpark, Advanced SQL , and Python.
Databricks Ecosystem: Deep experience with Delta Lake, Unity Catalog, Delta Live Tables (DLT), Auto Loader, and Databricks Workflows .
Cloud Infrastructure: Strong hands-on experience in at least one primary cloud provider ( AWS, Azure, or GCP ) integration with Databricks (S3/ADLS Gen2, IAM, Key Vaults/Secret Manager).
Data Modeling: Solid understanding of dimensional modeling (Kimball), One Big Table (OBT) strategies, and data vault patterns.
CI/CD & Software Engineering: Proficient in Git workflows, unit testing PySpark code (pytest), and deployment automation.
Preferred / Nice-to-Haves
Certifications: Databricks Certified Data Engineer Qualified.
Streaming: Hands-on with Apache Kafka, Event Hubs, or Kinesis integration via Structured Streaming.
GenAI / ML Ops: Familiarity with MLflow, Feature Store, or Vector Search within Databricks.
Infrastructure as Code (IaC): Experience using Terraform to provision Databricks workspaces and storage resources.
Performance Indicators (How success is measured)
Pipeline Reliability: Maintaining strict SLA thresholds on critical Gold-layer models.
Cost Efficiency: Measurable reduction in DBU costs through effective compute profiling and tuning.
Code Quality: High test coverage and zero-downtime CI/CD deployments.
#J-18808-Ljbffr
📌 Senior Databricks Data Engineer (Ontario)
🏢 KData AI
📍 Ontario