18 Aug
|
KData AI
|
Winnipeg
We are looking for a
Senior, Super Hands-On Databricks Data Engineer
who lives and breathes code, query optimization, and modern data architecture. In this role, you won't just design architectures on whiteboards—you will write production PySpark/SQL, optimize Databricks clusters, build streaming and batch pipelines, and enforce data governance. You will own end-to-end pipeline execution from raw ingestion to curated Gold layer models, playing a lead role in modernizing our Lakehouse platform. Key Responsibilities
1. Hands-On Pipeline Development & Lakehouse Architecture
Design, build, and maintain enterprise-scale
batch and real-time streaming pipelines
using
PySpark, SQL, Delta Live Tables (DLT), and Auto Loader
. Implement and refine
Medallion Architecture (Bronze Silver Gold)
to support downstream BI, reporting, and Machine Learning workloads. Enforce schema evolution, ACID transactions, and data compaction using
Delta Lake core constructs
. 2. Performance Tuning & Optimization (Deep Tech)
Diagnose and resolve Spark performance bottlenecks:
data skew, OOM errors, excessive shufflings, and memory spills
. Optimize queries using
Liquid Clustering, Z-Ordering, Data Partitioning, AQE (Adaptive Query Execution), and Photon engine tuning
. Benchmark and optimize Databricks compute workloads to minimize
DBU (Databricks Unit) consumption and cloud costs (FinOps)
. 3. Governance, Security & Quality
Implement end-to-end data governance, fine-grained access control (row/column-level security), and lineage tracking using
Unity Catalog
. Automate automated data quality validation checks and alert mechanisms across the pipeline life cycle. 4. Operations, CI/CD & DevOps
Automate pipeline orchestration using
Databricks Asset Bundles (DABs)
or
Databricks Workflows / Apache Airflow
. Build CI/CD pipelines (GitHub Actions, Azure DevOps, or GitLab) for automated testing, deployment, and code promotions. Required Skills & Qualifications
Must-Haves
Experience:
8+ years in Data Engineering
, with
4+ years of intensive, hands-on production experience on Databricks
. Programming Mastery:
Fluent in
PySpark, Advanced SQL
, and Python. Databricks Ecosystem:
Deep experience with
Delta Lake, Unity Catalog, Delta Live Tables (DLT), Auto Loader, and Databricks Workflows
. Cloud Infrastructure:
Solid hands-on experience in at least one primary cloud provider ( AWS, Azure, or GCP
) integration with Databricks (S3/ADLS Gen2, IAM, Key Vaults/Secret Manager). Data Modeling:
Solid understanding of dimensional modeling (Kimball), One Big Table (OBT) strategies, and data vault patterns. CI/CD & Software Engineering:
Proficient in Git workflows, unit testing PySpark code (pytest), and deployment automation. Preferred / Nice-to-Haves
Certifications:
Databricks Certified Data Engineer Professional. Streaming:
Hands-on with Apache Kafka, Event Hubs, or Kinesis integration via Structured Streaming. GenAI / ML Ops:
Familiarity with MLflow, Feature Store, or Vector Search within Databricks. Infrastructure as Code (IaC):
Experience using Terraform to provision Databricks workspaces and storage resources. Performance Indicators (How success is measured)
Pipeline Reliability:
Maintaining strict SLA thresholds on critical Gold-layer models. Cost Efficiency:
Measurable reduction in DBU costs through effective compute profiling and tuning. Code Quality:
High test coverage and zero-downtime CI/CD deployments.
#J-18808-Ljbffr
📌 Senior Databricks Data Engineer (Winnipeg)
🏢 KData AI
📍 Winnipeg