11 Sep
|
Appnovation
|
Toronto
11 Sep
Appnovation
Toronto
- As a Senior Data Engineer, AI &
- Agents, you will prepare enterprise domain data for consumption by AI agents, partnering directly with business stakeholders and subject-matter experts
- The role centers on building and registering domain-grounded data agents over governed datasets, performing lakehouse migrations, and onboarding new data domains — including commercial, finance, research and development, and real-world data — onto the enterprise data platform
- You will operate across Databricks and Snowflake, converting source tables to open formats and generating the catalogue metadata that powers downstream automation and data-contract workflows, all while ensuring high data quality, governed access, and reliable agent-based access to trusted data
- The ideal candidate brings deep, hands-on data engineering expertise, strong governance instincts, and excellent stakeholder-facing skills
- Build and register domain data agents at scale over governed tables across Databricks and Snowflake
- Perform lakehouse migrations, including converting source tables to open formats (e.g., Apache Iceberg) to enable agent-based access
- Generate and curate catalogue metadata that feeds downstream automation and data-contract workflows
- Partner with business stakeholders and subject-matter experts through iterative build, test, and validation cycles- Advanced SQL together with Spark / PySpark, and experience with pipeline orchestration (dbt, Apache Airflow, or Databricks Workflows)
- 5+ years of qualified experience in data engineering, with significant hands-on experience across modern data warehousing and lakehouse platforms (Databricks and Snowflake preferred)
- Strong data engineering background with genuine, hands-on fluency across both Databricks and Snowflake
- Bachelor’s or Master’s degree in Computer Science, Information Systems, Engineering,
or a related field
- Demonstrated experience building data agents or query interfaces over governed datasets (e.g., Snowflake Cortex or Genie)
- Experience implementing data quality, observability, and lineage, and applying governance controls such as masking and row- and column-level security
- Excellent stakeholder-facing skills, with a track record of translating business requirements into delivered data assets
- Pipelines and modelling: SQL, PySpark, dbt, Airflow / Databricks Workflows
- Governance and catalogue: Unity Catalogue, Horizon, Collibra
- Foundations: Python
- YAML data contracts
- Apache Iceberg
- Git and CI/CD
- AI and agents: MCP; vector databases and embeddings
- Data platforms: Databricks, Snowflake (Cortex, Genie)
- AWS and S3
- Quality-Focused: You are rigorous about data accuracy, lineage, and observability, ensuring high standards through validation before data reaches agents or the business
- Agent-Oriented Builder: You enjoy turning governed datasets into reliable, domain-grounded agents that business users can query with confidence
- Collaborative Partner: You thrive working directly with stakeholders and subject-matter experts through iterative build, test, and validation cycles
- Forward-Thinking: You are interested in the “big picture” of lakehouse architecture and open formats, eager to advance agent-based access patterns and best practices
- Governance-Minded: You understand the critical nature of data security in regulated domains and proactively apply masking and row- and column-level controls
- Experience with AWS and S3, in anticipation of onboarding native cloud data sources
- Familiarity with MCP-based data exposure and with embeddings or vector search for retrieval-augmented use cases
- Experience with regulated life-sciences data domains (clinical, commercial, or real-world data)
📌 Senior Data Engineer (AI & Agents) (Toronto)
🏢 Appnovation
📍 Toronto