Agentic AI Data Engineer (Canada)

Agentic AI Data Engineer (Canada)

12 Sep
|
One moment, please
|
Canada

12 Sep

One moment, please

Canada

Position Overview We are seeking a detail-oriented and scientifically-minded Agentic Data Engineer to join the Computational Sciences Center of Excellence (CS CoE) organization at Genentech Research and Early Development (gRED). This role addresses a critical need in scaling our AI models to address the needs in drug discovery by building largely automated, scalable, agent-driven data ingestion and curation pipelines for genomics data. This includes metadata inference, constructing performant query architectures, and transforming high-dimensional datasets (e.g., single-cell omics, clinical trials) into AI-ready training formats.

Key Responsibilities ● Build an agentic data ingestion pipeline.

Triage and prioritize incoming requests to ingest specific datasets.

Clean and organize the data. Build the first pass cleaning and organization steps into the agentic flow.

Validate cross-modal linkage. Add automated checks that catch when ingested data does not connect correctly and flag low quality or mismatched records.

Version every dataset. Retain and make prior versions addressable.

Preserve raw data and provenance.

Make agent workflows log validation and transformation steps so lineage is traceable. Make agents usable across teams.

Move beyond bespoke steps towards agents that teams can reliably use as a shared, deployed service.

Collaborate with AI, software engineering,



and computational biology groups to co-define data standards and conventions.

Qualifications &

Requirements Core (Required) Agentic AI engineering: Demonstrated experience building multi-agent workflows or LLM workflows using tools/frameworks such as LangGraph or LlamaIndex, including tool/function calling and asynchronous task execution.

Python data engineering: Solid Python for data manipulation, working with APIs and databases, and handling heterogeneous data formats.

Data versioning and provenance: Familiarity with dataset versioning approaches (e.g. DVC, lakeFS, or equivalent).

Working knowledge of scientific data structures: Comfortable or willingness to learn common omics data formats like AnnData, H5AD, TileDB.

Basic understanding of omics: No deep bioinformatics expertise required; just a basic understanding of different modalities (e.g. what is RNA-seq vs scRNA-seq vs WES; genomics vs transcriptomics vs proteomics vs metabolomics).

Unit testing: Comfortable writing unit and functional tests to ensure data processing workflows are reliable and reproducible.

Education: Degree in a technical field or equivalent practical experience.

Nice to have: Experience deploying agent workflows as a shared service (e.g., FastAPI or MCP endpoints).

Exposure to cloud (AWS, GCP) and containerization (Docker).

Familiarity with workflow managers such as Nextflow or Snakemake.

📌 Agentic AI Data Engineer (Canada)
🏢 One moment, please
📍 Canada

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: agentic ai data engineer (canada) / canada

Subscribe to this job alert:

Get the latest job offers by email for: agentic ai data engineer (canada) / canada