05 Sep
|
Ampstek
|
Ontario
You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures AI agents operate with fresh, trustworthy information.
What You Will Do Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency.
Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers.
Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines.
Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes.
Required Qualifications Deep experience with distributed data systems, SQL, and orchestration tools.
Experience tuning high-throughput database infrastructure.
Knowledge of Google’s GECX is a plus.
Familiarity with chunking strategies and embedding models.
Skillset Requirements ETL & Data Modeling:
Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency.
Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization.
Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures.
Data Governance: Zero-trust access, privacy controls, compliance, and auditability.
Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems.
Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection.
Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms.
Performance Tuning: High-throughput read/write optimization.
Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Outstanding Expectations, OpenLineage).
#J-18808-Ljbffr
📌 Data Engineer (Ontario)
🏢 Ampstek
📍 Ontario