30 Jul
|
E-Solutions
|
Toronto
30 Jul
E-Solutions
Toronto
Location: Toronto, ON (4 Days/Week Onsite)
Job Type: Contract
Note: Face-to-face (F2F) interview is required
Job Description:
We are looking for a Sr. Big Data Developer with experience in Big Data Engineering, API Integration, and AI-assisted development . The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in an enterprise environment.
Key Responsibilities
- Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters
- Build and optimize ETL/ELT pipelines using Spark Data Frames, Datasets and Spark SQL
- Tune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)
- Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms
- Work with Parquet, ORC, Avro file formats on HDFS
- Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations
- Design and maintain Hive external/managed tables and partitioned datasets
- Optimize slow-running queries and resolve correlated subquery issues
- Work with HDFS encryption zones and data governance requirements
Unix / Shell Scripting
- Develop and maintain bash shell scripts for job orchestration and automation
- Handle error management, return codes, logging and alerting in shell scripts
- Manage HDFS operations (hdfs dfs commands), file transfers, and data validation
API Extraction & Integration
- Build scripts and pipelines to extract data from REST APIs using curl and Python
- Parse and process JSON API responses and load into HDFS/Hive
- Manage pagination, error handling and retry logic for API calls
- Work with enterprise API gateways and URL parameter construction
- Leverage GitHub Copilot / AI coding assistants to accelerate development
- Use AI tools for code review, SQL generation, script debugging and documentation
- Contribute to AI-assisted data quality and anomaly detection pipelines
- Explore and implement LLM-based automation for repetitive data engineering tasks
Scheduling & Orchestration
- Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron
- Build and maintain Ansible playbooks for automated deployments
- Manage deployment pipelines including artifact versioning, Vault secret injection and setting-specific configuration
- Monitor job health, handle failures and implement alerting
Nice to Have
- Experience with Cloudera CDP (7.x) and migration from HDP
- Knowledge of Kerberos, Vault, HDFS encryption zones
- Familiarity with CI/CD pipelines (Helios, GitHub Actions)
- Experience with MSSQL / JDBC connectivity from Spark
- Understanding of AML / Financial regulatory data domains
#J-18808-Ljbffr
📌 Sr. Bigdata developer (Toronto)
🏢 E-Solutions
📍 Toronto