19 Aug
|
Themesoft
|
Toronto
Position: Big Data Developer
Location: Toronto – Hybrid
Skills:
- Big Data & Spark
- SQL
- Unix/Shell Scripting
Role Overview
We are looking for a Senior Backend Developer with 5 years of experience in big data engineering, API integration, and AI-assisted development. The ideal candidate will design, build, and maintain scalable data pipelines and backend systems in a enterprise environment.
Key Responsibilities
Big Data & Spark
- Design and develop Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters
- Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL
- Tune Spark jobs for performance (partitioning, caching, broadcast joins, shuffle optimization)
- Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms
- Work with Parquet, ORC, Avro file formats on HDFS
SQL & Data Engineering
- Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations
- Design and maintain Hive external/managed tables and partitioned datasets
- Optimize slow-running queries and resolve correlated subquery issues
- Work with HDFS encryption zones and data governance requirements
Unix / Shell Scripting
- Develop and maintain bash shell scripts for job orchestration and automation
- Handle error management, return codes, logging and alerting in shell scripts
- Manage HDFS operations (hdfs dfs commands), file transfers, and data validation
- Manage Kerberos authentication (kinit, keytab handling)
API Extraction & Integration
- Build scripts and pipelines to extract data from REST APIs using curl and Python
- Handle OAuth2 token generation, bearer token refresh and API health checks
- Parse and process JSON API responses and load into HDFS/Hive
- Manage pagination, error handling and retry logic for API calls
- Work with enterprise API gateways and URL parameter construction
AI & Copilot Capabilities
- Leverage GitHub Copilot / AI coding assistants to accelerate development
- Use AI tools for code review, SQL generation, script debugging and documentation
- Contribute to AI-assisted data quality and anomaly detection pipelines
- Explore and implement LLM-based automation for repetitive data engineering tasks
Scheduling & Orchestration
- Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron
- Build and maintain Ansible playbooks for automated deployments
- Manage deployment pipelines including artifact versioning, Vault secret injection and workplace-specific configuration
- Monitor job health, handle failures and implement alerting
Nice to Have
- Experience with Cloudera CDP (7.x) and migration from HDP
- Knowledge of Kerberos, Vault, HDFS encryption zones
- Familiarity with CI/CD pipelines (Helios, GitHub Actions)
- Experience with MSSQL / JDBC connectivity from Spark
- Understanding of AML / Financial regulatory data domains
Regards
Patrick Fernandez
Talent Acquisition Group - Strategic Recruitment Manager
📌 Big Data Developer (Toronto)
🏢 Themesoft
📍 Toronto