11 Aug
|
Cynet Systems
|
Toronto
11 Aug
Cynet Systems
Toronto
Job Overview:
Requirement/Must Have:
- 5+ years of experience in big data engineering, API integration, and AI-assisted development.
- Experience designing and developing Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters.
- Proficiency in building and optimizing ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL.
- Experience migrating Spark 2 applications to Spark 3 on Cloudera CDP platforms.
- Knowledge of Parquet, ORC, Avro file formats on HDFS.
- Ability to write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations.
- Experience designing and maintaining Hive external/managed tables and partitioned datasets.
- Proficiency in developing and maintaining bash shell scripts for job orchestration and automation.
- Experience managing HDFS operations, file transfers, and data validation.
- Experience managing Kerberos authentication (kinit, keytab handling).
- Ability to build scripts and pipelines to extract data from REST APIs using curl and Python.
- Experience handling OAuth2 token generation, bearer token refresh and API health checks.
- Proficiency in parsing and processing JSON API responses.
- Experience scheduling and managing jobs using AAP (Ansible Automation Platform), Control-M, or cron.
- Ability to build and maintain Ansible playbooks for automated deployments.
Responsibilities:
- Design and develop Spark-Scala applications for large-scale data processing.
- Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL.
- Tune Spark jobs for performance including partitioning, caching, broadcast joins, and shuffle optimization.
- Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms.
- Work with Parquet, ORC, Avro file formats on HDFS.
- Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations.
- Design and maintain Hive external/managed tables and partitioned datasets.
- Optimize slow-running queries and resolve correlated subquery issues.
- Work with HDFS encryption zones and data governance requirements.
- Develop and maintain bash shell scripts for job orchestration and automation.
- Handle error management, return codes, logging and alerting in shell scripts.
- Manage HDFS operations, file transfers, and data validation.
- Manage Kerberos authentication.
- Build scripts and pipelines to extract data from REST APIs using curl and Python.
- Handle OAuth2 token generation, bearer token refresh and API health checks.
- Parse and process JSON API responses and load into HDFS/Hive.
- Manage pagination, error handling and retry logic for API calls.
- Work with enterprise API gateways and URL parameter construction.
- Leverage GitHub Copilot / AI coding assistants to accelerate development.
- Use AI tools for code review, SQL generation, script debugging and documentation.
- Contribute to AI-assisted data quality and anomaly detection pipelines.
- Explore and implement LLM-based automation for repetitive data engineering tasks.
- Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron.
- Build and maintain Ansible playbooks for automated deployments.
- Manage deployment pipelines including artifact versioning, Vault secret injection and setting-specific configuration.
- Monitor job health, handle failures and implement alerting.
Nice to Have:
- Experience with Cloudera CDP (7.x) and migration from HDP.
- Knowledge of Kerberos, Vault, HDFS encryption zones.
- Familiarity with CI/CD pipelines (Helios, GitHub Actions).
- Experience with MSSQL / JDBC connectivity from Spark.
- Understanding of AML / Financial regulatory data domains.
Skills:
- Scala.
- Apache Spark.
- Spark SQL.
- DataFrames.
- Datasets.
- Big Data Engineering.
- Hadoop.
- HDFS.
- Cloudera CDP.
- Hive.
- HiveQL.
- ETL.
- ELT.
- SQL.
- Unix.
- Linux.
- Bash Shell Scripting.
- Python.
- REST APIs.
- API Integration.
- JSON.
- OAuth2.
- Ansible Automation Platform (AAP).
- GitHub Copilot.
- Control-M.
📌 Sr. Bigdata developer (Toronto)
🏢 Cynet Systems
📍 Toronto