Sr. Bigdata developer (Toronto)

Sr. Bigdata developer (Toronto)

11 Aug
|
Cynet Systems
|
Toronto

11 Aug

Cynet Systems

Toronto

Job Overview:

Requirement/Must Have:

- 5+ years of experience in big data engineering, API integration, and AI-assisted development.
- Experience designing and developing Spark-Scala applications for large-scale data processing on Hadoop/CDP clusters.
- Proficiency in building and optimizing ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL.
- Experience migrating Spark 2 applications to Spark 3 on Cloudera CDP platforms.
- Knowledge of Parquet, ORC, Avro file formats on HDFS.
- Ability to write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations.
- Experience designing and maintaining Hive external/managed tables and partitioned datasets.
- Proficiency in developing and maintaining bash shell scripts for job orchestration and automation.
- Experience managing HDFS operations, file transfers, and data validation.
- Experience managing Kerberos authentication (kinit, keytab handling).
- Ability to build scripts and pipelines to extract data from REST APIs using curl and Python.
- Experience handling OAuth2 token generation, bearer token refresh and API health checks.
- Proficiency in parsing and processing JSON API responses.
- Experience scheduling and managing jobs using AAP (Ansible Automation Platform), Control-M, or cron.
- Ability to build and maintain Ansible playbooks for automated deployments.

Responsibilities:

- Design and develop Spark-Scala applications for large-scale data processing.
- Build and optimize ETL/ELT pipelines using Spark DataFrames, Datasets and Spark SQL.
- Tune Spark jobs for performance including partitioning, caching, broadcast joins, and shuffle optimization.
- Migrate Spark 2 applications to Spark 3 on Cloudera CDP platforms.
- Work with Parquet, ORC, Avro file formats on HDFS.
- Write complex HiveQL / Spark SQL queries including window functions, CTEs, subqueries and aggregations.
- Design and maintain Hive external/managed tables and partitioned datasets.




- Optimize slow-running queries and resolve correlated subquery issues.
- Work with HDFS encryption zones and data governance requirements.
- Develop and maintain bash shell scripts for job orchestration and automation.
- Handle error management, return codes, logging and alerting in shell scripts.
- Manage HDFS operations, file transfers, and data validation.
- Manage Kerberos authentication.
- Build scripts and pipelines to extract data from REST APIs using curl and Python.
- Handle OAuth2 token generation, bearer token refresh and API health checks.
- Parse and process JSON API responses and load into HDFS/Hive.
- Manage pagination, error handling and retry logic for API calls.
- Work with enterprise API gateways and URL parameter construction.
- Leverage GitHub Copilot / AI coding assistants to accelerate development.
- Use AI tools for code review, SQL generation, script debugging and documentation.
- Contribute to AI-assisted data quality and anomaly detection pipelines.
- Explore and implement LLM-based automation for repetitive data engineering tasks.
- Schedule and manage jobs using AAP (Ansible Automation Platform) / Control-M / cron.
- Build and maintain Ansible playbooks for automated deployments.
- Manage deployment pipelines including artifact versioning, Vault secret injection and setting-specific configuration.
- Monitor job health, handle failures and implement alerting.

Nice to Have:

- Experience with Cloudera CDP (7.x) and migration from HDP.
- Knowledge of Kerberos, Vault, HDFS encryption zones.
- Familiarity with CI/CD pipelines (Helios, GitHub Actions).
- Experience with MSSQL / JDBC connectivity from Spark.
- Understanding of AML / Financial regulatory data domains.

Skills:

- Scala.
- Apache Spark.
- Spark SQL.
- DataFrames.
- Datasets.
- Big Data Engineering.
- Hadoop.
- HDFS.
- Cloudera CDP.
- Hive.
- HiveQL.
- ETL.
- ELT.
- SQL.
- Unix.
- Linux.
- Bash Shell Scripting.
- Python.
- REST APIs.
- API Integration.
- JSON.
- OAuth2.
- Ansible Automation Platform (AAP).
- GitHub Copilot.
- Control-M.

📌 Sr. Bigdata developer (Toronto)
🏢 Cynet Systems
📍 Toronto

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: sr. bigdata developer (toronto) / toronto

Subscribe to this job alert:

Get the latest job offers by email for: sr. bigdata developer (toronto) / toronto