13 Aug
|
Nityo Infotech
|
Mississauga
13 Aug
Nityo Infotech
Mississauga
Required Qualifications:
- Candidate must be located within commuting distance of Mississauga, Ontario or be willing to relocate to the areas.
- Bachelor’s degree or foreign equivalent required from an accredited institution. Will also consider three years of progressive experience in the specialty in lieu of every year of education.
- At least 4 years of Information Technology experience
- 4+ years of experience in Big Data technologies.
- Strong expertise in:
- Apache Spark (Core, SQL, DataFrames, RDDs)
- Scala programming
- PySpark
- Hands-on experience with:
- Kafka (real-time streaming)
- Hadoop ecosystem (HDFS, Hive, Impala)
- NoSQL Databases (HBase, MongoDB, Couchbase)
- Strong understanding of distributed computing concepts and data processing frameworks.
- Experience in building ETL/data pipelines for large-scale datasets.
- Proficiency in SQL and data modeling.
Preferred Qualifications:
- Hands-on experience with data lakes, data warehouses, and scalable ETL pipeline design, including batch and real-time processing architecture.
- Strong understanding and practical exposure to Agile software development methodologies (Scrum) and SDLC practices.
- Proven experience in Banking domain, supporting use cases such as fraud detection, risk analytics, regulatory reporting, and customer insights.
- Excellent analytical, problem-solving, and communication skills, with the ability to translate business requirements into scalable technical solutions.
- Demonstrated ability to work effectively in cross-functional, multi-stakeholder environments, collaborating with Business, Data Engineering, and Architecture teams.
- Experience with real-time data streaming frameworks such as Kafka and Spark Streaming for low-latency processing.
- Understanding data modeling concepts (dimensional modeling, snowflake schemas) to support analytics workloads.
- Experience and desire to work in a global delivery workplace.
Key Responsibilities:
- Design and develop large-scale data processing pipelines using Apache Spark (Scala &
- PySpark)
- Build and optimize batch and real-time data processing workflows using Spark, Kafka, and Hadoop ecosystem
- Develop Spark applications using RDDs, DataFrames, and Spark SQL for complex transformations
- Develop and optimize PySpark applications leveraging joins, Spark DAG execution flow, stage optimization, transformation techniques, and streaming with dynamic allocation and failover handling.
- Implement streaming pipelines using Kafka and Spark Streaming / Structured Streaming
- Develop and maintain HDFS, Hive, NoSql and Impala-based data lake solutions
- Convert existing SQL/Hive workloads into optimized Spark jobs for improved performance
- Work with ETL pipelines to ingest, cleanse, transform, and process large datasets
- Optimize performance through partitioning, caching, serialization, and tuning techniques
- Handle data formats such as Parquet, ORC, Avro, JSON
- Integrate multiple data sources including streaming systems, flat files RDBMS, and APIs
- Collaborate with cross-functional teams to understand business requirements and translate them into scalable technical solutions
- Ensure data quality, reliability, and performance monitoring across pipelines
- Participate in code reviews, design discussions, and best practices implementation
Key Skills:
- Distributed Data Processing.
- Spark Optimization &
- Performance Tuning.
- Real-time Data Streaming.
- Data Modeling &
- ETL Design.
- Problem-solving and Analytical Thinking.
- Strong Communication &
- Stakeholder Management.
📌 Spark Scala Developer (Mississauga)
🏢 Nityo Infotech
📍 Mississauga