Iris Software is looking to hire a Data Engineer for a full-time opportunity in Toronto, ON (hybrid position). Please respond back with your most recent resume if you would be interested...!
Job Title: Data Engineer
Location: Toronto, ON (2 days onsite)
Full-time with Iris Software for one of the banks in downtown Toronto The client is looking for:
- feature engineering
- data wrangling
- model automation
- Tool stack probably in this order – Python, PySpark, Glue, Sagemaker, SAS
We are seeking a highly skilled Data Engineer to help build a net-new PySpark data engineering capability from the ground up. The initial focus of this role will be designing and developing net-new feature engineering pipelines to support downstream data science initiatives.
As the platform capability matures, you will play a critical role in a large-scale modernization effort, tasked with translating and migrating legacy SAS-based workloads into a modern, scalable PySpark environment managed via Anaconda.
The ideal candidate possesses strong core data engineering expertise, experience building data preparation pipelines, and the analytical ability to reverse-engineer legacy business logic into optimized open-source code.
Key Responsibilities
- Foundation & Feature Engineering: Lead the foundational setup of Python/PySpark development environments, actively managing complex package dependencies and virtual environments using Conda/Anaconda.
- Pipeline Development: Design, build, and deploy net-new data pipelines focused heavily on feature engineering and data preparation to feed downstream machine learning and analytics use cases.
- Code Translation: Analyze and reverse-engineer existing SAS programs (developed by business users and data scientists) to accurately extract business rules, data transformations, and calculations.
- Platform Modernization: Translate extracted SAS logic into efficient, scalable, and functionally equivalent PySpark code.
- Performance Tuning: Optimize PySpark code for performance, scalability, and maintainability within distributed data processing environments.
- Stakeholder Collaboration: Collaborate closely with business users and data scientists to clarify requirements, validate feature outputs, and resolve discrepancies during the migration process.
- Quality Assurance: Perform robust unit testing, data reconciliation, and automated validation to ensure absolute data parity between legacy SAS outputs and the new PySpark pipelines.
- Documentation: Document technical designs, code lineage, testing results, and migration methodologies for future team scaling.
Required Qualifications
- 5+ years of hands-on experience in data engineering, data pipeline development, or building data platforms.
- Robust programming expertise in Python and PySpark for distributed data processing.
- Proven experience building data pipelines specifically for feature engineering, data curation, and advanced data preparation.
- Ability to read, interpret, and reverse-engineer legacy SAS code (such as SAS data steps, procedures, and macros) to extract complex business logic. (Note: Deep SAS development experience is a plus, but the ability to translate it is the core requirement).
- Experience establishing and managing Python environments, ensuring reproducibility, and handling library dependencies using Anaconda/Conda.
- Advanced SQL skills and experience working with large-scale structured datasets.
- Experience with code migration, platform modernization, or translating legacy codebases into modern open-source stacks.
- Strong analytical, problem-solving, and communication skills, with a track record of successfully interfacing directly with business stakeholders.
Preferred Qualifications
- Experience working with cloud-based data platforms (Azure, AWS, or GCP).
- Familiarity with version control tools (Git) and collaborative development workflows.
- Experience with CI/CD pipelines, automated testing, and MLOps principles.
- Exposure to financial services, banking, or other highly regulated industries.
- Knowledge of data governance, data cataloging, and data quality best practices.
Key Skills
- Python & PySpark
- Feature Engineering & Data Pipelines
- Anaconda / Conda Workplace Management
- SAS Interpretation & Code Translation
- Distributed Computing & Spark Optimization
- SQL & Data Reconciliation
- Legacy Platform Modernization
- Agile Delivery
About Iris Software Inc.
With 4,000+ associates and offices in India, U.S.A. and Canada, Iris Software delivers technology services and solutions that help clients complete fast, far-reaching digital transformations and achieve their business goals. A strategic partner to Fortune 500 and other top companies in financial services and many other industries, Iris provides a value-driven approach - a unique blend of highly skilled specialists, software engineering expertise, cutting-edge technology, and flexible engagement models. High customer satisfaction has translated into long-standing relationships and preferred-partner status with many of our clients, who rely on our 30+ years of technical and domain expertise to future-proof their enterprises.
Associates of Iris work on mission-critical applications supported by a workplace culture that has won numerous awards in the last few years, including Certified Great Place to Work in India; Top 25 GPW in IT & IT-BPM; Ambition Box Best Place to Work, #3 in IT/ITES; and Top Workplace NJ-USA.
Kind regards,
Amarpreet Singh
200 Bay Str. Toronto, ON, M5H4E9
200 Metroplex Drive, Suite #300 Edison, NJ 08817
[email protected] | www.irissoftware.com
Iris Software ranked 38th among 1200 companies spanning 20+ industries in the country to become India’s Best Companies to Work For 2022!
📌 Data Engineer (Toronto)
🏢 Iris Software
📍 Toronto