13 Aug
|
hireVouch
|
Vaughan
About Us
We are dedicated to building a cleaner, more sustainable future. By applying innovative technology and data-driven insights to green environmental initiatives, we help measure, analyze, and reduce environmental impact at scale. We are looking for a passionate, forward-thinking Junior Data Engineer to join our data team and help build the data pipelines powering our eco-focused solutions.
Position Overview
As a fresh graduate joining our team, you will work closely with senior data engineers and analysts to design, build, and maintain high-volume data pipelines. You will transform raw environmental datasets—such as energy metrics, carbon emissions data, and resource usage—into actionable insights. This role is ideal for a recent Computer Science graduate eager to apply modern data tooling (AWS, PySpark, Python) toward solving meaningful sustainability challenges.
Key Responsibilities
- Pipeline Development: Design, build, and maintain automated batch and real-time ETL/ELT pipelines to ingest, clean, and transform large-scale environmental data.
- Data Processing: Utilize Python and Apache Spark (PySpark) to process structured and unstructured datasets efficiently.
- Cloud Infrastructure: Help manage and expand our cloud data infrastructure using core AWS services (e.g., S3, Glue, EMR, Redshift, Lambda).
- Data Quality & Governance: Implement automated testing, validation, and monitoring to ensure data accuracy, reliability, and security.
- Cross-Functional Collaboration: Partner with Data Scientists, Business Analysts,
and Sustainability Specialists to deliver clean, structured data for reporting and machine learning applications.
Required Qualifications
- Education: Bachelor’s degree in Computer Science (or a closely related core computing field, such as Computer Engineering or Software Engineering) completed within the last 0–12 months.
- Core Programming: Strong foundation in Python and fundamental software engineering principles (OOP, data structures, algorithms, version control with Git).
- Distributed Computing: Academic or hands-on project experience using Apache Spark / PySpark to process large datasets.
- Cloud Fundamentals: Working knowledge or project experience with AWS core services (S3, EC2, IAM, Lambda, or managed data services).
- Databases & SQL: Solid grasp of relational databases, SQL query writing, data modeling concepts, and basic schema design.
Nice-to-Have / Preferred Qualifications
- Coursework, internship, or personal project focus on environmental data, sustainability, clean energy, or IoT telemetry data.
- Exposure to workflow orchestration tools (e.g., Apache Airflow, Dagster).
- Familiarity with containerization technologies (Docker, Kubernetes).
- Knowledge of CI/CD practices for data infrastructure.
What We Offer
- Mission-Driven Impact: Direct involvement in projects that combat climate change and advance sustainable practices.
- Mentorship & Growth: A collaborative setting with dedicated mentorship from experienced senior data engineers.
📌 Junior Data Engineer (Vaughan)
🏢 hireVouch
📍 Vaughan