If working with billions of events, petabytes of data, and optimizing for the last millisecond is something that excites you then read on! We are looking for a Senior Data Engineer who has seen their fair share of messy data sets and has been able to structure them for further fraud detection and prevention; anomaly detection and other AI products. You will be working on writing frameworks for real-time and batch pipelines to ingest and transform events from 100’s of applications every day.
These events will be consumed by both machines and people. Our ML and Software engineers consume these events to build current and optimize existing models to detect and fight new fraud patterns. You will also help optimize the feature pipelines for fast execution and work with software engineers to build event‑driven microservices.
You will get to put cutting‑edge tech in production and the freedom to experiment with new frameworks, try new ways to optimize, and resources to build the next big thing in fintech using data! This position will operate on a hybrid model and be based in our Toronto office, mandatory in‑office days are every Tuesday and one Friday every month. These designated days and frequency may change according to the company's discretion.
Work directly with the Platform Engineering Team to create reusable experimental and production data pipelines and centralize the data store. Keep the data whole, safe,
and flowing with expertise on high‑volume data ingest and streaming platforms (like Spark Streaming, Kafka, etc). Make the data available for online and offline consumption by machines and humans.
Shape the data by developing efficient structures and schema for the data in storage and transit. Explore new technology options for data processing, storage, and share them with the team. Degree in Computer Science, Engineering, or a related field Proficient in Spark/Scala/Python/Java.
You are passionate about producing clean, maintainable, and testable code as part of a real‑time data pipeline. You have experience implementing offline and online data processing flows and understand how to choose and optimize underlying storage technologies. You have worked or experimented with NoSQL databases such as You can connect different services and processes together even if you have not worked with them before and follow the flow of data through various pipelines to debug data issues.
You have previously worked on building serious data pipelines ingesting and transforming 10 ^6 events per minute and terabytes of data per day. You may not be a computer network expert but you understand issues with ingesting data from applications in multiple data centers across geographies, on‑premises and cloud, and will find a way to solve them. We ensure flexible hours outside of our core working hours Enrolment in the group health benefits plan right from day 1, no waiting period Team building events
📌 Data Engineer Python Senior F/H (Toronto)
🏢 Paytm
📍 Toronto