Data engineering · Case study
Toronto Open Data Pipeline
A reproducible route from City of Toronto source files to tested analytical tables.
At a glance
- CKAN / City files
- PostgreSQL raw tables
- 11 dbt models + 40 tests
- Airflow + marts
The problem
Police, TTC and census data arrive in different formats. The goal was to load them reliably, model them consistently and keep the transformations testable.
My role
Independent project. I designed the ingestion and star schema, wrote the dbt models and tests, scheduled the run with Airflow, and evaluated a delay model.
Key decision
I land source data as text so irregular values do not break ingestion. Typed staging models and dbt tests make those problems visible; full refresh keeps the workflow simpler for complete City snapshots.
Outcome
On a laptop, the 22 September 2026 run loaded 909,698 rows across four raw tables in about 5 seconds; the full pipeline took about 40 seconds. Eleven dbt models ran with 40 tests passed.
Limits
The hourly delay predictor is a prototype: its time-based evaluation did not beat the simpler historical-rate baseline. The pipeline is a local reproducible build, not a claim of production deployment.
See the work
The repository contains the source, setup instructions and the evidence behind these results.