← All projects

Data engineering · Case study

Toronto Open Data Pipeline

A reproducible route from City of Toronto source files to tested analytical tables.

At a glance

  1. CKAN / City files
  2. PostgreSQL raw tables
  3. 11 dbt models + 40 tests
  4. Airflow + marts

The problem

Police, TTC and census data arrive in different formats. The goal was to load them reliably, model them consistently and keep the transformations testable.

My role

Independent project. I designed the ingestion and star schema, wrote the dbt models and tests, scheduled the run with Airflow, and evaluated a delay model.

Key decision

I land source data as text so irregular values do not break ingestion. Typed staging models and dbt tests make those problems visible; full refresh keeps the workflow simpler for complete City snapshots.

Outcome

On a laptop, the 22 September 2026 run loaded 909,698 rows across four raw tables in about 5 seconds; the full pipeline took about 40 seconds. Eleven dbt models ran with 40 tests passed.

Limits

The hourly delay predictor is a prototype: its time-based evaluation did not beat the simpler historical-rate baseline. The pipeline is a local reproducible build, not a claim of production deployment.

See the work

The repository contains the source, setup instructions and the evidence behind these results.

← Back to projectsGet in touch →