TPS OPEN DATA STUDY · TORONTO · 2014 TO 2025
The citywide total hides what changed
These 464,571 Toronto Police Service records look different once they are separated by category, time, and place.
Share of analysed police-reported records, 2014 to 2025
01 · START WITH THE ROW
What does one row represent?
Each row is an offence or victim record. Several rows can belong to the same police-reported event.
The source file included report years from 2014 through partial 2026.
How to read the countA record is not a unique incident. Treating every row as a separate incident would overstate the number of events.
02 · Time
The categories did not move together
The citywide line combines five categories whose annual counts changed at different times and by different amounts.
-9.2% from 2024
What the total leaves outAuto Theft rose 240% from 2014 to 2023 and then declined. Assault remained the largest category every year, but followed a different pattern.
03 · Category patterns
Each category has a different schedule
Choose a category to compare how its own records are distributed by hour, weekday, and premises type.
More than half of Auto Theft records were associated with outside locations. The hourly distribution peaked at 22:00.
24-hour profile
Peak 22:00 · 8.8%
Weekday profile
Shared 0 to 18% scale
Premises composition
Within-category share
Two examplesAuto Theft peaks at 22:00, and most records are linked to outside locations. Break and Enter has its largest weekday share on Friday, with houses and apartments accounting for more than half of its records.
04 · Space
A high count is not a measure of risk
The map groups approximate coordinates into a grid. It shows where records are concentrated without displaying exact locations.
Relative spatial density
456,337 geocoded records
Relative density within selected category · logarithmic scale
Most records in this view
- 1West Humber-Clairville12,734
- 2Moss Park10,247
- 3Downtown Yonge East9,363
- 4York University Heights9,105
- 5Yonge-Bay Corridor8,862
Raw record counts, not population-adjusted rates.
What the map cannot showNeighbourhood counts have no denominator for population, visitors, vehicles, traffic, or land use. A higher count does not establish a higher risk for an individual.
05 · Model audit
Accuracy needs a baseline
The model classifies an existing record from its context. It does not forecast whether or where a future event will happen.
Compare the score
The random forest reached 59.2% accuracy. A model that always predicts Assault reached 58.2%.
Robbery
Precision17.8%
Recall24.5%
F10.206
Assault35%
Break and Enter13%
Auto Theft23%
Robbery24%
Theft Over5%
View full normalized matrix
| True \ Predicted | Assault | Break and Enter | Auto Theft | Robbery | Theft Over |
|---|---|---|---|---|---|
| Assault | 64% | 14% | 12% | 8% | 2% |
| Break and Enter | 24% | 60% | 11% | 2% | 2% |
| Auto Theft | 14% | 8% | 68% | 7% | 2% |
| Robbery | 35% | 13% | 23% | 24% | 5% |
| Theft Over | 32% | 25% | 25% | 9% | 10% |
How to read the resultThe random forest classified 59.2% of the 2025 test records correctly, compared with 58.2% for the Assault-only baseline. This experiment classifies existing records. It does not predict future crime.
06 · READ RESPONSIBLY
What this study can and cannot say
The charts describe police-reported records. They need to be read with the unit, denominator, and comparison baseline in view.
Records are not unique incidents
Several offence or victim rows can share one event ID. The row count must be read on its own terms.
Counts are not population-adjusted risk
The data has no denominator for residents, visitors, vehicles, traffic, or land use.
Classification is not future prediction
The experiment assigns a category to a recorded occurrence. It does not predict a future event.
Read the methods or reproduce the analysis
The paper documents the full methodology. The code package includes a 5,000-row stratified sample and the outputs used on this page.