Housing Market Classifier

Predicts where the Norwegian housing market is heading next quarter. For each of Norway's 15 counties, it labels the coming quarter as Hot, Stable or Cooling, using macroeconomic data from Statistics Norway (SSB). The short version of the result: the model does not beat a simple seasonal baseline.

Notice

Earlier versions of this project reported 86.5% test accuracy. That number came from a train/test split that leaked, and it has been retracted. See The leakage bug below.

results.json
645
County-quarters (2014+)
32
Features
7
Unseen test quarters
0.639
Random Forest test macro-F1
0.642
Seasonal lookup test macro-F1

The test set is 7 quarters the model never saw during training (2023K1 to 2024K3). Only the model with the best validation score is evaluated on the test set, so Gradient Boosting has no test score.

Model / baselineVal macro-F1Test macro-F1Test accuracy
Majority class0.2210.19741.9%
Persistence0.3220.35744.8%
Seasonal lookup0.6270.64292.4%
Random Forest0.5900.63991.4%
Gradient Boosting0.508n/an/a
what_this_means.txt

The Random Forest scores 0.003 macro-F1 below the seasonal lookup on the test set. A lookup table that only knows "which quarter of the year is it" matches the full model.

This is the honest result, and also the most interesting one. Norwegian house prices follow a strong yearly cycle: in the training window, 83% of Q1s are Hot and 77% of Q3s are Cooling. Since the label is next quarter's raw price change sorted into three bins, the calendar almost decides the answer by itself.

An ablation backs this up. Removing seasonal_factor and quarter_num drops the Random Forest to 0.555 test macro-F1. Nearly all of its apparent skill comes from seasonality, and the macroeconomic series add nothing on top of the calendar.

per_class_test.txt

Random Forest on the test set:

ClassPrecisionRecallF1Support
Hot0.981.000.9944
Stable0.000.000.008
Cooling0.880.980.9353

The 91.4% accuracy looks better than it is. The model never predicts Stable, and Stable is only 8 of the 105 test rows, so missing all of them barely moves accuracy. Macro-F1 is the number to look at.

confusion_matrix.png
Random Forest confusion matrices for validation (2020K4 to 2022K4) and test (2023K1 to 2024K3). On test, Hot and Cooling are almost always right and no Stable quarter is predicted correctly.

Chronological split. On the test quarters, Hot and Cooling are almost always right, while 7 of the 8 Stable quarters are predicted as Cooling.

feature_importance.png
Top 15 Random Forest features by Gini importance. seasonal_factor and quarter_num are the two most important.

The two most important features are seasonal_factor and quarter_num: the same finding, showing up in the importances.

the_leakage_bug.txt

The original evaluation was wrong, and fixing it changed the conclusion of the project. create_labels returned its data sorted by region and quarter, and train_model.py then took the first 65% of rows as training data under a comment that said "Time-based split". Since the rows were sorted by county name first, the data was actually split alphabetically by county, not by time.

Old (broken) splitNew split
TrainAgder to Telemark, all years2014K1 to 2020K3
ValidationTroms, Trøndelag, most of Vestfold2020K4 to 2022K4
TestRest of Vestfold, Vestland, Østfold2023K1 to 2024K3

This leaked information in two ways:

  1. Shared labels. The 15 counties map to only 10 SSB price regions. Buskerud, Telemark, Vestfold and Østfold share one price series, so their labels are identical every quarter. Most of the old test set had the same answer as a row the model had already trained on for that quarter.
  2. The same quarters in train and test. All 79 quarters showed up in both sets. National series like CPI, the policy rate and GDP are unique per quarter, so together they work like a timestamp. The model could learn "2016K2 was Hot" from one county and reuse it for another.

The fix assigns whole quarters to each set, so the model is only scored on quarters it has never seen. Two related bugs were fixed at the same time: several .diff() and .rolling() features are now grouped by region, and z-scored features now use per-region expanding statistics that only look at the past.

labels.txt

Each county-quarter is labelled by the price change in the next quarter. Features only use current and past data.

LabelNext-quarter price change
Hotabove +2.0%
Stablebetween −0.5% and +2.0%
Coolingbelow −0.5%

Compared against three baselines: majority class, persistence (repeat last quarter's label) and seasonal lookup (predict from the quarter of the year only).

C:\pipeline>
fetch_ssb_data.py (10 SSB tables) ↓ data_parser.py (JSON-stat2 to CSV, county harmonisation) ↓ enhanced_features.py (features + labels) ↓ train_model.py (chronological split, baselines, sklearn)

Norway redrew its county map in 2020 and 2024. The parser maps everything onto one consistent set of 15 modern counties, so each series means the same thing across the whole period.

Data Sources (SSB API)
SSB tableWhat it measuresFrequencyCoverage
03013Consumer price index (CPI)Monthly2005-2025
10701Norges Bank policy rateMonthly2014-2026
01222Population change by countyQuarterly2005-2024
07221House price index by regionQuarterly2005-2024
10187Property sales volumeQuarterly2008-2024
13760Unemployment rateMonthly2006-2026
03723Building starts by countyMonthly2005-2026
10748Mortgage interest ratesMonthly2014-2025
09171GDP volume changeQuarterly2005-2025
06944Household income by countyAnnual2005-2023
known_limitations.txt
  • 15 counties, but only 10 price series. Several counties share a price series and therefore identical labels, so the effective sample is smaller than 645 rows suggests.
  • Only 2014 onward by default. The policy and mortgage rate series start in 2014. Missing values are now imputed with medians fitted on the training window only, instead of filled with zeros.
  • Household income is dropped. SSB has no data for 7 of the 15 counties, so it is removed by the coverage rule.
  • Small test set. 7 quarters and 105 rows, with only 8 Stable rows. The Stable scores are indicative at best.
next_steps.txt

As the task is defined now, it is close to a seasonal lookup, which leaves the model very little room to show real skill. The most useful change would be a harder and more meaningful target: label on seasonally adjusted or year-over-year price changes instead of raw quarter-on-quarter changes.

That removes the yearly cycle from the target, so the macroeconomic features would have to carry actual signal. Since it changes what the project claims to predict, it is left as explicit future work.

Tech Stack
  • Python
  • scikit-learn
  • pandas
  • NumPy
  • Matplotlib
  • Seaborn
  • SSB API
background.txt

It started as a Random Forest written from scratch to learn how it works: Gini impurity, bootstrap sampling, random feature subsets and majority voting. Trained on 80 hand-downloaded samples, it reached 67% validation accuracy and could not predict Stable at all. The automated SSB pipeline expanded the data to all 15 counties, and the project moved to scikit-learn's ensemble classifiers.

Claude Opus 4.6 was used as a development tool throughout the project.