M4 forecasting competition

Generated by beam 0.2.0 at 2026-07-09T16:02:45.232412+00:00.

Under weighted sum aggregation with equal weights over the metrics smape and mase, Pawlikowski ranks first of 25 tools. Across 1000 weightings drawn at random from the metric simplex, Pawlikowski ranked first in 100 percent of draws. Its rank held in 1 of 2 leave-one-metric-out runs. Its rank held in 5 of 6 leave-one-dataset-out runs. The top rank is stable: the smallest single-metric weight change that would overturn it is about 0.50.

Consistency: the pairwise majorities are not transitive (0 circular triads), so no single order represents them, though Montero-Manso is the method preferred to every other one individually. See the critical-difference section below.

Input

layoutlong
tools25
metricssmape, mase
datasets6: Yearly, Quarterly, Monthly, Weekly, Daily, Hourly
weightingequal
aggregationsaw

Normalization

metricstrategy
smapemin_max
masemin_max

Ranking

composite score per tool
ranktoolcomposite
1Pawlikowski0.6305
2Montero-Manso0.6133
3Smyl0.6036
4Doornik0.6015
5Jaganathan0.6013
6Tartu M4 seminar0.5961
7Fiorucci0.5950
8Petropoulos0.5938
9Waheeb0.5930
10Pedregal0.5898
11Nikzad0.5855
12Roubinchtein0.5784
13ARIMA0.5763
14Darin0.5727
15Spiliotis0.5613
16Shaub0.5599
17Dantas0.5567
18Segura-Heras0.5550
19Trotta0.5525
20ETS0.5509
21Legaki0.5379
22Theta0.5259
23Ibrahim0.5210
24Damped0.5193
25Comb0.4698
normalized scores heatmap

Robustness at a glance

The glyph table shows the per-metric normalized scores as circle size and the overall composite as a bar, alongside panels for how far each tool's rank moves under the robustness checks: the span across leave-one-dataset-out runs, the SMAA rank acceptability, and the span across the five aggregations. A tool with a tight rank span is a stable recommendation; a wide span means the order at that position depends on a choice the analyst made.

funky heatmap with rank-robustness panels

Sensitivity

The ranking above commits to one weighting (equal) and one aggregation (saw), the way a reader would apply the tool in practice. The sections here re-run it under the alternatives, so you can read how much the recommendation depends on those analyst choices rather than on the data.

SMAA weight sampling

SMAA confidence per tool
toolranked first (percent of draws)
Pawlikowski99.5
Montero-Manso0.0
Smyl0.5
Doornik0.0
Jaganathan0.0
Tartu M4 seminar0.0
Fiorucci0.0
Petropoulos0.0
Waheeb0.0
Pedregal0.0
Nikzad0.0
Roubinchtein0.0
ARIMA0.0
Darin0.0
Spiliotis0.0
Shaub0.0
Dantas0.0
Segura-Heras0.0
Trotta0.0
ETS0.0
Legaki0.0
Theta0.0
Ibrahim0.0
Damped0.0
Comb0.0

Leave one metric out

The top-ranked tool kept its rank in 50 percent of the 2 leave-one-metric-out runs. Removing mase caused the largest rank change, a shift of 8.

Leave one dataset out

rank stability per tool across dataset omissions

The top-ranked tool kept its rank in 83 percent of the 6 leave-one-dataset-out runs. Omitting dataset Hourly caused the largest rank change, a shift of 13.

Aggregation agreement

Re-ranking under the 5 aggregations gives orderings that agree at a mean Kendall tau-b of 0.90. Every aggregation ranks Pawlikowski first, so the top recommendation does not depend on the aggregation choice.

Normalization agreement

Re-ranking under 5 normalizations (recommended, min_max, log_min_max, rank, zscore) gives orderings that agree at a mean Kendall tau-b of 0.90. The recommended column is the per-metric default each card declares (smape: min_max, mase: min_max); the others apply one strategy to every metric. Every normalization ranks Pawlikowski first, so the top recommendation does not depend on the normalization choice.

normalization agreement heatmap

What moves the ranking

Across the 120 combinations of weighting, aggregation and dataset, the rank variance splits into the dataset (96 percent), the weighting choice (0 percent), the aggregation choice (0 percent), and their interactions (3 percent). The largest is the dataset.

The top tool overall (Smyl) ranks first in 33 percent of combinations, with a rank span of 17. Its mean rank per dataset:

datasetmean rank of Smyl
Yearly1.0
Quarterly2.6
Monthly1.0
Weekly9.9
Daily17.8
Hourly5.4

Specification curve

Each of the 120 combinations of weighting, aggregation and dataset is one specification. The method that ranks first most often (Smyl) ranks first in 33 percent of them. The single most common ordering of all tools holds in 8 percent of specifications, and 5 distinct tool(s) rank first in at least one specification. The curve below plots that method's rank across the specifications, sorted from its best rank to its worst; the panel beneath marks which choices each specification used.

specification curve

Smallest weight perturbation

The top rank is stable: the smallest single-metric weight change that would overturn it is about 0.499 on mase.

Critical difference across datasets

Friedman test p-value 0.0001058 over 6 datasets. The Nemenyi critical difference is 15.543 average-rank units; tools closer than this are not separable at the chosen level.

average ranks with critical difference

Groups not significantly different:

Effect size: the top-ranked method (Montero-Manso) scores higher than the next (Pawlikowski) on 5 of 6 datasets (probability of superiority 83 percent).

Consistency: the pairwise majorities carry 0 circular triads out of 2300, so no single order agrees with all of them and the reported ranking depends on the aggregation rule (though Montero-Manso is still preferred to every other method individually).

pairwise dominance matrix ordered by methods outperformed

Posterior: under the Bayesian sign test (ROPE 0), the probability that Montero-Manso is practically better than Pawlikowski is 0.94, and the probability that the two are practically equivalent is 0.03.

posterior probability that the row method is practically better than the column method

Reproducibility

Input sha256: b1da4c264cb28c88b3e9d5dc19f010db67b33e680a89861d2cda3ced213fa562

packageversion
python3.12.13
beam0.2.0
numpy2.5.1
scipy1.18.0
pymcdm1.4.0
pyyaml6.0.3
jsonschema4.26.0

The full run manifest, written alongside this report, records the metric card versions and content hashes needed to reproduce the run.