Under weighted sum aggregation with equal weights over the metrics smape and mase, Pawlikowski ranks first of 25 tools. Across 1000 weightings drawn at random from the metric simplex, Pawlikowski ranked first in 100 percent of draws. Its rank held in 1 of 2 leave-one-metric-out runs. Its rank held in 5 of 6 leave-one-dataset-out runs. The top rank is stable: the smallest single-metric weight change that would overturn it is about 0.50.
Consistency: the pairwise majorities are not transitive
(0 circular triads),
so no single order represents them, though
Montero-Manso is the method preferred to every other one
individually. See the critical-difference section below.
| layout | long |
|---|---|
| tools | 25 |
| metrics | smape, mase |
| datasets | 6: Yearly, Quarterly, Monthly, Weekly, Daily, Hourly |
| weighting | equal |
| aggregation | saw |
| metric | strategy |
|---|---|
| smape | min_max |
| mase | min_max |
| rank | tool | composite |
|---|---|---|
| 1 | Pawlikowski | 0.6305 |
| 2 | Montero-Manso | 0.6133 |
| 3 | Smyl | 0.6036 |
| 4 | Doornik | 0.6015 |
| 5 | Jaganathan | 0.6013 |
| 6 | Tartu M4 seminar | 0.5961 |
| 7 | Fiorucci | 0.5950 |
| 8 | Petropoulos | 0.5938 |
| 9 | Waheeb | 0.5930 |
| 10 | Pedregal | 0.5898 |
| 11 | Nikzad | 0.5855 |
| 12 | Roubinchtein | 0.5784 |
| 13 | ARIMA | 0.5763 |
| 14 | Darin | 0.5727 |
| 15 | Spiliotis | 0.5613 |
| 16 | Shaub | 0.5599 |
| 17 | Dantas | 0.5567 |
| 18 | Segura-Heras | 0.5550 |
| 19 | Trotta | 0.5525 |
| 20 | ETS | 0.5509 |
| 21 | Legaki | 0.5379 |
| 22 | Theta | 0.5259 |
| 23 | Ibrahim | 0.5210 |
| 24 | Damped | 0.5193 |
| 25 | Comb | 0.4698 |
The glyph table shows the per-metric normalized scores as circle size and the overall composite as a bar, alongside panels for how far each tool's rank moves under the robustness checks: the span across leave-one-dataset-out runs, the SMAA rank acceptability, and the span across the five aggregations. A tool with a tight rank span is a stable recommendation; a wide span means the order at that position depends on a choice the analyst made.
The ranking above commits to one weighting (equal) and one aggregation (saw), the way a reader would apply the tool in practice. The sections here re-run it under the alternatives, so you can read how much the recommendation depends on those analyst choices rather than on the data.
| tool | ranked first (percent of draws) |
|---|---|
| Pawlikowski | 99.5 |
| Montero-Manso | 0.0 |
| Smyl | 0.5 |
| Doornik | 0.0 |
| Jaganathan | 0.0 |
| Tartu M4 seminar | 0.0 |
| Fiorucci | 0.0 |
| Petropoulos | 0.0 |
| Waheeb | 0.0 |
| Pedregal | 0.0 |
| Nikzad | 0.0 |
| Roubinchtein | 0.0 |
| ARIMA | 0.0 |
| Darin | 0.0 |
| Spiliotis | 0.0 |
| Shaub | 0.0 |
| Dantas | 0.0 |
| Segura-Heras | 0.0 |
| Trotta | 0.0 |
| ETS | 0.0 |
| Legaki | 0.0 |
| Theta | 0.0 |
| Ibrahim | 0.0 |
| Damped | 0.0 |
| Comb | 0.0 |
The top-ranked tool kept its rank in 50 percent of the
2 leave-one-metric-out runs. Removing mase
caused the largest rank change, a shift of 8.
The top-ranked tool kept its rank in 83 percent of the
6 leave-one-dataset-out runs. Omitting dataset
Hourly caused the largest rank change, a shift of
13.
Re-ranking under the 5 aggregations gives orderings that
agree at a mean Kendall tau-b of 0.90.
Every aggregation ranks
Pawlikowski first, so the top recommendation does not
depend on the aggregation choice.
Re-ranking under 5 normalizations
(recommended, min_max, log_min_max, rank, zscore) gives orderings that agree at a mean Kendall tau-b of
0.90. The recommended column is the per-metric
default each card declares (smape: min_max, mase: min_max); the others apply one
strategy to every metric.
Every normalization ranks
Pawlikowski first, so the top recommendation does not
depend on the normalization choice.
Across the 120 combinations of weighting, aggregation and dataset, the rank variance splits into the dataset (96 percent), the weighting choice (0 percent), the aggregation choice (0 percent), and their interactions (3 percent). The largest is the dataset.
The top tool overall (Smyl) ranks first in
33 percent of combinations, with a rank span of
17. Its mean rank per dataset:
| dataset | mean rank of Smyl |
|---|---|
| Yearly | 1.0 |
| Quarterly | 2.6 |
| Monthly | 1.0 |
| Weekly | 9.9 |
| Daily | 17.8 |
| Hourly | 5.4 |
Each of the 120 combinations of weighting, aggregation and
dataset is one specification. The method that ranks first most often
(Smyl) ranks first in
33 percent of them. The single most common ordering of all tools
holds in 8 percent of specifications, and
5 distinct tool(s) rank first in at least one specification.
The curve below plots that method's rank across the specifications, sorted from its best rank to
its worst; the panel beneath marks which choices each specification used.
The top rank is stable: the smallest single-metric weight change that would overturn it is
about 0.499 on mase.
Friedman test p-value 0.0001058 over 6 datasets. The Nemenyi critical difference is 15.543 average-rank units; tools closer than this are not separable at the chosen level.
Groups not significantly different:
Effect size: the top-ranked method (Montero-Manso) scores
higher than the next (Pawlikowski) on
5 of 6 datasets
(probability of superiority 83 percent).
Consistency: the pairwise majorities carry 0 circular triads
out of 2300, so no single order agrees with all of them and the reported ranking depends on the aggregation rule
(though Montero-Manso is still preferred to every other method individually).
Posterior: under the Bayesian sign test (ROPE 0), the probability that
Montero-Manso is practically better than Pawlikowski is 0.94,
and the probability that the two are practically equivalent is 0.03.
| package | version |
|---|---|
| python | 3.12.13 |
| beam | 0.2.0 |
| numpy | 2.5.1 |
| scipy | 1.18.0 |
| pymcdm | 1.4.0 |
| pyyaml | 6.0.3 |
| jsonschema | 4.26.0 |