Duo et al. 2018 single-cell clustering

Generated by beam 0.2.0 at 2026-07-09T16:02:42.164467+00:00.

Under weighted sum aggregation with equal weights over the metrics ari, runtime and shannon_entropy_diff, Seurat ranks first of 12 tools. Across 1000 weightings drawn at random from the metric simplex, Seurat ranked first in 100 percent of draws. Its rank held in 3 of 3 leave-one-metric-out runs. Its rank held in 12 of 12 leave-one-dataset-out runs. The top rank is stable: the smallest single-metric weight change that would overturn it is about 53.15.

Input

layoutlong
tools12
metricsari, runtime, shannon_entropy_diff
datasets12: Koh, KohTCC, Kumar, KumarTCC, SimKumar4easy, SimKumar4hard, SimKumar8hard, Trapnell, TrapnellTCC, Zhengmix4eq, Zhengmix4uneq, Zhengmix8eq
weightingequal
aggregationsaw

Normalization

metricstrategy
aribaseline_relative
runtimelog_min_max
shannon_entropy_diffmin_max

No normalization guard warnings.

Ranking

composite score per tool
ranktoolcomposite
1Seurat0.9470
2PCAKmeans0.8562
3PCAHC0.8555
4CIDR0.8303
5monocle0.8232
6RtsneKmeans0.8079
7TSCAN0.7716
8FlowSOM0.7229
9SC3svm0.6560
10SC30.6128
11pcaReduce0.5798
12RaceID20.5671
normalized scores heatmap

Robustness at a glance

The glyph table shows the per-metric normalized scores as circle size and the overall composite as a bar, alongside panels for how far each tool's rank moves under the robustness checks: the span across leave-one-dataset-out runs, the SMAA rank acceptability, and the span across the five aggregations. A tool with a tight rank span is a stable recommendation; a wide span means the order at that position depends on a choice the analyst made.

funky heatmap with rank-robustness panels

Sensitivity

The ranking above commits to one weighting (equal) and one aggregation (saw), the way a reader would apply the tool in practice. The sections here re-run it under the alternatives, so you can read how much the recommendation depends on those analyst choices rather than on the data.

SMAA weight sampling

SMAA confidence per tool
toolranked first (percent of draws)
Seurat99.9
PCAKmeans0.0
PCAHC0.0
CIDR0.0
monocle0.0
RtsneKmeans0.0
TSCAN0.0
FlowSOM0.0
SC3svm0.0
SC30.1
pcaReduce0.0
RaceID20.0

Leave one metric out

The top-ranked tool kept its rank in 100 percent of the 3 leave-one-metric-out runs. Removing runtime caused the largest rank change, a shift of 8.

Leave one dataset out

rank stability per tool across dataset omissions

The top-ranked tool kept its rank in 100 percent of the 12 leave-one-dataset-out runs. Omitting dataset KohTCC caused the largest rank change, a shift of 1.

Aggregation agreement

Re-ranking under the 5 aggregations gives orderings that agree at a mean Kendall tau-b of 0.62. Every aggregation ranks Seurat first, so the top recommendation does not depend on the aggregation choice.

Normalization agreement

Re-ranking under 5 normalizations (recommended, min_max, log_min_max, rank, zscore) gives orderings that agree at a mean Kendall tau-b of 0.64. The recommended column is the per-metric default each card declares (ari: baseline_relative, runtime: log_min_max, shannon_entropy_diff: min_max); the others apply one strategy to every metric. Every normalization ranks Seurat first, so the top recommendation does not depend on the normalization choice.

normalization agreement heatmap

What moves the ranking

Across the 240 combinations of weighting, aggregation and dataset, the rank variance splits into the dataset (71 percent), the weighting choice (4 percent), the aggregation choice (4 percent), and their interactions (21 percent). The largest is the dataset.

The top tool overall (Seurat) ranks first in 49 percent of combinations, with a rank span of 10. Its mean rank per dataset:

datasetmean rank of Seurat
Koh1.1
KohTCC1.4
Kumar4.7
KumarTCC4.9
SimKumar4easy2.0
SimKumar4hard2.0
SimKumar8hard1.1
Trapnell4.4
TrapnellTCC1.8
Zhengmix4eq1.0
Zhengmix4uneq1.0
Zhengmix8eq1.0

Specification curve

Each of the 240 combinations of weighting, aggregation and dataset is one specification. The method that ranks first most often (Seurat) ranks first in 49 percent of them. The single most common ordering of all tools holds in 3 percent of specifications, and 5 distinct tool(s) rank first in at least one specification. The curve below plots that method's rank across the specifications, sorted from its best rank to its worst; the panel beneath marks which choices each specification used.

specification curve

Smallest weight perturbation

The top rank is stable: the smallest single-metric weight change that would overturn it is about 53.145 on ari.

Reference levels

Chance baseline

How many tools score above a random method on each metric that declares a chance baseline.

metricchance scoretools above chance
ari0 12 of 12 (100 percent)

Every tool scores above chance on at least one metric.

Noise floor

The top two tools (Seurat and PCAKmeans) are separated above the noise floor on at least one metric.

Tool pairs the metric set cannot tell apart above the noise floor:

Critical difference across datasets

Friedman test p-value 1.093e-11 over 12 datasets. The Nemenyi critical difference is 4.810 average-rank units; tools closer than this are not separable at the chosen level.

average ranks with critical difference

Groups not significantly different:

Effect size: the top-ranked method (Seurat) scores higher than the next (PCAHC) on 9 of 12 datasets (probability of superiority 75 percent).

Consistency: the pairwise majorities are transitive, so one order is consistent with them .

pairwise dominance matrix ordered by methods outperformed

Posterior: under the Bayesian sign test (ROPE 0), the probability that Seurat is practically better than PCAHC is 0.97, and the probability that the two are practically equivalent is 0.00.

posterior probability that the row method is practically better than the column method

Reproducibility

Input sha256: 20585d88e44beeb39fffbdd206258b980f73578669dd04960e39d2cb2edab4a2

packageversion
python3.12.13
beam0.2.0
numpy2.5.1
scipy1.18.0
pymcdm1.4.0
pyyaml6.0.3
jsonschema4.26.0

The full run manifest, written alongside this report, records the metric card versions and content hashes needed to reproduce the run.