OpenProblems batch integration

Generated by beam 0.2.0 at 2026-07-09T16:03:19.144556+00:00.

Under TOPSIS aggregation with equal weights over the metrics ari, asw_batch, asw_label, cell_cycle_conservation, clisi, graph_connectivity, ilisi, isolated_label_asw, isolated_label_f1, kbet, nmi and pcr, combat ranks first of 14 tools. Across 1000 weightings drawn at random from the metric simplex, combat ranked first in 46 percent of draws. Its rank held in 11 of 12 leave-one-metric-out runs. Its rank held in 6 of 6 leave-one-dataset-out runs. The top rank is fragile: a weight change of about 0.04 on pcr is enough to overturn it.

Input

layoutlong
tools14
metricsari, asw_batch, asw_label, cell_cycle_conservation, clisi, graph_connectivity, ilisi, isolated_label_asw, isolated_label_f1, kbet, nmi, pcr
datasets6: cellxgene_census/dkd, cellxgene_census/gtex_v9, cellxgene_census/hypomap, cellxgene_census/immune_cell_atlas, cellxgene_census/mouse_pancreas_atlas, cellxgene_census/tabula_sapiens
weightingequal
aggregationtopsis

Normalization

metricstrategy
aribaseline_relative
asw_batchmin_max
asw_labelmin_max
cell_cycle_conservationmin_max
clisimin_max
graph_connectivitymin_max
ilisimin_max
isolated_label_aswmin_max
isolated_label_f1min_max
kbetmin_max
nmimin_max
pcrmin_max

No normalization guard warnings.

Ranking

composite score per tool
ranktoolcomposite
1combat0.8248
2harmonypy0.7563
3harmony0.7524
4scvi0.7397
5scalex0.7393
6pyliger0.7290
7liger0.7273
8scanvi0.7231
9batchelor_fastmnn0.6759
10scimilarity0.6273
11scgpt_zeroshot0.6096
12uce0.5859
13scanorama0.3061
14geneformer0.0882
normalized scores heatmap

Robustness at a glance

The glyph table shows the per-metric normalized scores as circle size and the overall composite as a bar, alongside panels for how far each tool's rank moves under the robustness checks: the span across leave-one-dataset-out runs, the SMAA rank acceptability, and the span across the five aggregations. A tool with a tight rank span is a stable recommendation; a wide span means the order at that position depends on a choice the analyst made.

funky heatmap with rank-robustness panels

Sensitivity

The ranking above commits to one weighting (equal) and one aggregation (topsis), the way a reader would apply the tool in practice. The sections here re-run it under the alternatives, so you can read how much the recommendation depends on those analyst choices rather than on the data.

SMAA weight sampling

SMAA confidence per tool
toolranked first (percent of draws)
combat45.9
harmonypy2.8
harmony5.7
scvi2.2
scalex0.9
pyliger3.3
liger0.8
scanvi31.0
batchelor_fastmnn6.0
scimilarity1.1
scgpt_zeroshot0.1
uce0.2
scanorama0.0
geneformer0.0

Leave one metric out

The top-ranked tool kept its rank in 92 percent of the 12 leave-one-metric-out runs. Removing pcr caused the largest rank change, a shift of 7.

Leave one dataset out

rank stability per tool across dataset omissions

The top-ranked tool kept its rank in 100 percent of the 6 leave-one-dataset-out runs. Omitting dataset cellxgene_census/gtex_v9 caused the largest rank change, a shift of 6.

Aggregation agreement

Re-ranking under the 4 aggregations gives orderings that agree at a mean Kendall tau-b of 0.55. They do not all rank the same tool first, so the top recommendation depends on which aggregation is used.

Normalization agreement

Re-ranking under 4 normalizations (recommended, min_max, rank, zscore) gives orderings that agree at a mean Kendall tau-b of 0.50. The recommended column is the per-metric default each card declares (ari: baseline_relative, asw_batch: min_max, asw_label: min_max, cell_cycle_conservation: min_max, clisi: min_max, graph_connectivity: min_max, ilisi: min_max, isolated_label_asw: min_max, isolated_label_f1: min_max, kbet: min_max, nmi: min_max, pcr: min_max); the others apply one strategy to every metric. They do not all rank the same tool first, so the top recommendation depends on which normalization is used.

normalization agreement heatmap

What moves the ranking

Across the 96 combinations of weighting, aggregation and dataset, the rank variance splits into the dataset (51 percent), the weighting choice (4 percent), the aggregation choice (16 percent), and their interactions (30 percent). The largest is the dataset.

The top tool overall (combat) ranks first in 36 percent of combinations, with a rank span of 8. Its mean rank per dataset:

datasetmean rank of combat
cellxgene_census/dkd5.8
cellxgene_census/gtex_v92.0
cellxgene_census/hypomap2.9
cellxgene_census/immune_cell_atlas5.8
cellxgene_census/mouse_pancreas_atlas2.9
cellxgene_census/tabula_sapiens3.4

Specification curve

Each of the 96 combinations of weighting, aggregation and dataset is one specification. The method that ranks first most often (combat) ranks first in 35 percent of them. The single most common ordering of all tools holds in 5 percent of specifications, and 9 distinct tool(s) rank first in at least one specification. The curve below plots that method's rank across the specifications, sorted from its best rank to its worst; the panel beneath marks which choices each specification used.

specification curve

Smallest weight perturbation

The top rank is fragile: a weight change of about 0.037 on pcr overturns it.

Reference levels

Chance baseline

How many tools score above a random method on each metric that declares a chance baseline.

metricchance scoretools above chance
ari0 14 of 14 (100 percent)

Every tool scores above chance on at least one metric.

Noise floor

The top two tools (combat and harmonypy) are separated above the noise floor on at least one metric.

Tool pairs the metric set cannot tell apart above the noise floor:

Reproducibility

Input sha256: c2dbf0368a99c29f80a229f9d2b321d2d2e288b9f0a81706b6e667ba6956acfc

packageversion
python3.12.13
beam0.2.0
numpy2.5.1
scipy1.18.0
pymcdm1.4.0
pyyaml6.0.3
jsonschema4.26.0

The full run manifest, written alongside this report, records the metric card versions and content hashes needed to reproduce the run.