Open reliability benchmark
What this is. A standardized, open table that scores reliability (uncertainty-quantification) methods for single-cell perturbation predictors across datasets and context-transfer settings. Each row is one method measured on one predictor in one transfer setting, with a pretraining-provenance flag so a possible train-test overlap is always visible next to the number. PertEMA is one method in this table, not the whole benchmark.
How to read it. Two metrics are reported. aurc is the
area under the risk-coverage curve (mean 1 - Pearson error when you abstain from the least-reliable
predictions first), where lower is better and oracle is the best reachable and
no_selection is the do-nothing baseline. reliability_spearman is the Spearman
correlation between predicted reliability and realized accuracy, where higher is better and the
random_feature_control and label_shuffle_control rows should sit near zero. The
provenance flag is CLEAN unless a predictor's pretraining may overlap the
benchmark data, in which case it reads UNKNOWN-OVERLAP. Gains are modest on the noisy primary
data and the clearest positive result is external and cross-cell-line (Replogle K562 to RPE1). See the
methods page for the honest scope.
| Dataset | Transfer setting | Predictor | UQ method | Metric | Value | 95% CI | n | Provenance |
|---|
Rows for the pertema method are shown in bold. A blank cell means the value was
not recorded for that row. Confidence intervals are half-width 95% intervals where available.