Open reliability benchmark

What this is. A standardized, open table that scores reliability (uncertainty-quantification) methods for single-cell perturbation predictors across datasets and context-transfer settings. Each row is one method measured on one predictor in one transfer setting, with a pretraining-provenance flag so a possible train-test overlap is always visible next to the number. PertEMA is one method in this table, not the whole benchmark.

How to read it. Two metrics are reported. aurc is the area under the risk-coverage curve (mean 1 - Pearson error when you abstain from the least-reliable predictions first), where lower is better and oracle is the best reachable and no_selection is the do-nothing baseline. reliability_spearman is the Spearman correlation between predicted reliability and realized accuracy, where higher is better and the random_feature_control and label_shuffle_control rows should sit near zero. The provenance flag is CLEAN unless a predictor's pretraining may overlap the benchmark data, in which case it reads UNKNOWN-OVERLAP. Gains are modest on the noisy primary data and the clearest positive result is external and cross-cell-line (Replogle K562 to RPE1). See the methods page for the honest scope.

Raw JSON

DatasetTransfer settingPredictorUQ method MetricValue95% CInProvenance

Rows for the pertema method are shown in bold. A blank cell means the value was not recorded for that row. Confidence intervals are half-width 95% intervals where available.