The point prediction worked. The confidence estimate did not.
A pretrained materials model ranked stable candidates efficiently on a static screening task, while the tested MC-dropout uncertainty signal failed the calibration test needed to justify uncertainty-aware acquisition.
Active learning usually assumes the acquisition function can trade off predicted value and model uncertainty. That only makes sense if the uncertainty score carries information about error.
Greedy mean ranking reached DAF 5.15 and low-weight UCB 5.17 at the tested 5% budget, while MC-dropout σ had Spearman −0.47 correlation with absolute error.
The public benchmark runs five seeds over 18,928 perovskite structures and includes direct calibration diagnostics: σ↔error association, reliability, miscalibration area, and empirical coverage.
This is a static labeled proxy benchmark, not a closed-loop DFT or laboratory discovery campaign. The result is specific to the tested pretrained model, readout-level MC-dropout configuration, dataset, and acquisition setup.