Fig1. of Mancinelli/Bach (2026)
- Comparison of the four measurement types examined in our study, across successive stages of dataset integration (ordered by acquisition date). A. Log(Bayes factor) for each measurement. B. Retrodictive validity scores, expressed as Pearson’s r, rather than the Cohen’s d values reported in Table 1. Both metrics are derived from t-statistics calculated on the accrued data. C. Differences in log Bayes factor between the highest-performing measure and the others, offering a focused comparison. Line width in C indicates the extent of subject overlap, serving as an index of the confidence in each comparison. In panels A and B, line width and alpha reflect the incremental sample size as datasets are accumulated. For clarity, the exact numbers of overlapping subjects are as follows: HPR-PSR: 114; HPR-RAR: 149; HPR-SCR: 175
© F. Mancinelli, D. Bach
Alle Bilder in Originalgröße herunterladen
Der Abdruck im Zusammenhang mit der Nachricht ist kostenlos, dabei ist der angegebene Bildautor zu nennen.
Experiment-based calibration is an emerging approach for measurement validation in the behavioural sciences. It allows comparing multiple measurement methods by how well they reproduce a known experimental manipulation, providing insight into their measurement accuracy.
Calibration entails questions unparalleled in classical validation approaches. The first is about inference: when should we conclude that one measurement method is truly more accurate than another? The second is about decisions: when should we decide that a method merits the investment of changing a measurement system? In this note, we review these questions in the context of the statistical challenges that arise in a calibration process: a potentially large and a priori unknown number of measurement methods; a requirement to integrate evidence across multiple calibration samples; and a possibility that some methods may not be available for all samples.