scieee AI-readable full text Open interactive document viewer

When Measures Become Targets: Lessons from Open Science and Machine Learning on the Fragility of Reform

Herrmann, Moritz

Full text

When Measures Become Targets: Lessons from Open Science and Machine Learning on the Fragility of Reform Moritz Herrmann1,2 1LMU Munich 2Munich Center for Machine Learning Insights from methodological ML research Current common but incomplete understanding of empirical research in ML leads to non-replicable and unreliable findings. One of the "main evidence gaps around current AI capabilities" identified in the International AI Safety Report (Bengio et al., 2025) Problem 1: Biased experiments and lack of scrutiny Most method comparisons are carried out as part of a paper proposing a new method and are usually biased in favor of the new method. Problem 2: Bias towards certain types of research There is a bias towards formal proofs & application improvements, while purely experimental research is much less incentivized. Problem 3: Conceptual clarity and operationalization A lack of 1) clarity about important concepts in ML research and 2) clear operationalization of experiments harms the validity of results. Scientifically, calls into question progress in the field. "[N]on-reproducible single occurrences are of no significance to science".a Practically, jeopardizes applied researchers’ trust in ML. May discourage applying ML methods, even though these can be beneficial. A long line of literature warned against this situation. Langley: Machine Learning as an experimental science Hooker: Testing heuristics: We have it all wrong 1995 Hand: Classifier technology and the illusion of progress Drummond: Machine Learning as an experimental science (revisited) Sculley et al.: Winner’s curse? On pace, progress, and empirical rigor McGeoch; Johnson 1988 2006 2018 2002 Drummond 2009 Drummond & Japkowicz 2010 Lones; Trosten Liao et al. Christodoulou et al.; Raff 2019 Henderson et al.; Melis et al.; Lucic et al. 2010 2019 2018 2021 2022 2023 2021 2022 2023 Mateus et al.; McElfresh et al.; Kapoor & Narayanan Ferrari Dacrema et al.; Marie et al.; Narang et al. Elor & Averbuch-Elor; van den Goorbergh et al. Mohammadmahdi et al.: Reporting bias when using real data sets to analyze classification performance WarningsEvidence Bouthillier et al. ML research For details & the list of references: Herrmann et al. (2024). Position: Why We Must Rethink Empirical Research in Machine Learning. Lessons from Open Science and ML The situation in ML research has a resemblance to replication crisis in applied research, with two striking similarities 1Neglect of epistemic foundations and warnings about it Applied research Statistical testing: Decades of dispute and warnings! 1942 1955 1994 2019 2Measures become targets Applied research p-value Measure of evidence becomes Indicator of scientific quality →Statistical significance becomes Target in (social) decision-making (publication) →p-hacking ML research Accuracy (w.l.o.g) Measure of prediction performance becomes Indicator of scientific quality →State-of-the-art (SOTA) performance becomes Target in (social) decision-making (publication) →“SOTA-hacking” (Gencoglu et al., 2019) Goodhart’s lawa “When a measure becomes a target, it ceases to be a good measure.” Bengio et al. (2025). International AI Safety Report. arXiv preprint arXiv:2501.17805 p. 45. Herrmann et al. (2024). Position: Why We Must Rethink Empirical Research in Machine Learning. Proceedings of the 41st International Conference on Machine Learning 235:18228-18247. LINK. aStrathern, M. (1997). ’Improving ratings’: audit in the British University system. European Review, 5(3), 305-321. LINK. p. 308. Gencoglu et al. (2019). HARK Side of Deep Learning-From Grad Student Descent to Automated Machine Learning. arXiv preprint arXiv:1904.07633 Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. LINK. p. 85. Syed, M. (2023). Some data indicating that editors and reviewers do not check preregistrations during the review process. PsyArXiv preprint LINK. p. 1. Klonsky, E. D. (2025). Klonsky, E. David. "Campbell’s law explains the replication crisis: Pre-registration badges are history repeating. Assessment 32.2, 224-234. LINK. p. 1. Consequences for Reform Reform efforts focus a lot on practices, in particular replacing quality indicators. Campbell’s law suggests this is no remedy. Campbell’s law “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.” Also applies to "alternative" quality indicators! Example: Preregistration “Another useful tool has been converted into an indicator of strong science and a goal in and of itself. [...] For example, there is already evidence that papers seeking PRBs [pre-registration badges] routinely violate the rules and spirit of pre-registration.” Survey of 201 articles from PLOS journals “At the editor/reviewer level things look much worse, with only 18% mentioning preregistration, 5% report-ing accessing the preregistration, and 3% discussing the relation between the preregistration and the manuscript.” Specifically problematic, as PreReg is not universally applicable. 1High potential for implicit bias for/against certain research types. 2E.g., PreReg is not well applicable to exploratory research, including some types of empirical ML. All Reform is Fragile •There are no one-size-fits-all solutions to the replication crisis. •We should not uncritically prescribe specific modes of operation. •Focus on epistemic aspects is as important as changing practices. Contact Bmo[email protected]