Speaker
Description
Computational reproducibility means others can obtain the same results using a study's code and data. It depends on both being available, documented, and preserved over time. The scientific methods relies on this. However, deposited code is often not executable years later because dependencies change, environments are lost, or model output is too large to archive. Other obstacles are not technical: little academic credit for software work, restrictive licenses, and a lack of staff to prepare and maintain code for release. Machine-learning methods add further problems, such as non-deterministic training, unavailable training data, and proprietary models that change without notice.
This presentation draws on years of editorial work at Geoscientific Model Development, a journal that has required public code and data for model descriptions since 2013. We discuss the requirements enforced, how they have worked or fallen short, and what they demand of authors and research infrastructures. We also present results from using GEMMA4, an open multimodal large language model, to assess the replicability of published studies, compared to human editorial assessment. The main finding is that the replicability of the papers published has largely improved over time. However, when we focus on what is fully replicable, problems persist. We discuss what these results imply for using AI in reproducibility checks.