Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

🖥️ Lecture slides — Session 01 (Wed Sep 30)

🖥️ Lecture slides — Session 02 (Fri Oct 2)

Open, reproducible science is the working standard for this course. Every homework and the final project are graded partly on whether someone else could rerun your analysis and get your result. This page explains what that means in practice: open science, the FAIR principles, licenses, data citation, preprints, and why reproducibility matters more now that AI writes a lot of the code.

Two complements to this lesson (not substitutes for it):

Lecture SlidesPresentation recording

Why a lecture on reproducibility?

Because most published computational results cannot be rerun. That holds across science and in the geosciences specifically, and it has been measured.

Across science. The best-known reproducibility studies come from outside our field:

StudyWhat was testedOutcome
Begley & Ellis (2012), Nature53 “landmark” preclinical cancer papers, repeated at Amgenfindings confirmed in 6 (11%)
Open Science Collaboration (2015), Science100 psychology experiments, rerun with new participants97% of originals significant; 36% of replications significant; effects half as large
Baker (2016), Nature survey1,576 researchers asked about their own experience72% had failed to reproduce someone else’s result, 56% their own
Stodden, Seiler & Ma (2018), PNAS204 computational papers in Science after its 2011 code-sharing policyartifacts obtained for 44%; results reproduced for 26%
Kapoor & Narayanan (2023), Patternspublished ML-based science in 17 fieldsdata leakage found in 294 papers

Baker’s survey numbers are recomputed from the raw survey data that Nature published on figshare (doi:10.6084/m9.figshare.3394951). Earth and environmental scientists made up 95 of the 1,576 respondents: 61 of them (64%) had failed to reproduce another group’s result and 39 (41%) their own. We are no exception.

In the geosciences. Reproducibility studies of the geoscience literature report similar results:

StudyCorpusOutcome
Stagge et al. (2019), Scientific Data360 of 1,989 articles from six hydrology and water-resources journals, 2017results reproduced for 1.6% of the articles tested; 95% confidence interval 0.6–6.8% for all 1,989
Nüst et al. (2018), PeerJ32 award-nominated GIScience conference papers (AGILE), 2010–2017none reached the top reproducibility level in any category; 19 of 32 at the lowest level for data
Konkol, Kray & Pfeiffer (2019), IJGISR code of 41 open-access papers (31 Copernicus geoscience articles from 2016–2017, plus the 10 most-cited Journal of Statistical Software R-package papers), rerun in a clean containercode ran without issues for 2; 33 ran after fixes, 2 partially, 4 not at all; 15 required contacting the authors; 173 issues in 39 papers; 46 of 97 regenerated figures differed from the published ones in content
Ireland et al. (2023), Seismica200 geophysics articles and the policies of 20 journals~86% carried a data availability statement, but the original data were accessible for 54% and the software was named in 49%
Massonnet et al. (2020), Geoscientific Model Developmentone Earth system model (EC-Earth3) run on different high-performance computing (HPC) systemsthe older model version gave statistically different climates on different machines

The Konkol et al. survey of 146 geoscientists recruited at the 2016 European Geosciences Union General Assembly shows the same gap from the authors’ side: 49% said they often or always publish results others can recompute, but 33% linked to their data and 12% to their code. Massonnet et al. conclude that “the default assumption should be that ESMs are not replicable under changes in the HPC environment, until proven otherwise.” A data availability statement is not the data, and a named model is not the run. Hutton et al. (2016) put the consequence bluntly in the title of a Water Resources Research commentary: “Most computational hydrology is not reproducible, so is it really science?”

Three geoscience cases worth knowing in detail.

The common thread is that reproducibility did not prove any of these results right or wrong. It made them checkable. The Turing Way lists “being reproducible does not mean the answer is right” as one of its barriers to reproducible research, and that is the point of the next section.

Reproducible, replicable, robust, generalisable

The Turing Way Community, 2025 organizes the vocabulary as a two-by-two table. Hold the research question fixed; then ask whether the data and the analysis are the same as the original or different.

Same dataDifferent data
Same analysisReproducible: the same steps on the same data give the same answerReplicable: the same analysis on different data gives a qualitatively similar answer
Different analysisRobust: a different workflow on the same data gives a qualitatively similar answerGeneralisable: the answer holds across new data and new analyses
The four combinations of same or different data and analysis. The Turing Way project illustration by Scriberia. Used under a CC-BY 4.0 licence. DOI: 10.5281/zenodo.3332807.

Figure 1:The four combinations of same or different data and analysis. The Turing Way project illustration by Scriberia. Used under a CC-BY 4.0 licence. DOI: The Turing Way Community & Scriberia (2024).

Read in machine-learning terms, for a model of GNSS displacement or seismic detection:

CellWhat changesWhat you would do in this course
Reproduciblenothingfresh clone, locked environment, documented command regenerates the figure
Robustthe analysischange the random seed, the train/validation/test split design, the preprocessing, or the model family, and compare with a simple baseline
Replicablethe dataapply the same pipeline to other stations, another region, or a later time period
Generalisablebothindependent groups, data and methods converge on the same conclusion

The phrase has a geophysical origin: Claerbout & Karrenbach (1992, SEG Annual Meeting), at the Stanford Exploration Project, argued that a seismic-imaging paper should ship with the programs and data that regenerate its figures. Other communities use different words. The National Academies report Reproducibility and Replicability in Science Engineering et al., 2019, used in Chapter 5.1, defines reproducibility as “obtaining consistent results using the same input data; computational steps, methods, and code; and conditions of analysis,” and replicability as “obtaining consistent results across studies aimed at answering the same scientific question, each of which has obtained its own data.” Those match the top row of the table. Some fields swap the two words entirely. When you write an assessment, state the operational test (same or different data, same or different analysis) rather than relying on the label.

What each test establishes

A strong result is one that keeps its conclusion as you move across the table. Each cell rules out a different way of being wrong:

Test passedWhat it rules outWhat it does not rule out
Reproduciblemissing steps, undocumented choices, an environment nobody else can builda bug, a leak, a bad metric — these reproduce perfectly
Robusta conclusion that depends on one seed, one split, one preprocessing choice, or a missing baselinesomething peculiar to this data set
Replicablea conclusion that depends on this station, this region, or this perioda shared blind spot in the method
Generalisabledependence on one data set and one pipelinenothing is ever settled for good; it is the best available evidence

So a strong result in this course:

  1. reruns from a fresh clone (reproducible);
  2. beats a simple baseline and keeps its conclusion under reasonable changes to seed, split design and preprocessing, with the spread reported as an uncertainty (robust);
  3. holds on data the model never saw, chosen to differ in location or time from the training data (replicable);
  4. states the conditions under which it is expected to fail.

The aftershock case above passed step 1 and failed step 2. The hockey stick passed all four, over two decades, through work by other groups. Homework in this course is graded on step 1. Hidden-test leaderboards (Chapter 3) and the final project push toward steps 2 to 4.

What is open science?

Open science is the practice of making the products of research — data, code, methods, and papers — available for others to inspect, reuse, and build on. In the geosciences this is not an abstract ideal. Most of the data we use in this book (seismic waveforms, GNSS positions, satellite imagery, climate reanalyses) exists because agencies and researchers published it openly. When you publish your own work the same way, you close the loop.

Open science has several components:

Reproducibility is the thread through all four: without open data, code and methods, nobody can occupy any cell of the table above except the original team.

FAIR principles

The FAIR principles (Wilkinson et al., 2016) describe what makes data useful to others:

When you assemble an AI-ready data set in Chapter 2, you will apply these to your own products: archive the data, document the provenance, attach a license, get a DOI.

Geoscience data is often data about place, and data about place can carry obligations that a license file does not capture — local regulations, community agreements, or national data policies. Before assuming open publication is appropriate, check the terms under which the data were collected.

Licenses

Without a license, others legally cannot reuse or modify your work, even if it sits in a public repository. “Public” does not mean “licensed.” So every repository you create in this course carries a license file.

Software licenses. The common open source choices:

choosealicense.com walks you through the choice. For coursework, MIT is a sensible default. The Turing Way licensing chapter has a longer discussion.

Data and text licenses. Software licenses are written for code and fit data poorly. For data sets, documentation, and figures, use Creative Commons:

A repository that contains both code and data can carry two licenses — say, MIT for the code and CC-BY for the data — with the split stated in the README.

What a reusable repository contains

Beyond the license, a repository that others (including future you) can use has:

The Software Carpentries Intermediate Research Software Development lessons cover this in depth Nenadic et al., 2022.

Data citation and DOIs

A DOI (Digital Object Identifier) is a persistent identifier that resolves to a data set, paper, or software release forever, even if the hosting URL changes. Data with a DOI can be cited like a paper, which is how data producers get credit — and why archives require you to cite the data you use. When this book downloads GNSS data from the Nevada Geodetic Laboratory in 1.7, we cite Blewitt et al. (2018). Do the same for every data set in your project.

Zenodo, operated by CERN, mints DOIs for free and integrates with GitHub: link your repository once, and every GitHub release is archived on Zenodo with its own DOI automatically. This is how you make your final project citable. GitHub documents the workflow here. Domain-specific archives (PANGAEA, EarthScope, national data centers) serve the same role for observational data.

Preprints

A preprint is the manuscript posted publicly before (or during) peer review. In the Earth sciences, the main servers are EarthArXiv and ESS Open Archive (ESSOAr). Preprints make results available months to years earlier than journals, carry DOIs, and are citable. Most geoscience journals permit preprinting; check the journal’s policy. Reading preprints is also part of the weekly literature work in this course, with the standard caution: a preprint has not yet been peer reviewed, so read it with the same critical eye you will learn to apply to AI output.

Reproducibility in the AI era

You might expect AI assistants to make reproducibility concerns obsolete: if the code can be regenerated on demand, why archive it? The opposite is true.

Chapter 5 develops this into a full workflow: environments as lockfiles, data versioning, and experiment tracking. For now, the rule is simple: everything that produced a figure or a number in your project is committed, licensed, and rerunnable.

Further reading

Cases and studies cited above

References
  1. Community, T. T. W. (2025). The Turing Way: A handbook for reproducible, ethical and collaborative research (1.2.3). Zenodo. 10.5281/zenodo.15213042
  2. The Turing Way Community, & Scriberia. (2024). Illustrations from The Turing Way: Shared under CC-BY 4.0 for reuse. Zenodo. 10.5281/ZENODO.3332807
  3. Engineering, M., on Behavioral, B., National Academies of Sciences, Engineering, Medicine, & others. (2019). Confidence in Science. In Reproducibility and Replicability in Science. National Academies Press (US).
  4. Nenadic, A., Crouch, S., Graham, J., Mangham, S., Laird, J., & Robinson, M. (2022). carpentries-incubator/python-intermediate-development: beta (beta). 10.5281/zenodo.6532057
  5. Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1). 10.1038/sdata.2016.18
  6. Begley, C. G., & Ellis, L. M. (2012). Raise standards for preclinical cancer research. Nature, 483(7391), 531–533. 10.1038/483531a
  7. Estimating the reproducibility of psychological science. (2015). Science, 349(6251). 10.1126/science.aac4716
  8. Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533(7604), 452–454. 10.1038/533452a
  9. Stodden, V., Seiler, J., & Ma, Z. (2018). An empirical analysis of journal policy effectiveness for computational reproducibility. Proceedings of the National Academy of Sciences, 115(11), 2584–2589. 10.1073/pnas.1708290115
  10. Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. 10.1016/j.patter.2023.100804
  11. Stagge, J. H., Rosenberg, D. E., Abdallah, A. M., Akbar, H., Attallah, N. A., & James, R. (2019). Assessing data availability and research reproducibility in hydrology and water resources. Scientific Data, 6(1). 10.1038/sdata.2019.30
  12. Hutton, C., Wagener, T., Freer, J., Han, D., Duffy, C., & Arheimer, B. (2016). Most computational hydrology is not reproducible, so is it really science?: REPRODUCIBLE COMPUTATIONAL HYDROLOGY. Water Resources Research, 52(10), 7548–7555. 10.1002/2016wr019285
  13. Claerbout, J. F., & Karrenbach, M. (1992). Electronic documents give reproducible research a new meaning. SEG Technical Program Expanded Abstracts 1992, 1, 601–604. 10.1190/1.1822162
  14. Nüst, D., Granell, C., Hofer, B., Konkol, M., Ostermann, F. O., Sileryte, R., & Cerutti, V. (2018). Reproducible research and GIScience: an evaluation using AGILE conference papers. PeerJ, 6, e5072. 10.7717/peerj.5072
  15. Konkol, M., Kray, C., & Pfeiffer, M. (2018). Computational reproducibility in geoscientific papers: Insights from a series of studies with geoscientists and a reproduction study. International Journal of Geographical Information Science, 33(2), 408–429. 10.1080/13658816.2018.1508687