Session 2 · Fri Oct 2 · Book sections 1.1–1.5, 1.9
🧰
By the end of the session, you should be able to evaluate whether a published computational result can be rerun, identify the computing resources a workflow requires, and distinguish an environment from a sandbox or container.
The class will produce a first reproducibility rubric. Each student will begin a repository that can supply evidence for that rubric.
Readings: 1.1–1.5 · Lab: 1.9 workbench setup
| Study | Tested | Outcome |
|---|---|---|
| Begley & Ellis 2012 | 53 landmark cancer studies | 6 confirmed |
| Open Science Collab. 2015 | 100 psychology experiments | 36% replicated (97% originally) |
| Stodden et al. 2018 | 204 computational Science papers | 26% reproduced |
| Stagge et al. 2019 | 360 hydrology articles | 1.6% reproduced |
| Konkol et al. 2019 | R code of 41 geoscience papers | 2 ran cleanly; 46 of 97 figures differ |
| Massonnet et al. 2020 | one Earth system model, two HPC systems | different climates |
Baker 2016 survey, recomputed from the raw data: 64% of 95 Earth and environmental scientists had failed to reproduce another group’s result.
Where reproducibility has been measured, most published computational results could not be regenerated. The geoscience studies report rates comparable to, or lower than, those in other fields.
✗ Hydrology: results reproduced for 1.6% of 360 sampled articles.
Stagge, J. H., Rosenberg, D. E., Abdallah, A. M., Akbar, H., Attallah, N. A., & James, R. (2019). Assessing data availability and research reproducibility in hydrology and water resources. Scientific Data, 6, 190030.
✗ Geophysics: 86% of 200 articles state data availability; 54% provide the data.
Ireland, M., Algarabel, G., Steventon, M., & Munafò, M. (2023). How reproducible and reliable is geophysical research? Seismica, 2(1). doi:10.26443/seismica.v2i1.278
✗ Aftershocks: a two-parameter model matches the AUC of a 13,451-parameter network.
Mignan, A., & Broccardo, M. (2019). One neuron versus deep learning in aftershock prediction. Nature, 574, E1–E3. On DeVries et al. (2018), Nature, 560, 632–634.
✓ Temperature reconstructions: independent methods and new proxies recover the same shape.
Wahl, E. R., & Ammann, C. M. (2007). Climatic Change, 85, 33–69 · PAGES 2k Consortium (2019). Nature Geoscience, 12, 643–649.
A rerun did not establish whether any of these results was correct; it allowed other researchers to examine and test them.
12 minutes in pairs
Then complete two sentences:
Another team could rerun this result because the authors provide ___
Another team would get stuck at ___ because ___
Then propose one observable criterion for the course rubric.
“Uses GitHub” describes a platform. “A fresh clone regenerates Figure 3 from documented commands” states an observable test.
| Same data | Different data | |
|---|---|---|
| Same analysis | Reproducible | Replicable |
| Different analysis | Robust | Generalisable |
Hold the research question fixed. The Turing Way definitions (book 1.1). The National Academies (2019) use the top row; some fields swap the two words. State the operational test, not the label.
Each cell corresponds to a distinct claim. Rerunning a notebook on the original data tests only reproducibility, the upper-left cell.
| Passed | Rules out | Still possible |
|---|---|---|
| Reproducible | missing steps, unbuildable environment | a bug, a leak, a bad metric |
| Robust | dependence on one seed, split, or choice; no baseline | a quirk of this data set |
| Replicable | dependence on one station, region, period | a shared blind spot in the method |
| Generalisable | dependence on one data set and pipeline | nothing is settled for good |
A conclusion gains support as it persists across these tests. A complete report also states the conditions under which the result is expected to fail.
| Dimension | Observable evidence |
|---|---|
| Scientific claim | one result and its success condition are stated precisely |
| Data and provenance | source, version, license, and processing history are recorded |
| Code and workflow | a fresh clone exposes the full path from data to result |
| Environment and compute | dependencies, hardware needs, and run commands are specified |
| Evaluation | split, baseline, metric, and uncertainty match the claim |
| Publication and credit | outputs are citable; people, data, software, and AI help receive credit |
0 missing · 1 present but incomplete · 2 another team can verify it
Rubric criteria should describe inspectable evidence and the scientific claim that the evidence supports.
Before using the rubric for grades, test whether two evaluators interpret it similarly.
Disagreement may reveal ambiguous wording, missing evidence, or a substantive difference in scientific judgment.
Hardware is the physical machinery that executes and stores the work.
| Resource | What it does | Typical bottleneck |
|---|---|---|
| CPU | general instructions; many scientific tasks | serial code or too few cores |
| GPU | many similar operations in parallel | model/data transfer and GPU memory |
| RAM | working memory for running processes | arrays or tables do not fit |
| Storage | persistent files and checkpoints | capacity and read/write speed |
| Network | moves data between systems | bandwidth, latency, access rules |
A computational requirement should identify the relevant resource, its scale, and the evidence used to estimate it.
| CPU | GPU |
|---|---|
| a small number of sophisticated cores | many simpler arithmetic units |
| strong performance for branching and sequential logic | strong performance for uniform parallel operations |
| usually accesses large system memory | usually uses separate, capacity-limited device memory |
| common target for preprocessing and general scientific code | common target for tensor operations and neural-network training |
Moving data between CPU memory and GPU memory has a cost. Small models, irregular algorithms, and slow data pipelines may gain little from a GPU.
| Context | What it means | Good fit | Main tradeoff |
|---|---|---|---|
| Local laptop | hardware you control directly | exploration, small data, offline work | limited memory, storage, and uptime |
| Back-end Linux server | remote shared machine reached through a network | longer jobs, shared data, stable services | permissions and shared resources |
| Cloud server | rented remote compute created on demand | temporary scale, public services, special hardware | cost and data governance |
| HPC system | many managed compute nodes plus a scheduler | large parallel jobs, many experiments | queues, modules, storage policy |
A computing center operates shared hardware, storage, networking, schedulers, and support. Hyak is UW’s HPC system.
| Stage | What happens |
|---|---|
| Login | the user connects to a login node and prepares code, data references, and a job request |
| Resource request | the job specifies cores, memory, GPUs, wall time, and sometimes node type |
| Scheduling | the scheduler waits until a compatible allocation is available |
| Execution | the job runs on one or more compute nodes, usually without an interactive desktop |
| Persistence | logs, checkpoints, and results must be written to approved storage before the allocation ends |
Requesting substantially more resources can increase queue time and allocation cost. Requesting too little can terminate the job.
| Term | Question it answers |
|---|---|
| Operating system | Which processes, files, devices, and users does the machine manage? |
| Computing environment | Which interpreter, packages, libraries, and exact versions run the code? |
| Sandbox | Which files, network services, tools, and actions may this process access? |
| Container | Which operating-system-level dependencies travel with the application? |
A lockfile fixes package versions, a sandbox limits what a process may access, and a container also carries operating-system dependencies.
A reproducible run should preserve enough information to reconstruct both software and execution context:
A published figure does not record the software, data, and parameters that produced it, so that information has to be preserved alongside the figure.
Choose a location, the likely limiting resource, and the rubric evidence you would record.
| Case | Decision prompt |
|---|---|
| 4 GB of GNSS data and exploratory plots | laptop, server, cloud, or HPC? |
| 2 TB of continuous waveforms searched overnight | where should data and computation meet? |
| CNN training on 100,000 satellite tiles | when does a GPU help? |
| restricted collaborator data | what should the sandbox prevent? |
Seven minutes in groups. State the assumptions behind one choice. Another group identifies a condition under which that choice would fail.
Open 1.9 and produce evidence for the rubric:
pixi install for the book environmentMLGEO2026_UWNETID and add pixi.toml plus pixi.lockpixi run command in the READMEThe repository should specify a runnable target, the environment that supports it, and an observable success condition.
Submit three lines before leaving:
HW1 due Mon Oct 12 · Ch 1 quiz opens Tue Oct 6 · Monday begins by scoring your paper and workbench with the revised rubric
ESS 469/569 · Machine Learning in the Geosciences · Autumn 2026