Your reproducible workbench

Session 2 · Fri Oct 2 · Book sections 1.1–1.5, 1.9

🧰

Learning objectives and class products

By the end of the session, you should be able to evaluate whether a published computational result can be rerun, identify the computing resources a workflow requires, and distinguish an environment from a sandbox or container.

The class will produce a first reproducibility rubric. Each student will begin a repository that can supply evidence for that rubric.

Readings: 1.1–1.5 · Lab: 1.9 workbench setup

Measured reproducibility of published research

Study Tested Outcome
Begley & Ellis 2012 53 landmark cancer studies 6 confirmed
Open Science Collab. 2015 100 psychology experiments 36% replicated (97% originally)
Stodden et al. 2018 204 computational Science papers 26% reproduced
Stagge et al. 2019 360 hydrology articles 1.6% reproduced
Konkol et al. 2019 R code of 41 geoscience papers 2 ran cleanly; 46 of 97 figures differ
Massonnet et al. 2020 one Earth system model, two HPC systems different climates

Baker 2016 survey, recomputed from the raw data: 64% of 95 Earth and environmental scientists had failed to reproduce another group’s result.

Where reproducibility has been measured, most published computational results could not be regenerated. The geoscience studies report rates comparable to, or lower than, those in other fields.

This lecture in the literature

✗ Hydrology: results reproduced for 1.6% of 360 sampled articles.

Stagge, J. H., Rosenberg, D. E., Abdallah, A. M., Akbar, H., Attallah, N. A., & James, R. (2019). Assessing data availability and research reproducibility in hydrology and water resources. Scientific Data, 6, 190030.

✗ Geophysics: 86% of 200 articles state data availability; 54% provide the data.

Ireland, M., Algarabel, G., Steventon, M., & Munafò, M. (2023). How reproducible and reliable is geophysical research? Seismica, 2(1). doi:10.26443/seismica.v2i1.278

✗ Aftershocks: a two-parameter model matches the AUC of a 13,451-parameter network.

Mignan, A., & Broccardo, M. (2019). One neuron versus deep learning in aftershock prediction. Nature, 574, E1–E3. On DeVries et al. (2018), Nature, 560, 632–634.

✓ Temperature reconstructions: independent methods and new proxies recover the same shape.

Wahl, E. R., & Ammann, C. M. (2007). Climatic Change, 85, 33–69 · PAGES 2k Consortium (2019). Nature Geoscience, 12, 643–649.

A rerun did not establish whether any of these results was correct; it allowed other researchers to examine and test them.

Paper sprint completion

12 minutes in pairs

  1. Confirm the paper, DOI, and scientific claim
  2. Locate the data and its access conditions
  3. Locate code, license, and dependency record
  4. Identify the exact figure or table you would rerun
  5. Record the first point where the evidence trail breaks

Then complete two sentences:

Another team could rerun this result because the authors provide ___

Another team would get stuck at ___ because ___

Then propose one observable criterion for the course rubric.

“Uses GitHub” describes a platform. “A fresh clone regenerates Figure 3 from documented commands” states an observable test.

Same or different data, same or different analysis

Same data Different data
Same analysis Reproducible Replicable
Different analysis Robust Generalisable

Hold the research question fixed. The Turing Way definitions (book 1.1). The National Academies (2019) use the top row; some fields swap the two words. State the operational test, not the label.

Each cell corresponds to a distinct claim. Rerunning a notebook on the original data tests only reproducibility, the upper-left cell.

What each test establishes

Passed Rules out Still possible
Reproducible missing steps, unbuildable environment a bug, a leak, a bad metric
Robust dependence on one seed, split, or choice; no baseline a quirk of this data set
Replicable dependence on one station, region, period a shared blind spot in the method
Generalisable dependence on one data set and pipeline nothing is settled for good

A conclusion gains support as it persists across these tests. A complete report also states the conditions under which the result is expected to fail.

Class rubric v0.1

Dimension Observable evidence
Scientific claim one result and its success condition are stated precisely
Data and provenance source, version, license, and processing history are recorded
Code and workflow a fresh clone exposes the full path from data to result
Environment and compute dependencies, hardware needs, and run commands are specified
Evaluation split, baseline, metric, and uncertainty match the claim
Publication and credit outputs are citable; people, data, software, and AI help receive credit

0 missing · 1 present but incomplete · 2 another team can verify it

Rubric criteria should describe inspectable evidence and the scientific claim that the evidence supports.

Rubric calibration

Before using the rubric for grades, test whether two evaluators interpret it similarly.

  1. Select one paper or repository
  2. Score each dimension independently
  3. Compare scores and identify the source of disagreement
  4. Revise the criterion or add an example

Disagreement may reveal ambiguous wording, missing evidence, or a substantive difference in scientific judgment.

Hardware resources

Hardware is the physical machinery that executes and stores the work.

Resource What it does Typical bottleneck
CPU general instructions; many scientific tasks serial code or too few cores
GPU many similar operations in parallel model/data transfer and GPU memory
RAM working memory for running processes arrays or tables do not fit
Storage persistent files and checkpoints capacity and read/write speed
Network moves data between systems bandwidth, latency, access rules

A computational requirement should identify the relevant resource, its scale, and the evidence used to estimate it.

CPU and GPU execution models

CPU GPU
a small number of sophisticated cores many simpler arithmetic units
strong performance for branching and sequential logic strong performance for uniform parallel operations
usually accesses large system memory usually uses separate, capacity-limited device memory
common target for preprocessing and general scientific code common target for tensor operations and neural-network training

Moving data between CPU memory and GPU memory has a cost. Small models, irregular algorithms, and slow data pipelines may gain little from a GPU.

Where computation runs

Context What it means Good fit Main tradeoff
Local laptop hardware you control directly exploration, small data, offline work limited memory, storage, and uptime
Back-end Linux server remote shared machine reached through a network longer jobs, shared data, stable services permissions and shared resources
Cloud server rented remote compute created on demand temporary scale, public services, special hardware cost and data governance
HPC system many managed compute nodes plus a scheduler large parallel jobs, many experiments queues, modules, storage policy

A computing center operates shared hardware, storage, networking, schedulers, and support. Hyak is UW’s HPC system.

How an HPC job reaches hardware

Stage What happens
Login the user connects to a login node and prepares code, data references, and a job request
Resource request the job specifies cores, memory, GPUs, wall time, and sometimes node type
Scheduling the scheduler waits until a compatible allocation is available
Execution the job runs on one or more compute nodes, usually without an interactive desktop
Persistence logs, checkpoints, and results must be written to approved storage before the allocation ends

Requesting substantially more resources can increase queue time and allocation cost. Requesting too little can terminate the job.

Software boundaries

Term Question it answers
Operating system Which processes, files, devices, and users does the machine manage?
Computing environment Which interpreter, packages, libraries, and exact versions run the code?
Sandbox Which files, network services, tools, and actions may this process access?
Container Which operating-system-level dependencies travel with the application?

A lockfile fixes package versions, a sandbox limits what a process may access, and a container also carries operating-system dependencies.

Minimum computational provenance

A reproducible run should preserve enough information to reconstruct both software and execution context:

  • source revision and uncommitted changes, if any
  • data identifiers, versions, and preprocessing parameters
  • environment specification and lockfile
  • command, configuration, random seed, and requested resources
  • logs, warnings, output checksums, and the result selected for reporting

A published figure does not record the software, data, and parameters that produced it, so that information has to be preserved alongside the figure.

Compute choice activity

Choose a location, the likely limiting resource, and the rubric evidence you would record.

Case Decision prompt
4 GB of GNSS data and exploratory plots laptop, server, cloud, or HPC?
2 TB of continuous waveforms searched overnight where should data and computation meet?
CNN training on 100,000 satellite tiles when does a GPU help?
restricted collaborator data what should the sandbox prevent?

Seven minutes in groups. State the assumptions behind one choice. Another group identifies a condition under which that choice would fail.

Workbench lab

Open 1.9 and produce evidence for the rubric:

  1. Start pixi install for the book environment
  2. Create MLGEO2026_UWNETID and add pixi.toml plus pixi.lock
  3. Record the tested operating system and machine type in the README
  4. Add one smoke test and name its pixi run command in the README
  5. Review the diff, commit with a real message, and push

The repository should specify a runnable target, the environment that supports it, and an observable success condition.

Exit ticket: input for rubric v0.2

Submit three lines before leaving:

  • one rubric criterion you would keep
  • one criterion you would sharpen, with observable evidence
  • one computing choice your project may require

HW1 due Mon Oct 12 · Ch 1 quiz opens Tue Oct 6 · Monday begins by scoring your paper and workbench with the revised rubric