Reproducible work with agents

Session 3 · Mon Oct 5 · Book sections 1.8 + 6.1

🤖

Learning objective

Evaluate an AI-assisted workflow as a scientific process: identify which decisions belong to the model, which operations belong to tools, which access constraints come from the sandbox, and which evidence supports the reported result.

Rubric checkpoint

Score the paper from your sprint using Friday’s six dimensions:

scientific claim · data · code/workflow · environment/compute · evaluation · publication/credit

For each dimension, assign:

  • 0 missing
  • 1 present but incomplete
  • 2 another team can verify it

A score whose interpretation is uncertain indicates where the rubric needs a clearer criterion or additional evidence.

From available artifacts to an executable claim

Code availability answers only one question: Can the files be inspected?

A computational claim also depends on:

  • the data object and preprocessing state
  • the software environment and external services
  • the order of operations and configuration
  • the evaluation target and tolerance for agreement
  • the hardware or numerical behavior that may affect the result

Agent, tools, environment, sandbox

Agent = language model + tools + a loop over observations

Layer Role in an agent run
Tools actions such as reading files, running Python, or querying a catalog
Environment exact Python, packages, libraries, and command-line programs
Sandbox allowed files, network destinations, tools, and write locations
Hardware CPU, GPU, memory, storage, and network underneath the run

The environment determines which software can execute. The sandbox constrains which resources that software may reach.

The agent loop

Each iteration contains three distinct sources of scientific responsibility:

Stage Example Question for the reviewer
Specification “Estimate the velocity of station P395” Were the target, units, interval, and success criterion defined?
Execution the agent reads data and runs regression code Which tools, data, environment, and permissions shaped the operation?
Interpretation the agent explains the estimated trend Does the conclusion follow from the computed output and domain assumptions?

A complete execution log does not compensate for a scientific question that was underspecified at the start.

Where errors enter

Failure mode What can happen Evidence or countermeasure
Attention important context receives little weight name required files; test cited details
Agreement the model follows a user’s wrong framing request alternatives; test opposing hypotheses
Judgment it chooses the wrong data, method, or interpretation specify success criteria; compare with ground truth
Hallucination fluent claims or citations have no source resolve citations; recompute quantitative claims

The rubric evaluates evidence and decision quality rather than confidence, fluency, or length of explanation.

Long context and incomplete evidence

An agent may receive an entire repository while using only a small fraction of it effectively.

  • Retrieval and attention determine which passages influence the next action
  • Files listed in a trace may not have influenced the conclusion equally
  • Summaries can omit exceptions, warnings, and negative results
  • A long reasoning trace records a proposed path, not a guarantee of causal faithfulness

An evaluation should test specific claims against source files and outputs rather than infer completeness from trace length.

Numerical calibration protocol

Before asking the assistant, define the reference calculation:

  • station component and date interval
  • treatment of missing data, offsets, and seasonality
  • regression model, units, and reported precision
  • tolerance for agreement and an independently executed implementation

Then record whether the assistant accessed the intended data, executed code, exposed warnings, and produced a value consistent with the reference.

Citation calibration protocol

For each suggested paper, verify four properties independently:

Property Verification source
Existence DOI resolver, publisher, or trusted bibliographic index
Bibliographic identity title, authors, year, and venue match
Relevance abstract or full text addresses the stated scientific question
Accessibility of evidence data, code, supplement, and licensing can be located

A real citation can still fail the relevance or evidence test.

Calibration experiment

20 minutes in pairs

  1. Give an assistant book/Chapter1-GettingStarted/data/gps_P395_relative_position.csv and ask for the east velocity of P395 in mm/yr
  2. Record whether it runs code or answers from text
  3. Recompute the value independently and compare
  4. Ask for three ML papers in your subfield
  5. Resolve every DOI and count unsupported citations

The class will report denominators, not anecdotes: number of attempts, valid tool executions, numerically consistent answers, and verified citations.

Course AI modes

🔒 By hand

The learning objective requires unaided execution or interpretation. These exercises establish the skills needed to inspect an assistant’s work.

🤝 Assistant allowed

The assistant may contribute to the workflow. The submission must disclose its contribution, preserve accepted artifacts, and document independent verification.

Students remain responsible for explaining the scientific and computational decisions represented in submitted work.

Sandbox boundaries and research data

Before giving an agent access to research material, classify both the data and the action.

Question Examples of relevant evidence
May the service receive these data? consent, data-use agreement, export controls, institutional policy
May the process contact the network? approved domains, API terms, firewall or sandbox configuration
May the process modify or delete files? write scope, version control, backups, confirmation policy
May outputs leave the environment? disclosure risk, licensing, embargo, human review

Credentials should enter through approved secret-management mechanisms rather than prompts, notebooks, or committed files.

AI-assisted rubric extension

Add these criteria to the class rubric:

Dimension Observable evidence
Tool disclosure tool and its contribution are named in the submission or commit
Human review the student can explain the accepted code and scientific choices
Independent verification numbers, citations, and API claims are checked outside the model output
Access boundaries sensitive data and credentials stay outside unapproved tools and sandboxes
Run record accepted code, environment, inputs, and outputs are preserved in the repository

0 missing · 1 present but incomplete · 2 another team can verify it

Same scale as v0.1. A claim without inspectable evidence, such as “verified the citations”, scores 1.

Peer rerun

Exchange repository URLs with another pair.

  1. Read only the README first
  2. Clone into a fresh directory
  3. Build the environment from the lockfile
  4. Run the smoke test the README names (pixi run smoke)
  5. Compare one output and log the first obstacle

The peer records the first obstacle before requesting help, because that obstacle provides evidence about documentation, environment, access, or computational assumptions.

Rubric v0.2

As a class, choose:

  • one criterion the paper review exposed
  • one criterion the calibration experiment exposed
  • one criterion the peer rerun exposed
  • the evidence required for a score of 2

Ch 1 quiz opens Tue Oct 6 and closes Thu Oct 8 · Wednesday: data modalities, formats, and the class project landscape