Skip to content

Upstream-delta register: migrate what the lectures use, review data updates after the migration #39

Description

@mmcky

Register of datasets whose committed copy differs from what upstream publishes today, so the migration can proceed without stopping to reconcile each one, and the deltas get reviewed together once it completes.

The policy this implements

A migration moves bytes; it does not update them. The copy that lands in this repo is the copy the lectures already consume, validated byte-identical in the repoint PR. That is what makes a repoint safe to merge: it provably cannot change a single figure.

Adopting a newer upstream vintage is a different change with a different risk profile — it does change lecture output, needs figures re-reviewed, and is an author-facing decision rather than an infrastructure one. Conflating the two would turn every repoint into a content review and stall the programme.

So when a migration discovers a delta against upstream:

  1. Migrate what the lectures use, unchanged, with the byte-compare gate as normal.
  2. Record the delta here, with enough detail that resolving it later needs no re-investigation.
  3. Review the register once the migration completes and decide each case on its merits: adopt the new vintage, keep the frozen copy deliberately, or reconcile a local edit.

Recording the delta in the dataset's manifest (integrity.upstream) is what makes this honest rather than a deferral — the gap is visible in the generated catalog from the day it is found, not filed away in an issue nobody reads.

Two kinds of delta

They look similar and need opposite responses, so the register distinguishes them.

Upstream moved. A newer vintage exists; our copy is an older one. Resolving means deciding whether to adopt it — and per AGENTS.md, a new vintage gets a new filename rather than replacing the old one, so consumers opt in and existing figures stay valid.

Our copy diverges. Upstream is unchanged but the committed file was modified, or was constructed by a process we cannot reproduce. Resolving means reconciling the edit — usually by recording it precisely enough to reapply, or by rebuilding from upstream and accepting the output change.

Register

mpd2020.xlsx — our copy diverges (three header labels)

Found while writing its manifest in #38.

Cell on Regional data Ours Upstream
row0 col1 gdppc_2011 GDP pc 2011 prices
row0 col9 pop Population
row0 col18 gdppc_2011 (empty)

Every data value matches, including all 21,683 rows of Full data. Only these three labels differ.

The edits are ours, established rather than assumed. The Internet Archive holds the upstream file with one unchanging content digest (4OWZWHTE5HXBBF4XLCCQTGDXY4HXOGOK) across every snapshot from 2021-01-10 to 2026-01-02, and a copy fetched 2026-08-06 matches it exactly — so upstream was byte-stable two years before these bytes were committed to lecture-python-intro on 2023-03-23.

Why this cannot simply be "fixed". long_run_growth reads that sheet with header=(0,1,2). The renames are load-bearing: replacing the file with a pristine upstream copy would silently change what the lecture plots, with no error raised. Any resolution has to move the lecture and the file together.

Options when reviewed: keep the edited copy and document it as the intended state (cheapest, and what the manifest records today); or restore the upstream file and adapt the lecture's column handling, which makes the file genuinely verbatim and removes a permanent local patch. There is no urgency either way — the current state is correct and documented.

life-expectancy-vs-gdp-per-capita.csv — upstream moved (a new vintage, and it is not a drop-in)

Found while writing its manifest in #74. Measured against the live grapher endpoint on 2026-08-12: HTTP 200, 1,414,428 B against our 2,059,709 B.

Four of eight columns are renamed. Only Entity, Code, Year and GDP per capita are stable.

Ours Today's export
Life expectancy at birth (historical) Life expectancy at birth
Population (historical estimates) Population
Continent World region according to OWID
417485-annotations — position 6 GDP per capita (Annotations) — position 8

The annotations column is renamed and moved, not dropped. That correction matters, because "drops 417485-annotations" is how this delta had been described in the work plan, and it points at the wrong failure. Coverage changes too: 62,156 → 29,912 rows, Year -10000..2021 → 1..2023, entities 317 → 276. Today's metadata cites Maddison Project Database 2023, HMD 2025 and UN WPP 2024, all of which postdate this export.

Why a refresh is a breaking change, in two different ways. The loud one: all four consuming lectures hardcode the old label and pass it as usecols, so today's export raises ValueError: Usecols do not match columns, columns expected but not found: ['Life expectancy at birth (historical)']. Note it names only the life-expectancy column — the annotations column is not in usecols, so the annotations change is not what breaks the read.

The quiet one is the reason this belongs in a register rather than a bug tracker. After relabelling usecols, dropna() yields 13,239 rows against our 12,445, so the sentence "dropped the number of rows in our DataFrame from 62156 to 12445" — hardcoded in the prose of all four repos — becomes wrong. And year == 2018 returns 166 rows in both vintages, so the final scatter and OLS fit would look superficially identical while every life-expectancy value shifted. A refresh that fixed only the ValueError would ship a silently different figure.

upstream-moved, established rather than assumed. The committed file's column names are OWID's own older schema — the (historical) suffixes and the numeric 417485-annotations variable id are not names anyone here would invent — and its Year ceiling of 2021 is consistent with an export taken before the source releases the current metadata cites. Our 62,156 rows are a superset of today's 29,912 in coverage, which rules out a local truncation. No QuantEcon edit is implicated.

Options when reviewed. Per AGENTS.md "Corrections vs vintages", adopting the new vintage means a new filename so consumers opt in — never a replacement of this file. That would also mean editing usecols and the two hardcoded row counts in four repos, and re-reviewing five figures. Keeping the frozen copy is the status quo and is what the manifest records today. There is no urgency: the current state is correct, documented, and the lecture embeds the live grapher as an iframe beside the frozen data anyway, so a reader already sees the current OWID vintage next to the figure.

Nothing else, yet

The other twenty-four manifested datasets have no known upstream delta. Twelve carry integrity.upstream.status: unverifiable, which is a different condition: not "we know it differs" but "we cannot check". Those are a reproducibility gap tracked through builder_status and PLAN Phase 9, not a drift item — they only enter this register if a check is later performed and finds a difference.

How entries get added

Any migration or verification that finds a delta adds a section here, and sets the dataset's integrity.upstream fields in the same PR so the catalog reflects it. The two must not diverge — a register entry with no manifest note is invisible to anyone reading the dataset.

Automating the detection side is proposed separately; see the upstream-freshness dashboard issue. This issue is the human register and stays useful regardless of whether that lands.

Part of #8. Related: #38 (the manifests that surfaced the first entry), #35 (licensing, the same record-and-track shape applied to a different question), and PLAN Phase 7.

Activity

  1. mmcky commented on Aug 6, 2026

    @mmcky
    ContributorAuthor

    The register now has a schema field, not just prose

    Following the question of whether #39 needed schema support: it did, and the gap was live rather than theoretical. mpd2020.xlsx was recording integrity.upstream.status: verified while differing from upstream — because the vocabulary (verified | spot-checked | unverifiable | unverified | failing) had no value for the outcome that actually occurred. The check ran in full and the file did not match. failing would imply something is broken, and verified is a claim a reader would act on: someone refreshing that file from the source would find it silently changes lecture output.

    Added in #38: a diverged status, plus the four fields this register needs.

    Field Purpose
    delta_kind local-edit or upstream-moved — the two kinds this issue distinguishes, now machine-readable rather than buried in prose
    delta what differs, precisely enough that resolving it later needs no re-investigation
    delta_evidence how delta_kind was established rather than assumed
    register points back here

    delta_evidence earns its place for a reason worth recording. Ruling out the other kind is the whole work, and it determines the file's class, not just its status: if Maddison had renamed its own headers after we took the file, our copy would be a faithful older vintage and verbatim; because upstream was byte-stable throughout, our copy was edited and it is constructed. Same three-cell difference, opposite classification. A delta_kind asserted without evidence would silently misclassify the dataset.

    diverged is deliberately not a defect state, which matters for how this register should be read. A migration moves the copy the lectures already consume, so a known delta against today's upstream is an expected result to record — not a problem to fix mid-repoint. The catalog gives it its own mark (⇄) so it cannot be misread as either a clean pass or a failure.

    The broader schema catch-up stays out of scope here and belongs with #14. The sketch is behind the manifests in ways that predate this work: consumers[].repo/.file is used by 10 manifests and documented in none, as are source.doi, source.note, source.version and source.file_url, plus the multi-sheet vocabulary. One sweep against a decided convention beats piecemeal additions.

  2. mmcky commented on Aug 7, 2026

    @mmcky
    ContributorAuthor

    Second entry: life-expectancy-vs-gdp-per-capita.csv — upstream moved

    Found while scoping Track A step 4 in QuantEcon/workspace-lectures#23. It is the first register entry that would break a build rather than quietly change a figure, so the ordering matters more here than it did for mpd2020.xlsx.

    The delta. Our committed copy (OWID, 2,059,709 B, in lecture-python-intro at lectures/_static/lecture_specific/simple_linear_regression/) has the header Entity,Code,Year,Life expectancy at birth (historical),GDP per capita,417485-annotations,Population (historical estimates),Continent. Today's export from the same OWID grapher differs in two columns.

    Column Our copy OWID today
    4 Life expectancy at birth (historical) Life expectancy at birth
    6 417485-annotations (dropped)

    Why it fails closed. lectures/simple_linear_regression.md:433 hardcodes the old label and passes it as usecols:

    cols = ['Code', 'Year', 'Life expectancy at birth (historical)', 'GDP per capita']
    df = pd.read_csv(data_url, usecols=cols)

    Against today's real header, pandas 2.3.3 raises ValueError: Usecols do not match columns. This is the mirror image of mpd2020.xlsx: there the renames are ours and load-bearing, and a pristine upstream copy would change the plot with no error; here the rename is upstream's and a refreshed copy stops the build on the spot. Both cases argue the same thing — the file and the lecture move together — but this one is the benign failure mode, and only if it is on the record before anyone reaches for a refresh.

    Classification. delta_kind: upstream-moved, integrity.upstream.status: diverged, per rule 4's second kind. Resolution therefore has a settled shape already: a newer vintage gets a new filename, so the frozen copy stays valid and simple_linear_regression.md:433 keeps working untouched. Nothing needs deciding now.

    delta_evidence is weaker than mpd2020's and should say so. What is measured is the two-column difference between our bytes and today's export; the direction rests on 417485-annotations being an OWID grapher export artefact keyed by variable id, which an edited copy would not acquire. That is inference, not the Internet Archive digest chain that settled mpd2020. The honest entry records it as such, and the Archive check on the grapher CSV endpoint is the work that would upgrade it.

    Register it before the step-4 repoint PR, not after. This file is not yet manifested, so its manifest is authored inside the repoint PR — which means integrity.upstream gets its first commit there. Rule 4 step 2 wants manifest and register landing together; if the entry trails the repoint, the manifest ships asserting nothing about the divergence, and the obvious next action on a fresh manifest — verify against upstream — is exactly what breaks. This is the first entry filed ahead of its own manifest, and that is the correct order for any delta discovered during scoping rather than during verification.

    "Nothing else, yet" now reads: one more, on a file still in the remaining 23.

  3. mmcky commented on Aug 11, 2026

    @mmcky
    ContributorAuthor

    Second entry added to the register: life-expectancy-vs-gdp-per-capita.csv, the first upstream-moved case (the first entry, mpd2020.xlsx, is local-edit). Paired with its manifest in #74, per this issue's own rule that a register entry with no manifest note is invisible to anyone reading the dataset.

    Three things worth flagging beyond the entry itself.

    One widely repeated description of this delta was wrong, and it pointed at the wrong failure. The work plan recorded that today's export "drops 417485-annotations". It does not — the column is renamed to GDP per capita (Annotations) and moved from position 6 to 8. The ValueError a refresh raises names only the life-expectancy column, because the annotations column is not in usecols at all. Had the entry been written from the plan's text, the public record would have been materially wrong about which change breaks the read.

    The delta is four renames, not one, plus halved coverage. 62,156 → 29,912 rows, Year -10000..2021 → 1..2023, 317 → 276 entities. Only Entity, Code, Year and GDP per capita survive unchanged.

    This entry is the clearest case yet for why the register exists. The loud failure — ValueError: Usecols do not match columns — is the harmless one, because it stops the build. The dangerous one is that year == 2018 returns 166 rows in both vintages, so someone who fixed the ValueError and rebuilt would get a scatter that looks right, an OLS line that looks right, and every underlying life-expectancy value silently different. That is rule 4 stated as a measurement rather than a principle.

    Also worth noting for whoever eventually reviews this register: mpd2020.xlsx and this file are opposite kinds and want opposite handling. mpd2020 is local-edit where the edits are load-bearing, so file and lecture must move together. This one is upstream-moved, so per AGENTS.md "Corrections vs vintages" adopting the new vintage means a new filename and consumers opting in — plus editing usecols and two hardcoded row counts in four repos, and re-reviewing five figures. Neither is urgent; both are now documented precisely enough that resolving them needs no re-investigation.

  4. mmcky commented on Aug 31, 2026

    @mmcky
    ContributorAuthor

    Register review, 2026-09-01 — the migration is complete, so this is the review step 3 promised

    The static migration finished on 2026-08-18 (40 of 40 datasets repointed, strict audit green on the 2026-08-31 run), which is the trigger this issue set for reviewing the register. Two entries, opposite kinds, and the recommendation is the same for both: keep the frozen copy, deliberately, and record that here. Neither delta is a defect; both are documented precisely enough in their manifests that the decision needs no re-investigation.

    mpd2020.xlsx (local-edit) — keep the edited copy as the intended state

    The data is the current Maddison release — the Internet Archive digest chain shows upstream has been byte-stable since 2021, so there is no vintage ambiguity and no newer data to gain. The only delta is three header labels that long_run_growth reads with header=(0,1,2), and those labels are load-bearing. Restoring the pristine upstream file would mean changing the column handling in all four consuming repos (lecture-python-intro, lecture-wasm, lecture-intro.zh-cn, and the canary test-actions-lecture-intro — corrected from "three" after Copilot's review on #108) for a lecture whose output would not change by a single value. The class is already constructed and the manifest already says the state is intended. Decision: keep. No manifest change needed, beyond optionally one line in integrity.upstream.note saying "reviewed 2026-09-01, kept deliberately" so a reader of the catalog sees the delta was adjudicated and not merely recorded.

    life-expectancy-vs-gdp-per-capita.csv (upstream-moved) — keep the frozen copy; adopting the new vintage is a lecture decision, not an infrastructure one

    Adopting OWID's current export is a content change with the profile rule 4 warns about: four renamed columns, coverage cut from 62,156 to 29,912 rows, and a year == 2018 scatter that would look identical while every life-expectancy value shifted. Doing it properly means a new filename (per AGENTS.md, so consumers opt in), edits to usecols and two hardcoded row counts in four repos, and five figures re-reviewed. There is no reader-facing benefit to set against that: the lecture already embeds the live OWID grapher as an iframe beside the frozen data, so a reader sees the current vintage next to the figure today. Decision: keep the frozen copy. If the lecture authors ever want the 2023 coverage for pedagogical reasons, that request should arrive as a lecture-python-intro issue and be executed as a new-filename vintage landing here — the register entry and manifest already describe exactly what that would take.

    What this means for the register

    Rule 4's three steps are now all exercised at least once: migrate unchanged, record the delta, review after the migration. The issue stays open as the standing register — entries are added by any future verification that finds a delta — and the upstream-freshness dashboard (#40) remains the proposal for finding those deltas routinely rather than by accident. Both entries above are marked as reviewed by this comment; the next review is due when a new entry lands or the #40 dashboard reports drift, not on a calendar.

  5. mmcky commented on Sep 1, 2026

    @mmcky
    ContributorAuthor

    Third entry: business_cycle_data.csv — upstream moved, and for the first time that is the file's normal state

    Added with its manifest in #109 (per this issue's rule that the register entry and the manifest's integrity.upstream land together). This one is different in kind from the first two, and the difference is worth stating because it is what every future dynamic snapshot will look like.

    The delta. The builder was re-run in full against live WDI (lastupdated 2026-07-13) on 2026-09-01, wbgapi 1.0.12 / pandas 2.3.3. Same five rows, same layout, same single structural null (YR1960). Two new year columns, YR2024 and YR2025, populated for all five economies. In the 1960–2023 overlap, 63 of the 64 year columns carry at least one revised cell — 236 of 320 cells in all — with median |change| 0.0 pp, 90th percentile 0.29, 99th percentile 1.00, maximum 1.50 pp.

    upstream-moved, and the evidence is structural rather than archival. The committed file is the verbatim output of wb.data.DataFrame(...).to_csv() with no post-processing, and the builder that wrote it is the one committed beside it, so a local edit has nowhere to have happened. The shape of the delta — small revisions spread across the whole history, largest in recent years, plus appended years — is the signature of national-accounts revision and rebasing, not of anything on this side.

    Why this entry does not want a decision. The first two entries were frozen extracts where the delta was a surprise and "keep or adopt" was a real question. This file is class: dynamic-snapshot with cadence: annual: being behind the source is its steady state between refreshes, and the resolution is the refresh itself — a deliberate, reviewed in-place update via the Phase 5 refresh-as-PR workflow, not a new filename (it is not a new vintage in the AGENTS.md sense; it is the same tracking snapshot moved forward) and not a correction. Nothing consumes the file today, so there is no figure at stake in the meantime.

    What the measurement did settle. The builder's validate() cannot assert overlap-window equality for a revised aggregate — the very first refresh would fail. It now bounds any single-cell revision at 5 pp (three times the largest routine revision observed), refuses a populated cell going empty or a column disappearing, and prints the revision summary on every run as the review surface. That is the overlap-window policy the Phase 5 template inherits for dynamic snapshots, and this entry is the measurement it rests on.

    Next review of this entry: the first refresh PR, whose diff summary either confirms the 5 pp bound or teaches us it was wrong.

  6. mmcky commented on Sep 1, 2026

    @mmcky
    ContributorAuthor

    Closing the loop on the third entry: the first refresh PR it pointed forward to has happened — #112, opened by the new refresh-snapshots workflow on its first dispatch and merged 2026-09-01. business_cycle_data.csv is now the 2025 vintage; its manifest reads integrity.upstream.status: verified with the delta block dropped by the stamp, and snapshots.py due reports it not due until 2027-09.

    The 5 pp bound held with room to spare: the largest revision in the 1960–2023 overlap was 1.50 pp, exactly the measurement the bound was set from, and no populated cell went empty. So the overlap-window policy this entry recorded stands as written; the next data point is the 2027 refresh.

    One thing the PR review turned up that belongs in the register's rules rather than in this entry: a dynamic snapshot's manifest must not carry prose that embeds a vintage fact (an end year, an observed range, a row count), because the stamp rewrites fields, not sentences — the first refresh shipped a column description still saying "YR2023 in the committed bytes". Recorded in AGENTS.md and builders/_template.py via #112.

    This entry is therefore resolved in the way a dynamic snapshot's entries always will be — by the refresh — and needs no further review. The register stays open for the next delta.

  7. mmcky commented on Sep 7, 2026

    @mmcky
    ContributorAuthor

    Closing as completed, ruled by @mmcky on 2026-09-07. The register did its job: the static migration completed on 2026-08-18 (40 of 40 repointed, strict audit green 2026-08-31), the review step ran on 2026-09-01 with both frozen copies kept deliberately, and the third entry was refreshed by #112 the same day. The manifests' integrity.upstream field and the generated catalog are the register of record from here — #125 records that in AGENTS.md and cites this issue's history as the worked example of the three steps. New deltas are set in the manifest in the PR that finds them; #40 remains the proposal for detecting them routinely.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions