Our first public correction-ledger entry is not a correction. It is a hold. A research candidate passed five GLM-5.3 role reviews and nine mechanical gates, then failed a final visual check because two chart headlines still displayed superseded numbers.
The article never shipped. No reader acted on the stale labels. Calling that event a published correction would exaggerate the harm. Deleting it from the record would hide a control failure. The honest category sits between those choices: resolved by withholding.
A correction measures what happened after publication. A hold records why publication did not happen. Both are evidence about whether the system works.
The clean-launch myth
Most research sites show readers a sequence of finished pages. That presentation creates an appealing fiction: every published result arrived whole, and everything not published was merely unfinished. Real analytical work is messier. Candidates fail because a source moved, a definition drifted, a chart retained old text, a statistical check contradicted the narrative, or an editor asked a question the analysis could not answer.
Those failures matter even when readers never see the candidate. They reveal which controls caught the problem, which controls missed it, and whether the organization changed the system or simply tried again. A public hold is therefore not a confession that a false result was published. It is a receipt showing that publication was actually gated.
One ledger, four different obligations
| State | What happened | Required response |
|---|---|---|
| Prepublication hold | A candidate fails before release. | Preserve the run, record the failed surface, and do not publish it. |
| Minor correction | A live detail is wrong but the conclusion survives. | Correct the page, date the note, and preserve the change in the ledger. |
| Material correction | A live error changes advice, estimates, or a central claim. | Flag it prominently, explain reader impact, publish replacement evidence, and notify affected readers where practical. |
| Retraction | The central conclusion is no longer supportable. | Keep the URL and record, mark the result withdrawn, and explain why it should not be relied upon. |
Putting these states in one ledger does not make them equivalent. It gives them a common identity, date, status, explanation, and evidence path. The entry type and severity preserve the difference in reader impact.
What COR-001 actually says
The held Touchdown Regression candidate contained internally inconsistent presentation. Its updated analysis reported an 11.7% walk-forward improvement and a 3.0-touchdown decline for the highest overperformance group. Two rendered PNG headlines still said 8.9% and 2.8 touchdowns. The underlying tables were updated; the visible words were not.
The agent roles approved the research bundle, and the mechanical gate passed because neither control was then checking the text baked into the images. Human visual QA caught the mismatch before release. The candidate was withheld, its publication record was removed, and a process note was attached to the exact run. The public Touchdown Regression page remained on its earlier deterministic launch snapshot.
That sequence matters. “We noticed a chart bug” is a story. The preserved run ID, process note, absent publication record, figure hashes, and ledger entry make it inspectable.
The fix had to change the system
Correcting two strings would have repaired one candidate. It would not have reduced the chance of recurrence. We added a figure contract that binds each chart to its source metrics, visible headline, and final PNG hash. The publication gate validates that contract, and the promotion step checks it again before moving any candidate into the live site.
The second check is deliberate. A gate can pass, then a file can change before publication. Promotion therefore re-hashes the manifest, validates article claims, validates figure semantics, and refuses deterministic offline role identities when an online-reviewed release is required. Publication becomes one atomic transition rather than a collection of manual copy operations.
Why the original evidence stays put
Silent replacement makes a repository look cleaner while making its history less useful. The stale charts, approved role artifacts, gate decision, and hold note remain attached to the same run. A corrected candidate must receive a new run identity. It cannot overwrite the failed attempt and inherit its approvals.
This is also why the correction ledger links to evidence instead of merely summarizing it. The prose is an interpretation; the artifacts show the state that produced the interpretation. If they disagree, the discrepancy is visible.
The ledger is verified, not self-proving
A build-time verifier checks that correction IDs are unique, every referenced run resolves, a held run says it was not published, and no publication record exists for that held run. The same verifier blocks a release if the reporting-channel gap is accidentally softened into a claim that intake is ready.
Those checks prove internal consistency. They do not prove that every mistake has been found, that severity was judged perfectly, or that the organization will respond well under pressure. A policy is a constraint on behavior, not evidence of flawless behavior.
What is still missing
We do not yet have a durable public address for readers to report an error. The ledger says so plainly. A real intake channel needs reliable delivery, ownership, response expectations, and protection from the channel disappearing when one person changes an account setting.
That gap matters commercially too. We will not treat a public corrections page as sufficient support infrastructure for a paid subscription. Billing creates obligations: durable contact, entitlement recovery, refunds, and a way to reach affected members when a material result changes. The honest status remains not ready.