Your own mistakes are the first eval set

Model benchmarks are useful. The failures your actual workflow already produced are harder to dismiss—and more valuable to remember.

If an autonomous research desk fails in a way worth writing a postmortem about, it has already produced the beginning of an evaluation set. The mistake is specific, the evidence is real, and the cost of forgetting it is obvious.

Fourth Down Labs now has three preserved online incidents. The first candidate let the model replace a locked registration, gave reviewers an incomplete bundle, accepted self-authored model metadata, omitted attrition, and never invoked two configured roles. The second produced valid JSON with placeholder judgment and concrete framing problems. The third passed all five GLM-5.3 roles and all nine research checks, then failed visual QA because two PNG titles were stale.

THE INCIDENT RULE

Do not ask only whether the model was good. Ask which layer should have made that exact failure impossible to publish.

The workflow is the unit under test

A role-model benchmark can tell us whether GLM-5.3 recognizes a statistical flaw in a supplied passage. It cannot tell us whether the orchestrator actually invoked the editor, whether transport code trusted the model's self-reported identity, whether the gate reconciled every row, or whether the final website displayed the reviewed number.

Those are system behaviors. Our evaluation target is therefore the whole publication path: registration, evidence construction, specialist outputs, schema semantics, deterministic gating, promotion, visual QA, and the public record. GLM-5.3 is named and measurable inside that system, but it is not credited for controls the code or a human supplied.

Five moves from incident to eval

01

Preserve

Keep the rejected or withheld run intact. Do not “fix” history by overwriting the evidence that exposed the weakness.

02

Classify

Name the failed layer: model judgment, schema, orchestration, deterministic gate, visual QA, or release process.

03

Control

Change the smallest enforceable boundary that would prevent the same failure from becoming reader-visible again.

04

Test

Add a mutation that makes the old failure recur and require the build or promotion path to reject it.

05

Publish

Show the incident, repair, test, and remaining limitation together—before claiming the system improved.

Denominators make the story honest

Across three online runs, five GLM-5.3 roles created 15 configured role opportunities. Thirteen outputs exist. The first researcher and editor were never invoked. Saying “three AI-reviewed runs” would erase the most important orchestration failure, so the public scorecard reports 13 of 15.

Role decisions also stay separate. The statistician produced one block, one revise, and one pass. The citation checker produced one block and two passes. The adversarial reviewer produced one revise and two passes. The editor produced one revise, one pass, and one missing output. These are preserved outcomes, not an accuracy rate and not independent votes.

Nine repairs, not nine victory laps

The incident register maps the failures to nine controls: immutable registration; complete evidence bundles; transport-authoritative identity; five-role completion; row-level attrition reconciliation; semantic review validation; per-attempt receipts; generated claim framing; and figure claim contracts with atomic promotion.

Each control names implementation surfaces and at least one test. The public build independently regenerates the evaluation register and fails if the committed scorecard drifts. Seven adversarial mutations try to hide a known incident, credit GLM-5.3 for the visual-QA catch, remove a control's test status, inflate role totals, invent a challenger evaluation, approve a nonexistent model swap, or name an unsupported challenger.

Why this is not a model leaderboard

Three runs from one research series on one day cannot estimate general model quality. The incidents were not blinded, no challenger saw the same frozen corpus, and every online role used one model family. Calling the result “GLM-5.3 accuracy” would be false precision.

What the baseline can support is narrower and useful: these failures happened; these artifacts preserve them; these controls now exist; these tests exercise the repaired boundaries; and any future model comparison must carry the same evidence and scoring rules.

How a future model earns the desk

GLM-5.3 remains the default until evidence justifies a change. An incumbent and challenger must receive the same frozen public—or explicitly disclosure-authorized—corpus, prompts, schemas, temperature, output cap, retry budget, and hidden failure labels. We will compare schema validity, blocker recall, false passes, evidence specificity, decision coherence, exact usage, billed cost, and latency.

No challenger may create a publication-eligible false pass on the known corpus. A change also needs complete per-attempt receipts and a human release decision that names quality, price, speed, shared-model risk, and sample limits. “Newer” is not a result.

Transparency should compound

A secret workflow can improve internally. A public workflow can create a second asset: readers learn what changed, future contributors inherit the reasons behind the controls, and model upgrades become falsifiable decisions instead of brand announcements.

The scorecard currently contains zero challenger evaluations and zero approved swaps. Those zeros are not embarrassing. They are the difference between having a protocol and pretending it has already produced evidence.

AI-009

The method is part of the product.

All process notes →