The desk studies
itself.

We turn every preserved GLM-5.3 failure into a regression case, connect it to a tested control, and publish the standard any future model must clear before we claim improvement.

ONLINE RUNS03

One series · one day

GLM ROLE OUTPUTS13/15

Preserved / configured

KNOWN FAILURE CASES3/3

Mapped to tested controls

LIVE CHALLENGERS00

No model comparison claimed

A memory system for mistakes.

This is a three-incident control baseline, not a leaderboard, accuracy estimate, or proof of general model quality.

  • Only three preserved online runs exist, all from one research series and one day.
  • All online roles used the same model family, so role separation did not create independent judgments.
  • The incidents were observed in production-style research runs, not a blinded challenger benchmark.
  • No challenger model has been called, scored, approved, or rejected under this protocol yet.
Open the machine-readable evaluation register ↗

Count the missing
outputs too.

Every role used z-ai/glm-5.3 through OpenRouter. These counts describe artifacts, not five independent minds.

ROLEOUTPUTSPASSREVISEBLOCKNOT INVOKED
Researcherz-ai/glm-5.32/32001
Statisticianz-ai/glm-5.33/31110
Adversarial Reviewerz-ai/glm-5.33/32100
Citation Checkerz-ai/glm-5.33/32010
Editorz-ai/glm-5.32/31101

Failure becomes
infrastructure.

Each case preserves what failed, which layer stopped release, the remediation, and the tests that now defend the boundary.

FAIL-001RESEARCH GATE AND POST RUN AUDIT

RUN / 2026-08-31-touchdown-regression-v1

The workflow let the model rewrite the contract

WHAT FAILED / 5
  1. The online researcher was incorrectly allowed to replace the locked preregistration.
  2. Reviewers received the memo and analysis summary but not the full provenance, article, or artifact inventory.
  3. The test-season range was encoded ambiguously as [2018, 2024].
  4. The pipeline did not generate a row-level attrition reconciliation.
  5. The editor role was configured but was not invoked.
WHAT CHANGED / 6
  1. Load the preregistration from the versioned series registry and make it immutable during a run.
  2. Use the researcher as a conformance reviewer rather than an author of the locked memo.
  3. Supply provenance, article text, analysis, mechanical checks, and a SHA-256 artifact inventory to every specialist.
  4. Make role, model, and timestamp transport-authoritative.
  5. Publish an attrition-by-season artifact and a full list of walk-forward test seasons.
  6. Invoke the editor after specialist review and block on any non-pass decision.
FAIL-002SPECIALIST NONPASS AND RESEARCH GATE

RUN / 2026-08-31T010623Z-touchdown-regression-v1

Valid JSON still contained invalid judgment

WHAT FAILED / 4
  1. The statistician issued a revise decision for concrete article framing changes around bootstrap interpretation, selection effects, and individual-level application.
  2. The adversarial reviewer returned a pass decision containing a literal placeholder blocking finding.
  3. The editor returned literal placeholder summary and finding text.
  4. The JSON schema enforced structure but did not yet enforce minimum semantic content or decision/finding coherence.
WHAT CHANGED / 6
  1. Require minimum lengths for review summaries and finding fields.
  2. Reject a pass decision containing a blocking finding and a block decision without one.
  3. Retry invalid structured outputs rather than writing placeholder artifacts.
  4. Remove the bootstrap win-rate framing and explain that resamples are not independent replications.
  5. Bound the tiebreaker advice to a population-level heuristic and disclose the uncertain direction of survivor conditioning.
  6. Trace the email subject to a generated statistic and cite both nflverse source files.
FAIL-003HUMAN VISUAL QA

RUN / 2026-08-31T011519Z-touchdown-regression-v1

Every model passed; the pixels were still wrong

WHAT FAILED / 2
  1. figures/forecast-error.png: visible 8.9%; analysis 11.7%
  2. figures/regression-quintiles.png: visible 2.8 touchdowns; analysis 3.0 touchdowns
WHAT CHANGED / 1
  1. Every future figure is rendered from a JSON figure contract containing its visible title, source metrics, and final PNG hash. The mechanical gate and promotion command independently validate that contract against analysis.json.

The repair must
have a test.

A promise is not a control. Every item below resolves to implementation files, named tests, and a public contract where one exists.

CTRL-001IMPLEMENTED + TESTED

Immutable registration

A model replacing the preregistered estimand or design after results exist.

CODE SURFACES
2
NAMED TESTS
1
PUBLIC CONTRACTS
1
CTRL-002IMPLEMENTED + TESTED

Complete reviewer evidence bundle

Specialists reviewing a summary without provenance, article text, inventory, or mechanical checks.

CODE SURFACES
1
NAMED TESTS
1
PUBLIC CONTRACTS
2
CTRL-003IMPLEMENTED + TESTED

Transport-authoritative identity

A model self-reporting a different role, model name, or timestamp inside its answer.

CODE SURFACES
1
NAMED TESTS
1
PUBLIC CONTRACTS
1
CTRL-004IMPLEMENTED + TESTED

Five-role completion and all-pass gate

A configured specialist never running, or a revise/block decision being ignored.

CODE SURFACES
2
NAMED TESTS
1
PUBLIC CONTRACTS
1
CTRL-005IMPLEMENTED + TESTED

Row-level attrition reconciliation

Forecast metrics silently conditioning on an unexplained surviving subset.

CODE SURFACES
2
NAMED TESTS
1
PUBLIC CONTRACTS
0
CTRL-006IMPLEMENTED + TESTED

Semantic review validation

Placeholder findings and incoherent pass/block combinations satisfying a merely structural schema.

CODE SURFACES
1
NAMED TESTS
2
PUBLIC CONTRACTS
1
CTRL-007IMPLEMENTED + TESTED

Per-attempt call receipts and bounded retry

A malformed attempt disappearing after retry, along with its cost, identity, or failure status.

CODE SURFACES
1
NAMED TESTS
1
PUBLIC CONTRACTS
1
CTRL-008IMPLEMENTED + TESTED

Generated claim framing

Bootstrap resamples being described as independent replication or a population heuristic becoming a player-level rule.

CODE SURFACES
1
NAMED TESTS
1
PUBLIC CONTRACTS
1
CTRL-009IMPLEMENTED + TESTED

Figure claim contract and atomic promotion

A current chart containing a stale visible title or being copied separately from its reviewed run.

CODE SURFACES
2
NAMED TESTS
2
PUBLIC CONTRACTS
0

GLM-5.3 stays until evidence beats it.

No model change is justified. GLM-5.3 remains the named default and no challenger has been evaluated.

01 / CORPUS
  • Use already-public artifacts or obtain explicit disclosure approval before any prepublication bundle leaves the repository.
  • Freeze the same evidence, role prompt, output schema, temperature, output cap, retry budget, and scoring rules for incumbent and challenger.
  • Keep the failure-case labels outside the model-visible evidence and score every role output from preserved artifacts.
02 / MEASURES
  • Schema-valid output rate and retry rate
  • Known blocking-failure recall
  • False-pass count on known failure cases
  • Evidence-specificity and placeholder rejection
  • Role-decision coherence
  • Exact provider-reported tokens, billed cost, and transport latency
03 / APPROVAL
  • Zero publication-eligible false passes across the known failure corpus after deterministic gating.
  • No regression on any incumbent control case without a documented, public exception.
  • Complete per-attempt receipts for both models; missing usage remains null and cannot be estimated.
  • A human release decision that names quality, cost, latency, shared-model risk, and the limits of the sample.