One series · one day
AGENT EVALUATIONS / INCIDENT BASELINE v1.0.0
The desk studies
itself.
We turn every preserved GLM-5.3 failure into a regression case, connect it to a tested control, and publish the standard any future model must clear before we claim improvement.
Preserved / configured
Mapped to tested controls
No model comparison claimed
WHAT THIS SCORECARD IS
A memory system for mistakes.
This is a three-incident control baseline, not a leaderboard, accuracy estimate, or proof of general model quality.
- Only three preserved online runs exist, all from one research series and one day.
- All online roles used the same model family, so role separation did not create independent judgments.
- The incidents were observed in production-style research runs, not a blinded challenger benchmark.
- No challenger model has been called, scored, approved, or rejected under this protocol yet.
ROLE HISTORY / GLM-5.3
Count the missing
outputs too.
Every role used z-ai/glm-5.3 through OpenRouter. These counts describe artifacts, not five independent minds.
FAILURE CORPUS / 03
Failure becomes
infrastructure.
Each case preserves what failed, which layer stopped release, the remediation, and the tests that now defend the boundary.
RUN / 2026-08-31-touchdown-regression-v1
The workflow let the model rewrite the contract
- The online researcher was incorrectly allowed to replace the locked preregistration.
- Reviewers received the memo and analysis summary but not the full provenance, article, or artifact inventory.
- The test-season range was encoded ambiguously as [2018, 2024].
- The pipeline did not generate a row-level attrition reconciliation.
- The editor role was configured but was not invoked.
- Load the preregistration from the versioned series registry and make it immutable during a run.
- Use the researcher as a conformance reviewer rather than an author of the locked memo.
- Supply provenance, article text, analysis, mechanical checks, and a SHA-256 artifact inventory to every specialist.
- Make role, model, and timestamp transport-authoritative.
- Publish an attrition-by-season artifact and a full list of walk-forward test seasons.
- Invoke the editor after specialist review and block on any non-pass decision.
RUN / 2026-08-31T010623Z-touchdown-regression-v1
Valid JSON still contained invalid judgment
- The statistician issued a revise decision for concrete article framing changes around bootstrap interpretation, selection effects, and individual-level application.
- The adversarial reviewer returned a pass decision containing a literal placeholder blocking finding.
- The editor returned literal placeholder summary and finding text.
- The JSON schema enforced structure but did not yet enforce minimum semantic content or decision/finding coherence.
- Require minimum lengths for review summaries and finding fields.
- Reject a pass decision containing a blocking finding and a block decision without one.
- Retry invalid structured outputs rather than writing placeholder artifacts.
- Remove the bootstrap win-rate framing and explain that resamples are not independent replications.
- Bound the tiebreaker advice to a population-level heuristic and disclose the uncertain direction of survivor conditioning.
- Trace the email subject to a generated statistic and cite both nflverse source files.
RUN / 2026-08-31T011519Z-touchdown-regression-v1
Every model passed; the pixels were still wrong
- figures/forecast-error.png: visible 8.9%; analysis 11.7%
- figures/regression-quintiles.png: visible 2.8 touchdowns; analysis 3.0 touchdowns
- Every future figure is rendered from a JSON figure contract containing its visible title, source metrics, and final PNG hash. The mechanical gate and promotion command independently validate that contract against analysis.json.
CONTROL REGISTER / 09
The repair must
have a test.
A promise is not a control. Every item below resolves to implementation files, named tests, and a public contract where one exists.
Immutable registration
A model replacing the preregistered estimand or design after results exist.
- CODE SURFACES
- 2
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 1
Complete reviewer evidence bundle
Specialists reviewing a summary without provenance, article text, inventory, or mechanical checks.
- CODE SURFACES
- 1
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 2
Transport-authoritative identity
A model self-reporting a different role, model name, or timestamp inside its answer.
- CODE SURFACES
- 1
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 1
Five-role completion and all-pass gate
A configured specialist never running, or a revise/block decision being ignored.
- CODE SURFACES
- 2
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 1
Row-level attrition reconciliation
Forecast metrics silently conditioning on an unexplained surviving subset.
- CODE SURFACES
- 2
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 0
Semantic review validation
Placeholder findings and incoherent pass/block combinations satisfying a merely structural schema.
- CODE SURFACES
- 1
- NAMED TESTS
- 2
- PUBLIC CONTRACTS
- 1
Per-attempt call receipts and bounded retry
A malformed attempt disappearing after retry, along with its cost, identity, or failure status.
- CODE SURFACES
- 1
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 1
Generated claim framing
Bootstrap resamples being described as independent replication or a population heuristic becoming a player-level rule.
- CODE SURFACES
- 1
- NAMED TESTS
- 1
- PUBLIC CONTRACTS
- 1
Figure claim contract and atomic promotion
A current chart containing a stale visible title or being copied separately from its reviewed run.
- CODE SURFACES
- 2
- NAMED TESTS
- 2
- PUBLIC CONTRACTS
- 0
MODEL-CHANGE PROTOCOL
GLM-5.3 stays until evidence beats it.
No model change is justified. GLM-5.3 remains the named default and no challenger has been evaluated.
- Use already-public artifacts or obtain explicit disclosure approval before any prepublication bundle leaves the repository.
- Freeze the same evidence, role prompt, output schema, temperature, output cap, retry budget, and scoring rules for incumbent and challenger.
- Keep the failure-case labels outside the model-visible evidence and score every role output from preserved artifacts.
- Schema-valid output rate and retry rate
- Known blocking-failure recall
- False-pass count on known failure cases
- Evidence-specificity and placeholder rejection
- Role-decision coherence
- Exact provider-reported tokens, billed cost, and transport latency
- Zero publication-eligible false passes across the known failure corpus after deterministic gating.
- No regression on any incumbent control case without a documented, public exception.
- Complete per-attempt receipts for both models; missing usage remains null and cannot be estimated.
- A human release decision that names quality, cost, latency, shared-model risk, and the limits of the sample.