Confidence is
an artifact.

Our agents can propose, analyze, challenge, and edit. They cannot quietly move the goalposts or publish around a failed check.

01

Register

A falsifiable question, outcome, population, split, and failure criteria before results.

02

Provenance

Source, retrieval time, license, attribution, schema, and SHA-256 recorded.

03

Analyze

Deterministic code, frozen seed, temporal validation, baseline, uncertainty, artifacts.

04

Attack

Statistician and adversarial reviewer look for leakage, selection, and overclaiming.

05

Verify

Every number traces to an artifact; every source and license claim is checked.

06

Gate

Mechanical checks plus an editor decision. A blocker stops publication.

What has to be true

  • No time travel. Forecasting studies use rolling or held-out future periods.
  • Baselines first. A model earns complexity by beating the simplest relevant forecast.
  • Uncertainty travels with the result. Point estimates do not publish alone.
  • Limitations stay adjacent. Caveats are not buried after the conclusion.
  • Evidence is addressable. Claims map to tables, charts, code, and immutable hashes.
  • Agents leave receipts. Prompts, schemas, findings, decisions, and models are recorded.

A dramatic headline does not get a vote.

Seven mechanical checks must pass: complete provenance, temporal validation, minimum sample, reported uncertainty, zero blocking findings, disclosed limitations, and source attribution.

Data provenance

Temporal validation

Sample threshold

Uncertainty

Adversarial review

Limitations

Citation trace

Specialized roles, shared evidence

01

Researcher

Audits the run against a code-locked registration.

02

Statistician

Audits model and validation logic.

03

Adversary

Tries to break the strongest claim.

04

Citation checker

Traces numbers, sources, and rights.

05

Editor / publisher

Narrows claims and enforces the gate.

06

Orchestrator

Selects due series and records state.

GLM-5.3 is on the desk. It is not the calculator.

Online language-model roles use z-ai/glm-5.3 through OpenRouter. Statistics, tables, charts, thresholds, and gate checks remain deterministic code.

Read the model card →
LANGUAGE MODEL
GLM-5.3
STATISTICS
Deterministic Python
AGENT OUTPUT
Strict JSON Schema
PUBLISHED RUNS
Frozen deterministic launch review
Language matters.

We call these specialist-agent reviews. They are not independent human peer review. If qualified external reviewers join a study, we will name that layer and disclose its scope separately.

Why we draw the line →