Corrected public evidence

One 600-run performance record.

All public performance claims use the same 600 valid runs: 353 wins, 247 losses, and a 58.8% result against TabFM Ensemble.

Current corrected record

Computed Elo

The four-model ranking starts every rating at 1000 and applies 60 normalized batch-update rounds with K=64. Lower RMSE wins each pairwise comparison; numerical ties follow the 1e-12 convention. Inspect every comparison and source hash.

Published data

Corrected evidence v1 · Corrected evidence v2 · Corrected discovery · Elo 600 v1 · 600-run atlas · Focal cases · Filtered provenance v1 · Filtered provenance v2.

Withdrawn surfaces

The former estimates and campaign endpoints still return explicit withdrawal records with a link to the corrected Elo endpoint. Their former public bytes are retained only in the private archive and are not part of public provenance.

Limits

The 58.8% result describes these 600 runs and is not a prospective universal success rate. Focal cases remain post-hoc and descriptive. The 317,156,096 generated algorithms are a separate generation-scale fact, not additional performance runs.

Evidence claim

Claim detail

Status
Source
Pointer
Limitation

Open the normalized ledger