MOVE 37 → MOVE 42

MOVE42 / INNOVATIONZERO

The next move is an algorithm.

Move 37 answered a Go position with one coordinate no human expected. Move42 asks the same question in a larger action space: when the board is a requirement, can the move be a complete executable algorithm?

POSITION → COORDINATEbecomesREQUIREMENT → ALGORITHM
Editorial thesis briefVEOX · 2026

01 · Same abstraction, larger action space

02

A requirement is a board. The contract defines legal play. The referee decides what worked.

MOVE 37

Go positionQ16one legal coordinate

MOVE42

Technical requirementExecutable algorithmone independently runnable method
state → legal action → external consequence → preserved experience

The proposed output changes from an answer or score to an independently runnable method. The learning target is the distribution of later executable proposals across requirements.

02 · Where search occurs

03

Search now. Learn for later. Or combine both.

A

Search on the current requirement

AutoResearch improves through a proposal–execution–inspection–retry loop on the current requirement.

propose → execute → inspect → retry ↻
B

Learn across prior requirements

Move42 aims to improve the proposal distribution across requirements, so the first executable move itself becomes learnable.

receipts → later proposal distribution → first executable move

Hybrids remain possible; the distinction is where search occurs and what persists between tasks.

The current record does not measure superiority over research agents, task-time speedup, or an admitted prospective strict first move.

03 · How to build the game

04

Board. Move. Rules.
Referee. Memory.

01

Board

The requirement plus every permitted input, state transition, and resource boundary.

Encode the requirement and permitted state before any proposal is made.
02

Move

One complete executable algorithm whose outputs can be independently reproduced.

Define the runnable artifact, interface, and terminal outputs that count as a move.
03

Rules

The legality contract that rejects shortcuts, leakage, hidden state, and invalid resources.

Close shortcut and leakage paths before play, then make every violation an explicit failure.
04

Referee

An external consequence-based evaluator that executes the move and records success or a named failure.

Build the referee outside model self-assessment and bind its identity to every outcome.
05

Memory

Canonical success and failure receipts that can supervise later proposals across requirements.

Preserve receipts, train later proposals, and separately freeze a true one-proposal prospective evaluation.

04 · Current evidence

05

A finite, inspectable 600-run record.

600valid runs
353wins
247losses
58.8%vs TabFM Ensemble
Four-model Elo ranking over 600 valid runsStatic final four-model Elo ranking computed from the 600 valid runs.90095010001050jope-prime1038.1TabPFN v31028.9TabFM Ensemble996.5TabFM936.6
Final computed Elo ranking
RankModelFinal Elo
1jope-prime1038.1
2TabPFN v31028.9
3TabFM Ensemble996.5
4TabFM936.6
TabFM, TabFM Ensemble, TabPFN v3, and jope-prime from the jAIn foundation family.Simulated learning path from 900 Elo; only the final four-model Elo ranking is computed from the 600 valid runs.Source: move42.elo-600.v1.

The 58.8% result describes these 600 runs and is not a prospective universal success rate.

05 · Three requirements, three moves

06

Algorithms respond to the board they receive.

Named within-dataset win

Cpu

Nested polynomial interactions and sigmoid transforms feed a ridge-GCV head; the receipt describes the executable structure without assigning a causal explanation.

vs TabFM
6.66×
vs TabFM Ensemble
5.13×
vs TabPFN v3
11.34×

Named within-dataset win

Greenhouse

A Gaussian-process regressor combines radial-basis and periodic components for the two-feature table; the match is case-specific, not a climate claim.

vs TabFM
3.40×
vs TabFM Ensemble
4.66×
vs TabPFN v3
5.03×

Named within-dataset win

Spectrometer

A compact ridge-GCV calibration algorithm is the recorded quality outlier; its longer complete workflow keeps it outside the fast focal pair.

vs TabFM
5,434.54×
vs TabFM Ensemble
4,373.89×
vs TabPFN v3
4,530.65×

Each value is a labeled within-dataset RMSE comparison. Exact algorithms and source-bound records remain available on the atlas.

06 · Receipts become memory

07

Every outcome can teach the next proposal.

01Proposecomplete executable algorithm
02Executeexternal consequence-based referee
03Preservesuccess, loss, or named failure
04Train laterstored supervision across requirements

The update path

Execution produces the learning signal.

Gradients update the proposing model after execution; each algorithm remains independently runnable, and frozen validation decides what persists.

07 · What does this mean?

08

Four possible games.
All explicitly vision.

Vision · Compiler game

A source algorithm, target semantics, architecture, and optimization contract.

Move. A complete transformation or optimization algorithm that produces executable output.

Referee. Correctness suites, resource limits, and consequence-based performance measurements.

A design horizon, not a finding of the current regression study.

Vision · Control game

A control requirement, observable plant state, hard constraints, and permitted actuators.

Move. A complete control algorithm that can be executed against the declared interface.

Referee. Constraint violations, stability, resource use, and measured physical consequence.

A design horizon, not evidence of autonomous control performance.

Vision · Experimental game

An apparatus, measurement protocol, budget, safety envelope, and experimental objective.

Move. A complete executable experiment schedule and analysis method.

Referee. Predeclared measurements, explicit failure outcomes, and independently preserved receipts.

A design horizon, not an experimental result reported here.

Vision · Scientific-design game

A scientific question, admissible evidence, interventions, and falsification criteria.

Move. A complete executable design for collecting and analyzing the next evidence.

Referee. Preregistered consequence tests that can reject as well as support the proposed method.

A design horizon, not a claim of causal or scientific discovery.

08 · Two scales, kept separate

09

317,156,096 generated algorithms.

A hash-bound generation-scale fact. It is not a performance-run total.

Generation scale

317,156,096generated algorithms

Source: algorithm-generation receipt.

Performance evidence

600valid performance runs

Source: 600-run atlas and Elo receipt.

The 58.8% result and final four-model Elo ranking use only the 600 valid runs.

09 · Limits and source record

10

Bold thesis.
Visible limits.

Established here

An inspectable 600-run result, a computed final four-model Elo ranking, named focal comparisons, executable algorithms, and a separate generation-scale receipt.

Not established

Expected future win rate, focal-case causality, matched comparator speed, universal superiority, safe autonomous use, or prospective one-proposal performance.

Active paper

Move42: InnovationZero Learns to Invent Executable AlgorithmsSHA-256 7a80bbfac4f20a9db7b94bd46111fa6262d72bc65dbcf33c3605a484f68dabce

The immutable v1 research-paper bytes remain unchanged. Former public evidence bytes are preserved privately with their existing hashes.