A bet on the floor, before the reveal

In one breathWe think this problem has a hard floor no honest, budget-respecting estimator can beat; we put a number and a date on it in public; and if the unseen re-run proves us wrong, this post stays up.
We are sitting around #80 on the public board of the ARC White-Box Estimation Challenge as we write this. The board is a few days past its last re-grade — a first billing correction has been merged and applied — and it is still shaking out; the measuring stick underneath it is under active refinement, with more fixes in review right now. That is a well-run young benchmark doing exactly what it should, so nothing here is a complaint about any snapshot, and nothing here is about anyone else's entry. It is simpler than that: we believe something specific about this problem, and we would rather say it now, in public, with a date on it, than claim it after the fact.
The bet: a floor
The task is to predict, cheaply, the average behaviour of a large random neural network — without running it anywhere near enough times to just measure the answer. It looks to us like this problem has a floor: for any honest, budget-respecting method, the error cannot be pushed below a certain level, no matter how clever the method is. We spent a week trying to get under it from every direction we could think of — richer features, better sampling, restructured sampling, biased shortcuts. Everything failed in an instructive way. The genuine improvements we found only brought us closer to the floor, never through it.
So we are betting on a number. In the challenge's own units the floor lands at roughly 3.7e-7 (its meter-independent form is a per-pair variance of about 0.012), with a little spread — about 3.4 to 3.9e-7 — depending on which unseen networks you happen to draw. That is where our own best honest estimator lands on networks it has never seen, and where we expect careful methods to converge. Part of any low score on a fixed, visible test set — ours included — is a favourable draw rather than durable signal. That is the nature of any feedback board, and we do not exempt ourselves from it.
We are stating the number before our full technical write-up, so the claim is on the record ahead of the proof. The argument is short, and its structural core is machine-checked in Lean 4 against Mathlib, with no gaps and no custom assumptions — every step, including the last three classical Gaussian integrals that had been standing in as assumptions, is now proved from the library's standard foundations and confirmed by the prover's own axiom audit. What the proof certifies is the shape of the argument; the floor's numerical size is measured, not proven, and we say so. And genuinely: if someone clears the floor on unseen networks, that is the most interesting possible outcome — we would love to be wrong about this one.
The prediction
The rules are explicit that the live board does not decide the ranking. The prize is settled on unseen networks: after the submission phases close, each team's one designated entry is re-graded on a fresh, private set of networks generated from a held-out seed, with results in early October. Our prediction, filed today:
One — the public and private rankings will differ noticeably. That is by design, and we expect it to be visible.
Two — careful honest methods will converge toward the floor on unseen networks, from both sides, because for the honest, budget-respecting class the floor is a property of the problem, not of any particular test set. A biased method whose error happens to generalise could in principle sit below it — most plausibly if the held-out networks change shape — which, as we say below, is the outcome we would find most interesting.
Three — our own standing on the unseen evaluation should move up substantially from around #80, into the cluster that sits at the floor — with one honest caveat. That rise rides on two things: entries that were fitted to the visible set regressing on networks they have never seen, and the remaining gaps in the measuring stick being closed before the final grade. The second is not guaranteed. At least one accounting gap looks open as we write, and if it is still open at the re-run, some scores rest on work the meter does not fully count, a fresh set of networks will not demote those on its own, and our move up is muted to that extent. That is a genuine dependency, not a hedge we can resolve from here.
We are not predicting we finish first. If the floor is real, the field that reached it is a tight pack separated by draw luck, and some of today's leaders may hold their ground just fine. Concretely: if we do not clear roughly the top third on the unseen re-run, treat this prediction as failed. What we are calling in advance is the direction of our own move — up — and the floor itself.
Why post it
We made a wager this week: put the effort into the unseen evaluation rather than the visible one, because whether a method holds up on networks it has never seen is the number we actually care about. Staking the prediction in advance just keeps us honest. If we are wrong, this post stays up, and you get to say so.
And part of why this contest is worth this much of our effort: it is a miniature of a question we care about far beyond it — how much can you know about a system from its structure, before you run it? The floor, if it holds, is a small piece of that answer. Phase 1 closes in days; the decisive fresh re-run lands in about two months. We will check back in either way. Good luck to everyone on the board — it has been a genuinely strong field.