The score is the harness
A model score without a harness name
is a press release. ARC Prize now
reports two numbers on purpose. The
Standard harness is the comparable
one: every provider gets the same
interface, and the model has to
decide what to keep in visible notes.
The Provider Adapter is a different
system under test — weights plus
opaque reasoning state plus the
vendor’s compaction. 99.9% is real.
So is 62.7%. Putting only the first
on a slide is how you nerd-snipe
yourself into buying memory you did
not budget.
The same split showed up in Codex.
Compaction is lossy on purpose. It
is also where failed patches go to
die. Notes that survive a window
boundary, plus search over earlier
tool output, are an adapter you can
see. Treat them as part of the
product: version the config, log
whether notes were on, and do not
compare a notes-on Astra run to a
compact-everything Fable run. Flash
3.8 is the other side of the same
coin. A 13-second HTML loop at 1.8
cents is a harness you will actually
repeat. A gated cyber twin is not
“the same model with more thinking.”
Stop when you can name the harness,
say whether hidden state crossed a
request boundary, and write down the
score without that adapter. Hunting
for one more leaderboard cell after
that is collecting stickers.
Checklist
-
Never cite a percentage without
the harness name (Standard,
Adapter, PRO-LONG, your CI).
-
If a vendor publishes one number
and the eval lab publishes two,
file both.
-
Log whether opaque state, notes,
or compaction ran. That flag is
part of the result.
-
Prefer searchable prior windows
over a single summary when a
failed fix still matters.
-
Return tool results as objects
with names, not arrays the model
has to count.
This week: pick one score you have
repeated in chat, write the harness
name next to it, and find the number
without the adapter.