Essay · The evidence layer
Benchmarks Are Ads. Records Are Evidence.
A benchmark score is a snapshot, taken by an interested party, on questions that leak into training data. A record is longitudinal, adversary-aware, and boring. Buyers keep confusing the brochure for the audit.
Every model launch arrives with a chart, and every chart says the same thing: we are above the line that matters. I do not think the charts are fraudulent. I think they are advertisements, in the precise sense: claims produced by the seller, at a moment of the seller's choosing, on terms the seller helped set. Markets know what to do with advertisements, enjoy them and verify elsewhere. The strange thing about the AI market is that elsewhere barely exists, so the ads are being read as if they were audits, and procurement decisions worth millions rest on the difference.
What a benchmark score actually certifies
Read a leaderboard number carefully and its jurisdiction shrinks. It certifies performance on a fixed question set, at one moment, on a serving configuration the vendor prepared for the occasion. The famous sets leak into training data, which is not cheating so much as osmosis: the internet is the training corpus and the benchmark is on the internet. Vendors tune for the boards they know matter, Goodhart doing its patient work. And the number is silent about everything a buyer actually lives with: variance, the tail, degradation under load, and whether the model scored is the model served, which readers of The Model on the Invoice will recognize as never guaranteed.
The snapshot problem, again
Even a perfectly clean benchmark shares the defect of every snapshot: it tells you about a day. Models behind APIs change, quietly and often, and a launch-day score has no expiry printed on it. A buyer comparing model A's March number to model B's July number is comparing two photographs taken in different weather and calling it a race. The question that matters for anyone building on a model is not where it ranked once but how it behaves over time: is it stable, is it drifting, did the step change last month land in the domain your product depends on. Snapshots cannot answer questions shaped like that. Only a series can.
What a record does differently
A record, in the sense I use across this series, has four properties a benchmark lacks. Longitudinal: the same measurements, repeatedly, forming a series that shows change rather than position. Adversary-aware: batteries rotated and refreshed against contamination, because the instrument assumes it is being gamed. Independent: run by a party with no model in the race and no launch to promote. And tamper-evident: results chained so the history cannot be quietly improved, which matters exactly as much as you think it does the first time a series turns unflattering. This is the design of Modelometer: behavior across eight domains on a steady cadence, identity checked daily, every run on a hash-chained public ledger. Boring by design. Evidence usually is.
What buyers should ask for instead
Next procurement cycle, skip the chart argument entirely and ask for series. Show me this model's behavioral record for the last two quarters, from a source you do not control. Show me its stability in the domains my product touches. Show me the change events, dated. A vendor who can point to an independent record is showing you evidence; a vendor who can only re-send the launch chart is showing you an ad, twice. And if the answer is that no independent record exists for their model, that is itself the finding: you are being asked to build on a foundation nobody is watching.
The chart and the ledger
None of this makes benchmarks worthless; ads carry information, and a model atop every board is probably strong. The error is category, not content: the chart answers how impressive, the ledger answers how dependable, and a business runs on the second. We learned this distinction everywhere else. Nobody buys a company because its own deck says revenue is up; they ask for the audited statements. The AI market is simply young enough that the decks still work. They will stop working the way they always stop working: one expensive dispute at a time.
The transition will be generational rather than dramatic. Nobody will announce the end of leaderboard procurement; it will simply become slightly embarrassing, the way citing a company's own press release as due diligence is slightly embarrassing, and the buyers who moved to records first will quietly stop being surprised by their vendors. Markets do not abandon advertisements. They just stop mistaking them for audits, one burned budget at a time.
Read on
Why your own eval is also a snapshot: Your Eval Passed. Your Product Failed. The institution the market is missing: The Assay Office for AI. Whether size itself proves anything: model size is not trust.