Flagship essay · The evidence layer
Your AI Changed Last Night. Prove It.
The model beneath your product is not the one you tested. Providers swap, tune, and reroute in silence, and your only record is a forum thread full of vibes. The case for an independent record of machine behavior.
You did not build on a model. You built on an endpoint. The difference matters more than any benchmark. A model is a fixed artifact; an endpoint is a promise that whatever answers today will behave like whatever answered yesterday. Nobody notarizes that promise. The provider can quantize it, distill it, reroute it to a cheaper sibling, retune its refusals, and your product inherits the change at the exact moment your users do. When it happens, and it happens, your evidence is a feeling: it got worse. Feelings do not survive a dispute, and they certainly do not survive an audit.
The silent swap
Every builder on frontier APIs has lived this. A workflow that passed evaluation in spring starts failing in summer. The tone shifts. The refusals move. The same prompt returns a different shape of answer, and nothing in your code changed. Ask the provider and you get a version string that has not moved and a changelog that does not mention it. The community calls it nerfing. I call it what it is: an unwitnessed change to a system other people have built businesses on. Not malice, usually. Cost pressure, safety tuning, capacity routing, a quiet quantization that saves them money and costs you accuracy. But unwitnessed is the operative word, because everything downstream of that change, every decision your product made with it, every log you kept about it, now describes a system that no longer exists.
The changelog is a press release
Here is why every existing answer fails the moment money is on the table. The provider's changelog is written by the party with the strongest incentive to say nothing changed. The status page measures whether the endpoint is up, not whether it is the same. And your internal evaluations, however good, watch only your own traffic: they cannot distinguish your prompt drift from their model drift, and when you take them to the provider it is your word against theirs, your harness against their denial. Evidence you produce about your own dependency is testimony. What a dispute needs is a witness.
What the EU AI Act expects you to already know
This stops being a quality complaint and becomes a compliance problem the moment your system falls under the EU AI Act. Deployers of high-risk AI systems are required, under Article 26, to monitor the operation of the system and to keep the logs it generates; Article 12 exists so that the record of what the system did actually gets kept, and Article 72 obliges providers to run post-market monitoring of how the system behaves in the real world. Every one of those duties quietly assumes something nobody checks: that the system being monitored today is the system that was assessed yesterday. There is a sharper edge still. The Act treats a substantial modification as an event that changes who bears provider obligations. If the model beneath your product materially changed and nobody recorded it, you cannot say whether your compliance file still describes the system you are running, and you cannot even establish when the divergence began. Monitoring a system nobody witnesses is not monitoring. It is paperwork about a memory.
Oversight has two legs, and both are unmeasured
I have argued that most human oversight of AI is theater: a person present, nothing measured, and the Meaningful Override Rate is my proposal for measuring the human leg. This essay is about the other leg. Oversight assumes you can see the thing you oversee, and today the machine leg is unwitnessed: the system under your oversight can become a different system overnight, with no record that it did. Every model is an opinion, and the opinion is being revised without a hearing. You cannot oversee what you cannot see change.
What an independent record looks like
The instrument this requires is boring on purpose, and it has five properties. Independent: run by neither the provider nor the buyer, so its record is admissible to both. Continuous: identity checked daily, behavior measured on a steady cadence, because a snapshot proves nothing about drift. Comparable: the same questions asked of every model, so change is visible against a baseline and against peers. Tamper-evident: every measurement chained to the last, so the record itself cannot be quietly rewritten. Public: because evidence that only one party can read is just another changelog. Think seismograph, not survey. Nobody asks a seismograph how it feels about the ground.
This is what Modelometer is
So I am building the witness. Modelometer is an independent measurement bureau for AI model behavior, live in beta: it checks serving identity daily, measures behavior weekly across eight domains, runs more than fourteen thousand model calls a week, and writes every run to a hash-chained public record. It does not know why a model changed; it establishes that it did, when it did, and by how much, which is the part no one else in the dispute can credibly supply. Honest scope: a beta, an instrument still earning its calibration, and I publish its methods so you can attack them. That is what separates a measurement from a marketing claim.
If this is your desk
Three desks carry this risk today. If you deploy a high-risk system under the EU AI Act, your monitoring duty needs evidence that is not your own testimony. If your product carries an SLA that depends on a model you do not control, your exposure changes every time that model does, and today you find out from your customers. And if your audit trail cites a model by name, an independent behavioral record is the difference between a fact in your file and an assumption in it. The companion problem, whether the model on your invoice is the model answering your calls at all, is its own essay: the model on the invoice.
What would prove me wrong
If providers began shipping signed behavioral diffs with every change, verified by third parties, and procurement accepted those as evidence, the case for an outside bureau weakens to a formality. I would take that outcome happily: it is the same argument won by other means, the way financial audit stopped being optional. What I do not expect is the status quo surviving contact with the first serious dispute over a silent model change. Somebody will ask what changed and when, and the honest answer today, from everyone in the room, is that nobody kept the record.
The bet
My bet is procurement. Today buyers ask which model you use. Soon they will ask the question that actually carries risk: show me the record of what has changed beneath your product, and who keeps it. Teams that can answer with an independent record will wear the change and move on. Teams that cannot will wear the blame, because a system nobody witnessed is a system nobody can defend. The oversight argument and the evidence argument are one argument: keep a human in command, and keep a record the human can command with. I intend to be holding that record.
Read on
The companion piece: The Model on the Invoice, on whether you are even being served the model you pay for. The human leg of the argument: human oversight is mostly theater and the Meaningful Override Rate. The premise underneath both: every model is an opinion. The instrument: modelometer.com.