Chenna, M.Essays

Essay · The accountable human

Human-in-Command

Regulators are converging on a phrase: human-in-command. It is better than human-in-the-loop, and it will be diluted within a year unless somebody pins it down. Command has four conditions, and all four are measurable.

scroll to read

Essay · The accountable human

Human-in-Command

Regulators are converging on a phrase: human-in-command. It is better than human-in-the-loop, and it will be diluted within a year unless somebody pins it down. Command has four conditions, and all four are measurable.

Chenna, M. · Founder, Sanctity · Amsterdam · August 20, 2026

Watch the language of AI oversight evolve and you can see an idea trying to be born. First it was human-in-the-loop, which promised only presence: a person somewhere in the pipeline, doing something. Then meaningful human oversight, which admitted presence was not enough but declined to say what would be. Now the sharper phrase is arriving in supervisory and standards conversations: human-in-command. I want to claim it before it gets diluted the way its predecessors were, because command, unlike presence, can actually be specified. And anything that can be specified can be measured.

Presence is not command

A person watching a dashboard is present. A person clicking confirm is present. Neither is in command, and every incident review that ends with a human was in the loop is exploiting the gap between the two. I have written about the confirm button as liability laundering: the decision was made by the model, the accountability was routed to the person nearest the button. Loop language cannot distinguish that from real control. Command language can, because command has conditions, and they either hold or they do not.

The four conditions

A human is in command of an AI system when four things are simultaneously true. Authority: they can reverse or halt the decision, and the reversal sticks without heroics. Information: they can see what the system saw and what it concluded, in a form a human can actually judge, not a wall of confidence scores. Time: the decision waits for them, or can be unwound after them; a five second window before auto-approval is presence wearing a costume. Consequence: their name attaches to the outcome, so the judgment is exercised with the care that ownership produces. Remove any one condition and command quietly degrades back into theater. Most deployments I have examined fail at least two.

Every condition gets faked

The faking is rarely malicious; it is what happens when a compliance requirement meets a delivery deadline. Authority gets faked with an override switch nobody has ever used in production. Information gets faked with explainability dashboards that explain everything except what mattered. Time gets faked with review queues sized so that real scrutiny would collapse the backlog, which guarantees scrutiny is not real. Consequence gets faked with committee sign-offs, where a shared name is no name. The system passes its audit. The command does not exist.

Command can be measured

Here is where this stops being philosophy. Each condition throws off evidence if you bother to collect it. Authority shows up in the Meaningful Override Rate: of the decisions a human could have reversed, how many did they reverse, and did the reversal hold. Information and time show up in the texture of the log: how long the human actually had, what they were actually shown. Consequence shows up in whether a named person appears in the record at all. And all of it presumes one more thing, the thing this site keeps returning to: that the system being commanded is the system you think it is. A human cannot command a model that changed last night without a record. The machine leg of command is an independent record of what the machine was.

Why the phrase matters now

Phrases set procurement. Human-in-the-loop appeared in a thousand contracts and obligated approximately nothing, because nobody defined the loop. If human-in-command enters contracts the same way, undefined, it will die the same death. But if buyers write the four conditions into their requirements, authority, information, time, consequence, each with its evidence, then the phrase becomes a specification, and specifications get built. That is the entire difference between the words we have had and the word we need. Command is not a vibe. It is four tests, and your deployment either passes them or it does not.

One prediction to hold me to. Within a few procurement cycles, some standards body or supervisor will publish a definition of human-in-command, and the deployments that measured the four conditions early will discover they are already compliant, while the deployments that treated the phrase as decoration will discover they have been accumulating evidence of its absence. Language in this field has a way of hardening into requirements. Choose your words while they are still cheap, and be able to show your evidence when they are not.

Read on

The measurement of the human leg: the Meaningful Override Rate. Why the confirm button fails: the confirm button is not oversight. The machine leg: Your AI Changed Last Night. Prove It.