← Back to the essay
Working paper · Proposed open standard · v0.1

Meaningful Human Oversight of AI: A Proposed Measurement Standard

Chenna, M.
Sanctity · Amsterdam, Netherlands
2026 · Working paper · License: CC BY 4.0

Abstract

Regulation increasingly requires a human to oversee high-risk artificial-intelligence systems, yet provides no way to determine whether that oversight is ever exercised. Because an oversight power that is never used is externally indistinguishable from one that does not exist, compliance can be satisfied while command is absent. We propose the Meaningful Override Rate (MOR): among decisions a human was positioned to change, the fraction the human actually changed and the change was honored. We specify two companion measures, contestation latency and reversal validity; a five-part bar a reported figure must clear to be citable; and a fixed reporting format. We argue the standard should be published before any measurements, so the instrument does not accrue to whoever it flatters. This is version 0.1, released for use, critique, and revision.

Keywords: human oversight, AI governance, EU AI Act Article 14, human-in-the-loop, automation bias, accountability, measurement.

Introduction

Frameworks for trustworthy AI, including Article 14 of the European Union's AI Act as it comes into force, require that certain systems remain under human oversight.1 Such provisions specify that a human be able to intervene. They do not, and largely cannot, establish that intervention ever occurs. A reviewer who holds a veto but never exercises it presents, to any external audit, the same evidence as a reviewer who holds no veto at all. Oversight therefore risks becoming a certified property of a system rather than an observed behavior of a person.

This paper proposes a minimal, auditable measure of exercised oversight, together with the conditions under which a reported value should be believed. The aim is not to replace legal requirements but to give them an instrument: a way to distinguish oversight that is real from oversight that is performed.

Definitions

Reviewable decision

A decision produced or proposed by an automated system that a designated human was both positioned and permitted to change before it took effect. Reviewable decisions form the denominator of the measure. Decisions a human could not, in practice, have altered are excluded.

Meaningful override

A reviewable decision in which the human changed the system's output, on the human's initiative, and the change was honored by the deploying organization. A recorded presence, an elapsed interval, or a countersignature is not, by itself, a meaningful override. Authorizing is not overseeing.

The metric

Meaningful Override Rate (MOR). The fraction of reviewable decisions that were meaningful overrides.
MOR = (meaningful overrides) / (reviewable decisions) (1)

Companion measures

MOR reported alone is fragile: it can be inflated by trivial changes or depressed by conditions that make disagreement impossible. Two companions constrain its interpretation.

Contestation latency

The realistic time and context a reviewer had to form and register disagreement. As latency approaches zero, review approaches a turnstile, and MOR loses meaning at any value.

Reversal validity

The rate at which human overrides were, in retrospect, correct relative to the system's original output. A high override rate that is frequently wrong is a distinct failure, not evidence of oversight.

Validity criteria

A reported MOR value should be treated as citable only if it satisfies all of the following. A figure short of these is an anecdote in the form of a percentage.

  1. Stated sample. At least approximately one hundred reviewable decisions.
  2. Pre-registered meaning. The criterion for a meaningful override is fixed and published before results are examined.
  3. One named workflow. A specific task in a specific deployment, not a category of organizations.
  4. Explicit denominator. Overrides divided by reviewable decisions, without selection on the outcome.
  5. Inter-rater agreement. More than one adjudicator concurs on what counted as meaningful.

Reporting format

To keep figures comparable and self-limiting, report MOR in the following shape, with its scope caveat attached:

On [workflow X], MOR was [N] percent across [M] reviewable decisions.
Meaningful override was defined as [criterion]. Inter-rater agreement: [k].
This result is not claimed to generalize beyond [workflow X].

Presence is not command

A recurring design shortcut is to evidence oversight by logging human presence or human confirmation. Such logs are tidy and auditable, and they measure the wrong quantity. A confirmation that is always given carries no information about command. Where a reviewer can technically object but never does, and is nonetheless accountable for the system's errors, the human occupies what Elish terms a moral crumple zone: a locus of blame without a locus of control.2 Sustained non-use of a veto is also the expected result of well-documented automation bias, the tendency of operators to defer to reliable automation.3 MOR is constructed to detect exactly this condition.

Status and versioning

This is version 0.1. It is published before any measurements are reported, by design, so that the definition of the instrument cannot be adjusted to suit results already in hand. Reference implementations and measured values will be released separately, each accompanied by its method and limitations, and only when they meet the criteria of Section 4. A measurement standard is fair only if it is fixed before the measurement.

A falsifiable claim

We state one testable expectation. Across real deployments, measured under the criteria above, the meaningful override rate of systems currently marketed as offering human oversight will be found to sit near zero. A body of deployments in which humans routinely and validly overrule the system would falsify this claim, and such evidence, if produced, will be recorded here in revision.

References

  1. European Parliament and Council of the European Union (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), Article 14, Human oversight.
  2. Elish, M. C. (2019). Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction. Engaging Science, Technology, and Society, 5, 40-60.
  3. Parasuraman, R., and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2), 230-253.
  4. National Institute of Standards and Technology (2023). AI Risk Management Framework (AI RMF 1.0). NIST.
Cite as: Chenna, M. (2026). Meaningful Human Oversight of AI: A Proposed Measurement Standard (MOR v0.1). Sanctity, Amsterdam.
Version 0.1 · Licensed CC BY 4.0 · Free to use, run, and cite · Revisions recorded at manjchenna.com/essays/meaningful-override-rate