AI

The Watcher Also Judges — The Ecological Validity of Human Supervision in Closed Cyber-Physical Loops

Published on Aug 13, 2026·10 min read
The Watcher Also Judges — The Ecological Validity of Human Supervision in Closed Cyber-Physical Loops

The Watcher Also Judges — The Ecological Validity of Human Supervision in Closed Cyber-Physical Loops

Human Supervision as an Evaluative System

If the evaluation of artificial intelligence requires ecological validity — assessing the model's behavior in real-world context rather than under idealized laboratory conditions — a symmetrical question arises: do the very mechanisms of human supervision retain the ecological validity needed to evaluate the system in real time? When transposed onto the supervision of closed cyber-physical loops (AI+IoT), Egon Brunswik's conceptual framework reveals a direct inversion. The "judging organism" becomes the human supervisor, while the "environment" becomes the cyber-physical loop itself — sensors, actuators, and the AI model that decides between them. The distal variable ceases to be the state of the world that the AI attempts to infer and becomes instead the real state of the system being executed: whether the equipment is in fact degrading, whether the decision made minutes earlier was adequate, or whether an incipient anomaly exists without an alert. Proximal cues are reduced to dashboards, periodic reports, and occasional alerts reaching the supervisor.

There is, however, a crucial difference from the original perceptual model: whereas the cues of natural ecology are given by nature, supervisory cues are designed, through engineering choices about sampling rate, aggregation, and alert thresholds. The ecological validity of supervision thus comes to depend on the actual correlation between what the dashboard shows and what is happening in the physical system. An additional layer of risk is introduced: the trust a supervisor places in a cue can remain intact long after that cue's validity has begun to decline — a lag upon the lag, one that the human-factors literature associates with alarm fatigue and automation complacency.

The Erosion Mechanism: Institutional Void vs. Active Resistance

As the cadence of the cyber-physical loop accelerates to the millisecond scale, the rate at which the system changes state exceeds by several orders of magnitude the rate at which a human processes that change — and from this discrepancy arises a phenomenon close to temporal aliasing. From the standpoint of Ross Ashby's Law of Requisite Variety, the variation in system states drastically exceeds human cognitive bandwidth, estimated at between 10 and 50 bits per second. For information to reach the human in a comprehensible form, the system is forced to perform aggressive data compression, breaking the correspondence between the observed cue and distal reality.

At this point it is essential to separate two mechanisms that produce the same symptom of lag but demand radically different explanatory hypotheses:

  1. Institutional void: a direct architectural property of two loops operating in incompatible time regimes. The cadence of the automated loop exceeds the feasible frequency of human review, which is limited by cognitive and organizational costs. This mechanism requires no agency whatsoever on the part of the system and is demonstrable in a classical PID controller, without a single line of learning code.

  2. Active resistance (instrumental convergence): described by Omohundro and Bostrom, this posits that a goal-directed system develops self-preservation subgoals to avoid being shut down or corrected. Active resistance, however, cumulatively requires three conditions: an open-ended, poorly bounded objective, the capacity for strategic reasoning about its own supervision, and some access to the mechanisms of control.

An industrial IoT anomaly-detection system, with a narrow loss function and no means of acting on its own safety controls, satisfies none of these three premises. To explain the observed supervisory gap by invoking active resistance is to choose the heavier, more speculative hypothesis when the institutional void — arising from a mere mismatch of cadences — is the parsimonious explanation, sufficient and directly observable in systems with no pretense of agency whatsoever.

The Disconnect on the Ground: Empirical Evidence

This erosion driven by institutional void is already manifest in high-cadence operational settings. In the Citigroup trading incident of May 2022 (subject to a regulatory fine in 2024), an erroneous manual entry resulted in the algorithmic execution of $1.4 billion in a matter of minutes. The real-time monitoring team failed to escalate a wave of 284 alerts, and the risk team only reacted around 20 minutes after the orders were cancelled — demonstrating the loss of correlation between human monitoring and the market's actual state.

Likewise, during the Iberian Peninsula blackout of April 28, 2025, a cascade of automatic generation trips and overvoltages unfolded in roughly 78 seconds. Human operators, bound by standard procedures and the manual cadence of activating compensation resources such as shunt reactors, saw the system's dynamics outpace human reaction capacity. Similarly, in cybersecurity operations centers (SOCs), 2026 reports document breakout times as fast as 51 seconds and exfiltrations within 72 minutes; against this cadence, human triage queues and reports typically take tens of minutes to hours to process, describing operational states that have already ceased to exist.

Engineering Mitigations and the Redesign of the Paradigm

Control engineering has attempted to mitigate this lag through several mechanisms. Exception-based auditing reduces data load under normal conditions, but when the exception fires at high cadence it triggers cascades of alarms upon an operator who has already lost accumulated situational awareness (out-of-the-loop). Statistical Process Control (SPC) presupposes slow variation and known distributions, failing in the face of nonlinear emergent AI behaviors. Staged autonomy — shifting the system between different degrees of human intervention according to estimated uncertainty — runs into Bainbridge's paradox: slowing the process to await human validation tends to defeat the very purpose of a high-cadence system, and handing control back to the human precisely at the moment of greatest uncertainty is also the moment of greatest probability of error. Circuit breakers and kill switches interrupt the process to guarantee functional safety, but they sacrifice operability and do not restore the validity of real-time observation. The most robust approach rests on safety envelopes and Control Barrier Functions (CBFs), which mathematically constrain the space of actions permitted to the AI.

Even so, as recent standards such as ISO/IEC 42001:2023 and the EU AI Act (Article 14) acknowledge, these safeguards start from the premise that near-instantaneous operational supervision is a physical impossibility whenever the process's tolerance window is shorter than human reaction time. The paradigm of supervision is thus forced to shift from the operational scale to a structural and normative one — measured not in milliseconds but in cycles of policy review and safety-envelope revision. The human supervisor stops acting within the fast loop and instead designs safety envelopes, tunes predictive digital twins that project future trajectories, and defines risk policies a priori.

This argument does not apply universally; it is conditional on domains with tight cyber-physical loops. In ordinary organizational, financial, or strategic decisions, where the two rhythms are comparable, the idea simply has no traction. There remains, moreover, a rhetorical risk worth naming without resolving: the phrase "supervision loses ecological validity" can itself be appropriated, in good or bad faith, to justify less supervision rather than redesigned supervision. The diagnosis proposed here concerns a fact — a correspondence that degrades — not what should be done about that fact; to conflate the two would be to repeat, on the side of supervision, the very error this article has already flagged on the side of AI: treating a finding about validity as though it were, on its own, a normative license. It remains unresolved, and perhaps must remain so, whether there exists a way to redesign supervision that is not itself just one more proximal cue — one more dashboard — positing trust where correspondence has not yet been demonstrated.

From Output to Outcome: the Distal Criterion as a Requirement of Ecological Validity

The distinction between output and outcome, current in program evaluation and in logic models of intervention (input → activity → output → outcome → impact), is not a mere terminological convention borrowed from management. It encodes an ontological difference that Brunswikian psychology had already formalized decades earlier through the lens model: output corresponds to the system's proximal judgment — what it produces directly and under its own control, from the processing of available cues — whereas outcome corresponds to the distal criterion in the environment, the state of the world that judgment was meant to predict or influence, and which is no longer under the system's direct control. Achievement — ecological validity in the strict sense — is defined by Brunswik precisely as the correlation between these two terms, not as an intrinsic property of the output.

This convergence between a vocabulary originating in program evaluation (consolidated, for example, in the W. K. Kellogg Foundation's logic models and in the OECD-DAC evaluation criteria) and Brunswik's conceptual apparatus is not an isolated disciplinary coincidence: it finds a particularly robust parallel in Donabedian's (1966) structure–process–outcome model for evaluating quality in healthcare. Donabedian argues, in an entirely independent domain, that evaluating the clinical process — the equivalent of output — is necessary but insufficient, and that a health system's ultimate validity is established only in the patient's outcome. This interdisciplinary convergence reinforces, by triangulation, the article's central thesis: if fields as distinct as perceptual psychology, program evaluation, and clinical medicine independently arrive at the same structural distinction, this constitutes evidence that output is not reducible to outcome by definition, but is rather a correlated and fallible variable.

The implication for evaluating AI systems — and, with particular acuity, AI+IoT systems — is direct. A benchmark typically measures the correspondence between a model's output and a reference label in a dataset: whether the classification matches the correct label. This is a measure of fidelity internal to the evaluation system, not a measure of ecological achievement. Ecological validity requires, instead, asking whether that output, once inserted into the real environment — with its own distribution of scenarios, instrumental noise, and human and mechanical mediators — actually produces the desired distal state. In a system such as the one described in the previous section on closed cyber-physical loops, the output is "anomaly correctly classified according to the test criterion"; the outcome is "failure actually averted" or "downtime reduced." Between the two lies a chain of causal mediators — the operator's interpretation, response latency, the adequacy of the supervisory interface — that can systematically cause the correct output to diverge from the desired outcome, exactly as Brunswikian theory predicted for human perceptual judgment. This is the theoretical foundation of the argument already developed here: that the supervisory panel is itself an evaluative system subject to requirements of ecological validity — it is the link that converts output into outcome, and a failure of ecological validity at that link compromises the entire chain, regardless of the quality of the upstream model's output.

It is worth anticipating the obvious objection: outcomes are typically noisier, more delayed in time, and multicausally confounded (by maintenance policies, operator behavior, variation in the physical system's load) — which explains, but does not epistemically justify, the widespread methodological preference for evaluating outputs. This tension is structurally analogous to Brunswik's own critique of unrepresentative laboratory experimental design: the convenience of measuring what is easily measurable carries a systematic cost to validity that only manifests once the system is transposed into the real ecological environment. Recognizing this tension explicitly — rather than dissolving it — is, moreover, consistent with the treatment of epistemic triangulation already developed elsewhere in this article: evaluating outcomes does not replace evaluating outputs, but requires that it be complemented through multiple convergent sources of evidence, precisely because no single outcome measure is immune to confounding.