Formal Architecture of the Chain
The symmetric comparison between human judgment and AI judgment — placing an for the human judge and an for the system side by side, as though they were two estimates of the same kind of quantity observed at equivalent positions in the causal chain — fails for a structural reason that no statistical correction can fix: the human and the AI typically do not occupy the same place between the environment and the consequence. The human receives cues from the environment and can, typically, act on it directly; the AI receives cues from the environment and produces a recommendation that only becomes action through a human who accepts, modifies, or rejects it. Written minimally, the first case is Human → Environment; the second is AI → Human → Environment. Treating these two chains as term-by-term comparable imposes a symmetry that the system's own architecture already contradicts.
The alternative proposed here does not abandon Brunswik's formal apparatus — on the contrary, it extends it, chaining two successive lens models rather than comparing two isolated ones:
We adopt the classic Lens Model notation, where measures the linear predictability of the distal criterion from the available cues, the consistency of the agent's response policy — the degree to which its responses are predictable from the cues — and the matching between the linearly predicted components of the environment's and the agent's models: technically, , the correlation between the values predicted by the two linear models, not a direct comparison of their weight vectors. Throughout this article, we read informally as "cue-weighting correspondence" — but that reading is an interpretive shorthand that holds only under additional conditions (common model specification, comparable scale, treatment of cue correlation); under strong collinearity, very different weight vectors can produce similar predictions, and vice versa. The term , already defined in Section 0, captures the residual correlation between the two models' residuals — never the agent's consistency, which is 's role.
In the first stage, the AI system occupies the position of the classic Brunswikian judge: it receives a set of cues from the environment — sensor readings, time series, derived signals — and produces a judgment about that environment. This stage admits the usual decomposition: measures the predictability of the environment given the set of cues available to the AI; measures the matching between the linearly predicted components of the AI's model and the ecological model — interpretively, the extent to which the system's observable policy aligns with the cue structure that predicts the criterion; measures the degree to which the system's output is predictable from the modeled set of cues. Paraphrase-invariance tests, or tests across semantically equivalent configurations of the same cues, can complement this measure — notably as a safeguard against an inflated merely by deterministic decoding — but they are not identical to it. A note on language is worth making, and it is not decorative: we call here the consistency of an observable input-output decision policy, never "the AI judge's consistency" — this vocabulary difference is precisely what avoids reimporting, through the back door, the assumption of cognitive agency that the ontological objection identifies as illegitimate. We are not claiming that the system judges; we are claiming only that its input-output behavior does, or does not, exhibit the statistical properties this decomposition describes.
It is in the second stage that the chain stops being a mere repetition of the first. The AI's output — the recommendation, the confidence score, the generated report — is not, for the human, the distal criterion to be inferred; it is, itself, an additional proximal cue, sitting alongside the other cues that reach the human directly from the environment through the supervision dashboard. In Lens Model 2, we treat the AI's output as an additional cue within the set of cues available to the supervisor. The human forms their judgment from this composite cue. The same family of parameters can be estimated at the second stage, though its interpretation is distinct because is a cue produced by an earlier stage of the chain, not a direct measurement of the environment: measures the predictability of the distal criterion given the set of cues available to the human, including ; measures the consistency of the supervisor's response policy; measures the matching between the linearly predicted components of the human's policy and the ecological model, considering the set of cues — including .
It is precisely here, in this distinction between real and perceived reliability, that the automation bias and complacency already identified in the previous article find, for the first time, a formal expression. Over-reliance on the AI's recommendation is not measured at the level of global , but at the level of the specific weight assigned to the cue in the human's regression policy. This comparison requires a qualification that is not decorative: since is typically derived, at least in part, from the same environmental cues the human already observes directly, is a conditional coefficient — its value depends on the collinearity between and the remaining cues, and on the exact specification of the model. Over-reliance should, therefore, be assessed at the level of the conditional use of : the weight estimated in the supervisor's policy, relative to the ecologically valid incremental value of , conditioned on the same set of direct cues from which the recommendation is derived. A systematically higher than that conditional ecological weight is evidence of over-weighting; a systematically lower value is evidence of under-weighting. The correct reading is, therefore: over-reliance is a phenomenon at the level of cues' conditional weights, not at the level of the global matching captured by .
This distinction between aggregate correspondence and cue-level correspondence invites us to explicitly separate two failure modes in Lens Model 2, which the human-AI teaming literature frequently treats as a single phenomenon but which the decomposition proposed here allows us to distinguish precisely. Aggregate matching failure occurs when is globally low — the supervisor fails, on the whole, to get the environment's weight structure right, distributing their attention poorly across all available cues, including but not limited to . This is the kind of failure a conventional auditing dashboard, reporting only , is well-equipped to detect.
Localized weighting failure is structurally different and more insidious: the supervisor gets the global weight structure reasonably right — remains moderate or even high — but systematically distorts the conditional weight assigned to a single cue, typically , compensating for that distortion through adjustments to the others. A supervisor who consistently over-weights the AI's recommendation but compensates by under-weighting direct environmental cues can produce a perfectly healthy while accumulating a structural dependence on the AI that only reveals itself once that compensation becomes impossible — for instance, when a scenario drifts outside the usual distribution, at precisely the moment when direct environmental cues become less informative and the AI's own reliability also degrades. This is, in systems-safety language, a correlated, latent failure: invisible to the aggregate indicator until the conditions that were masking it cease to hold.
The implication for the design of the auditing tool is direct, and not merely academic: an audit report that reports only is structurally insufficient to detect latent over-reliance. The audit proposed in this article requires, by construction, the reporting of conditional weights at the level of individual cues — not just aggregate matching — precisely because it is at the cue level, not the summary level, that excessive dependence on the algorithmic recommendation becomes visible before it becomes consequential.
The closing of the chain is what this architecture gains, and what the symmetric comparison could never offer — not an algebraic decomposition already derived for the joint achievement, , but the possibility of formulating separable diagnostic hypotheses about where correspondence with the real environment began to fail. Among the most central: (1) low — the AI fails to correspond to the real structure of the environment; (2) low — aggregate matching failure at the human stage; (3) misaligned from 's conditional ecological weight, with preserved — localized weighting failure, consistent with over-reliance or under-reliance; (4) a poorly defined, unobservable, or endogenous criterion from the outset, whose discussion — in its weak and strong forms — is the subject of the following section. This list is not exhaustive: low (environment poorly predictable from the AI's cues), low (unstable system policy), low (inconsistent supervisor), and each stage's residual components remain relevant diagnostic dimensions that a complete audit should also report.
It is worth flagging here, even though its development is reserved for the following section: the final arrow, from Human Decision to , should be read as a potential causal relationship, not as a guarantee that the observed outcome is a clean, exogenous criterion. When the action taken modifies the very outcome used to evaluate it — stopping a machine may be precisely what prevents the failure that would otherwise serve as the criterion — stops being a simple achievement criterion.
Formally, the chain's achievement is . The classic LME applies to each stage in isolation; an extension of the equation to two chained stages, in which the output of the first becomes a cue for the second, has not, to date, been derived in the literature. The structural intuition is that introducing can alter the criterion's predictability at the second stage, — which remains a property of the entire set of cues available to the human, not of alone — but the effect of that introduction depends both on the predictive quality of the AI's output at the first stage (involving not just and , but also and, where relevant, ) and on the incremental information adds relative to the direct cues already accessible to the supervisor — information that may be small if is highly redundant with those cues. This relationship remains, for now, qualitative.
The question this architecture allows us to ask is not "who judges better, the human or the AI?" — a question the five gathered critiques show to be misframed — but rather: "at which link in this particular chain, for this particular system, did correspondence with the real environment begin to fail?"
It remains deliberately unresolved in this section whether and the conditional weights are even estimable in practice without a sufficiently long and sufficiently varied history of the supervisor's decisions — a data problem, not a theoretical one, particularly acute in high-stakes domains where critical events are rare.
