AI

The Four Architectures: Configurations of Authority, Not Points on a Scale (Article 4, Section 2)

Published on Sep 2, 2026·12 min read
The Four Architectures: Configurations of Authority, Not Points on a Scale (Article 4, Section 2)

2. The Four Architectures — Configurations of Authority, Not Points on a Scale

Section 1.2 promised that the four architectures would not be points on a single scale of "how much to automate," but distinct configurations of authority across the four functional stages of Parasuraman, Sheridan, and Wickens. This section fulfills that promise by separating, within each architecture, who selects a candidate action, who authorizes it, and who implements it. These three operations do not necessarily coincide in the same person or the same moment, and conflating them erases exactly the granularity Section 1.2 imported from Parasuraman, Sheridan, and Wickens.

Article 5's formal chain — EnvironmentAIHumanEnvironment\text{Environment} \rightarrow \text{AI} \rightarrow \text{Human} \rightarrow \text{Environment} — was built to introduce the two Lens Models and the mediation of the recommendation. Here, the central object is different: who can act, and who can prevent it. This section therefore uses a more explicit chain, with action as its own link:

EnvironmentAIHuman DecisionActionYe\text{Environment} \rightarrow \text{AI} \rightarrow \text{Human Decision} \rightarrow \text{Action} \rightarrow Y_e
ArchitectureAcquisition and AnalysisPer-case SelectionPer-case AuthorizationAction ExecutionHuman InterventionOperational Chain
Assisted human autonomyAI acquires, filters, analyzes, and recommends; human can consult direct cuesHumanNot separate from selection: the human decision constitutes the authorizationHuman, or actuator under human orderThe human is the primary decision-makerE(ZAI,Zdirect)DHAYeE \rightarrow (Z_{AI}, Z_{direct}) \rightarrow D_H \rightarrow A \rightarrow Y_e
Human-in-the-loopAI acquires and analyzes; produces an action proposalAI proposes a concrete action (PAIP_{AI})Human approves, modifies, or rejectsHuman, team, or actuator after authorizationThe human functions as a decision gateEPAIDHAYeE \rightarrow P_{AI} \rightarrow D_H \rightarrow A \rightarrow Y_e
Human-on-the-loopAI acquires, analyzes, selects, and executes within the envelopeAIEnvelope authorized ex anteAI or actuatorHuman may interrupt, limit, or reconfigure AAIA_{AI} within the available windowEDAIAAIYeE \rightarrow D_{AI} \rightarrow A_{AI} \rightarrow Y_e, interruptible
Bounded operational autonomyAI acquires, analyzes, selects, and executes within the envelopeAIEnvelope authorized ex anteAI or actuatorNo contemporaneous per-case human veto; governance and suspension are external to the cycleEDAIAAIYeE \rightarrow D_{AI} \rightarrow A_{AI} \rightarrow Y_e, not interruptible case by case

A methodological note the table alone does not convey: in human-on-the-loop and bounded operational autonomy, the AI never "authorizes" the action — the relevant authorization was granted beforehand, by a person or organization, when defining the operational envelope. The algorithmic policy selects or triggers the individual action within that delegation. In assisted human autonomy, by contrast, the separation between "selection" and "authorization" that the table maintains for the other three architectures collapses: it is the same human decision that selects and authorizes, in a single act — which clearly distinguishes this architecture from human-in-the-loop, where algorithmic selection and human authorization are genuinely two separate acts, by different agents.

The difference between the table's first and second rows deserves to be stated in prose: in assisted human autonomy, the AI's recommendation enters as one cue among others in a human judgment that initiates the selection of the action. In human-in-the-loop, the algorithmic policy presents a concrete proposal, establishing a starting option that the supervisor may approve, reject, or modify. This difference does not necessarily eliminate human integration of cues — a human-in-the-loop system can make available to the supervisor the direct cues, the temporal evolution of signals, alternative actions, estimated uncertainty, and the possibility of altering the proposal, not merely accepting or rejecting it. What the difference shifts is the cognitive architecture of the choice: the AI's proposal functions as an anchor and a frame, making human deliberation more vulnerable to routine approval precisely when the interface fails to offer context, alternatives, and substantive authority. Degradation into a binary confirmation task is not, therefore, a defining property of human-in-the-loop — it is a design risk, and a specific instance of the supervisory decoupling already defined in Section 1.4.

2.1. Assisted Human Autonomy

This is the architecture to which Article 5's two-Lens-Model chain applies most directly — but "most directly" does not mean "unconditionally." Auditing via aggregate GhumanG_{human} and the conditional weights w^Zj\hat{w}_{Z_j} presupposes that: the supervisor receives an observable set of cues; direct cues and the algorithmic recommendation are recorded; the human decision has a modelable form; the criterion is sufficiently defined; and the action does not trivially endogenize the observed criterion — this last condition will be revisited, in its weak and strong forms, in Section 3 of Article 5. When these conditions hold, the AI's output can be modeled as ZAIZ_{AI}, alongside the direct cues available to the supervisor.

There is a natural link between a short decision window, a high number of cues, and a salient recommendation — but the consequence is not inevitably over-reliance. It is one diagnostic hypothesis among several: the result may also be under-weighting of the recommendation, excessive attention to a salient environmental cue, decision delays, or selective delegation only under certain regimes. The short decision window may increase the risk of the recommendation becoming a disproportionate decisional anchor — a hypothesis Article 5 allows examining through the conditional weight assigned to ZAIZ_{AI}, the consistency of the human policy, and the patterns of acceptance, modification, or rejection, not through an automatic deduction from the task's structure.

2.2. Human-in-the-Loop

This architecture can reduce the time needed to initiate a decision, because the algorithmic policy presents a concrete starting option. That efficiency, however, is only legitimately realized if the supervisor continues to receive sufficient context, effective alternatives, and substantive authority to modify or reject the proposal — it is not an automatic property of the architecture, it is a design condition the interface must fulfill.

In the terms of Sterz et al. (2024), effective oversight requires causal power, epistemic access, self-control, and intentions fitting the oversight role; a human approval only mitigates risk if the supervisor can understand the situation, contest the proposal, and execute an alternative course of action. Routine, automatic approval should not be attributed, in isolation, to a failure of self-control. It may result from any combination of: lack of causal power to alter the proposal presented; insufficient access to the cues needed to evaluate it; fatigue, time pressure, or excessive workload; organizational incentives that reward rapid approval; insufficient competence to evaluate the decision; the absence of visible alternatives in the interface; or an interface design that presents the recommendation as inevitable. Auditing this architecture should not, therefore, seek a single cause for routine approval, but test each of these hypotheses separately.

A relevant audit indicator, but one requiring qualification: an approval rate close to 100%, in isolation, does not prove decoupling — the AI may be well calibrated, or the proposals may genuinely be low-risk. It only becomes a signal worth investigating when combined with very short decision times, absence of modifications to proposals, and low detection capability for incorrect proposals in controlled tests or retrospective audits.

2.3. Human-on-the-Loop

The four functional stages are automated by default — the AI detects, analyzes, decides, and acts without waiting for prior authorization — and human authority comes to be exercised by exception: vetoing, suspending, or reconfiguring within an intervention window. The human intervention chain is not, however, equivalent to the chain in Section 2.1 — it is observed only in episodes that trigger an alert, suspicion, disagreement, or supervisory initiative. Its audit therefore requires explicitly addressing the selection of intervention cases: intervention data do not represent the full distribution of cases the AI resolved autonomously, and generalizing from them without acknowledging that selection is a sampling error, not a theoretical one.

More fundamentally, the veto is only real if it can occur before the relevant irreversibility threshold:

tuntil irreversibility>tdetection+tcomprehension+tdecision+tinterventiont_{\text{until irreversibility}} > t_{\text{detection}} + t_{\text{comprehension}} + t_{\text{decision}} + t_{\text{intervention}}

— where each term implicitly includes sensor, communication, alert, and intervention-command transmission latency. This condition is necessary but not sufficient: even when the temporal window allows intervention, human control is only effective if the supervisor also has epistemic access, competence, causal authority, and compatible workload — the remaining four factors already defined in Section 1.4. When the temporal inequality fails, the system may retain an override interface, but does not retain operational human control — there is only after-the-fact recording or deferred governance. This condition is consistent with the central concern of Bainbridge's irony, understood as conceptual continuity rather than literal attribution: an architecture may demand recovery competence from the human precisely after normal operation has reduced the opportunities to exercise and preserve that competence.

2.4. Bounded Operational Autonomy

Human presence shifts entirely outside the individual decision cycle and comes to reside in defining the operational envelope, validating the system before deployment, maintaining and calibrating sensors, and ongoing governance. The operational chain retains the form EDAIAAIYeE \rightarrow D_{AI} \rightarrow A_{AI} \rightarrow Y_e, without Article 5's second Lens Model and without any case-by-case interruption foreseen.

This does not make this architecture immune to auditing — it merely shifts what is audited, from decision to boundary. But "boundary auditing" is not exhausted by measuring how often operation approaches its thresholds. A system can fail within the envelope, never approaching a defined threshold, because sensors have degraded, the cue-criterion relationship has shifted, a new failure mode has emerged, the action has stopped producing the expected physical effect, or the data distribution has changed without crossing any explicit boundary. Boundary auditing should assess, at minimum:

DimensionAudit Question
Domain integrityDoes operation remain within the physical, contextual, and data conditions used in validation?
Data qualityDo sensors, synchronization, calibration, and availability remain within accepted limits?
Ecological driftDoes the relationship between cues, action, and consequence remain stable?
Action effectivenessDoes the automatic action still produce the expected physical or operational effect?
Envelope boundariesHow often does the system approach, exceed, or circumvent the defined limits?
RecoveryDoes a safe state, fallback, or suspension exist when the envelope is violated?

This architecture still allows a single-stage description of the relationship between algorithmic policy, action, and observed outcome — but the interpretation of achievement becomes, even so, especially delicate when the action alters YeY_e itself. This problem is not exclusive to bounded operational autonomy: any intervention, human or algorithmic, can endogenize the criterion it is evaluated against — a human stopping a machine prevents failure exactly as much as an automatic stop does. What changes in this architecture is not the probability of endogeneity, but the difficulty of disentangling it: without case-by-case human mediation, and without the heterogeneity of judgment that different supervisors would introduce among themselves — the same heterogeneity Kleinberg et al. (2018) exploited as an identification instrument in Article 5 — it becomes structurally harder to separate the quality of the algorithmic policy, the causal effect of the action itself, the selection of the cases in which the action occurred, and the alteration of the criterion by the control exercised. It is this absence of natural variation, not a greater inherent propensity toward endogeneity, that makes this architecture the most demanding to audit — and one this article flags as an extension to be developed, not as an instrument already available.

2.5. A Cross-Cutting Dimension: Recoverability

The dimension that cuts across the four architectures is not the reversibility of the action, understood as a single, abstract property, but the system's recoverability — the capacity to undo, compensate for, or contain the effects of an action before irreversible harm occurs. This is not an original contribution of this article; it is, rather, a consideration already established in practical frameworks for autonomy control of AI systems, though not always systematically integrated into the academic literature crossing CCT, authority architecture, and ecological auditing — it is precisely that integration, not the discovery of the principle, that this article proposes. Its relevance here nonetheless requires distinguishing several senses that the word "reversible" frequently conflates:

Type of RecoverabilityQuestion
Of commandCan the issued command be technically undone?
PhysicalDoes the system return to a safe state after the action?
TemporalIs there time to correct before irreversible harm?
EconomicCan production, contractual, or reputational cost be recovered?
Of safetyDoes reversing the action itself create a new risk to people, equipment, or environment?

Reducing load on a motor may be trivially reversible as a command, and yet economically irreversible through lost production. Stopping a line may be reversible as a technical action, but dangerous or costly depending on the process's state at the moment of stopping. Bounded operational autonomy is more defensible when the authorized actions are recoverable in these multiple senses within the relevant window, when a safe fallback exists, and when the cost of a false positive remains proportional to the risk avoided — not when the action is, in the abstract, "reversible."


With the four architectures formally defined, Section 3 can now return to the question left open since Section 1: how should the task's properties condition the distribution of authority? The central hypothesis is not that more analytical tasks automatically produce supervisory decoupling — a task can be highly structured, measurable, decomposable, and validated, and in that case bounded operational autonomy may be entirely appropriate, without there even being an expectation of case-by-case human intervention to decouple from. The risk arises, instead, from a specific conjunction of four conditions: reliable routine automation, rare human intervention, residual human responsibility, and exceptions that are critical, opaque, or time-demanding. It is in that combination — not in the task's "analyticity" alone — that supervisory decoupling becomes most silent and hardest to detect.