3. The Criterion Problem — Weak Form and Strong Form
All the preceding sections presupposed, without saying so, that — the distal criterion against which the entire chain is evaluated — exists as a well-defined, observable quantity, independent of the decision using it as a standard. The five adversarial reviews gathered for this article converge, with notable unanimity, in pointing out that this presupposition is the most fragile part of the whole edifice.
The criterion problem is not merely about whether exists and is observed. It is about knowing what represents, when it is measured, for whom, under which action it is observed, and whether it remains the same once the system starts acting on the environment. Before distinguishing unobservability from endogeneity — the two forms that give this section its name — it is worth acknowledging an even more elementary question: a can be perfectly recorded and still be an imperfect proxy for the construct one intends to audit.
| Stated Objective | Possible Observed Criterion | Problem |
|---|---|---|
| Industrial safety | Number of emergency stops | Many stops may indicate high risk, but also effective safeguards |
| Asset reliability | Failure recorded within 30 days | Depends on usage intensity, maintenance, and observation period |
| Correct diagnosis | Agreement with a clinical panel | Measures agreement with specialists, not clinical truth |
| Fairness in bail decisions | Recorded recidivism | Measures detected crime, policing, and prosecution — not fairness |
| Good algorithmic advice | Acceptance by the supervisor | Measures human adherence, not the recommendation's ecological validity |
At Argus, "no failure after the alert" can mean that the alert identified a real risk and the action prevented the failure; that there was no real risk and a false positive occurred; that there was risk, but the machine was not stressed enough for the failure to manifest; that a partial failure went unrecorded; or that the equipment was replaced before the observation horizon. The same operational label can correspond to entirely different mechanisms. The chain formalized in Sections 1 and 2 does not resolve this construct validity — it requires that it be specified before the audit, not after.
A second question, distinct but equally prior to the two forms, is that of the temporal horizon. This article deals explicitly with closed cyber-physical loops, and therefore there is not just one criterion — there is , where is the horizon at which the outcome is measured. The same decision can improve an immediate outcome and worsen a later one: stopping a machine prevents a failure today, but may introduce startup stresses that increase future risk; deferring maintenance improves availability in the current shift but increases the probability of failure within weeks. The criterion therefore needs to declare an explicit prediction horizon ("failure within the next 24 hours," "before the next inspection"), a time zero (the alert, the decision, the execution of the action, or the return to operation), and an aggregation rule (first failure, any failure, maximum severity). Without a previously declared temporal window, the audit risks retrospectively choosing the horizon that favors the observed policy — and this risk connects directly to the temporal aliasing already formalized in Article 2, and to the presentation and duration properties already classified in Article 3: the same erosion of correspondence between the supervision's cadence and the system's cadence can manifest here as an erosion of correspondence between the window in which the criterion is defined and the window in which the decision actually produces its effect.
With the semantic condition and the temporal horizon declared, the distinction between the two forms of the criterion problem becomes operational. In its weak form, the problem is one of unobservability: exists, in principle, as a fact of the world, but is not directly accessible for all the cases the audit would need to compare. In medical diagnosis, the gold standard is frequently the consensus of an expert panel — using one human judgment to validate another. In judicial sentencing, the usual proxy, recidivism, is contaminated by policing practices and by the bail system itself, which determines who is observed at liberty. In industrial IoT, failure is frequently censored by the very maintenance intervention: if the system prevents the breakdown, the counterfactual of failure is never observed.
In the language of potential outcomes: let be the outcome that would occur under an action . For each case, one only observes — the outcome under the action actually taken — never the counterfactual .
There is also a distinct risk that should not be confused with good or poor observability of the criterion: leakage, or tautological overlap between cues and criterion. If a cue includes information that is contemporaneous with, subsequent to, or nearly identical to the very outcome one intends to predict, can be artificially inflated. The problem, in that case, is not that there is little ecological uncertainty; it is that the audit stops testing genuine prediction or correspondence and starts measuring a circular relationship between variable and criterion — particularly relevant in this article's AI-IoT chains, where derived cues and the AI's own output can, without care, contaminate one another over time.
This is precisely the selective labels problem, named and formalized by Lakkaraju, Kleinberg, Leskovec, Ludwig, and Mullainathan (2017): we only observe the outcome for the cases in which the decision allowed that outcome to manifest. The authors do not treat the problem as simple missing data to be corrected by reweighting — they propose a specific technique, contraction, which compares human decisions and predictive models under selective labels by exploiting heterogeneity among decision-makers, without relying on traditional counterfactual imputation.
The weak form does, however, admit a partial and technically honest response. Kleinberg, Lakkaraju, Leskovec, Ludwig, and Mullainathan (2018), in the exact context of bail decisions, exploit the approximately random assignment of cases to judges and the systematic variation in their leniency to estimate effects and compare policies under explicit causal assumptions. The design does not make observable for all cases; it creates exogenous variation in the decision that allows learning about counterfactual consequences in comparable groups. Ludwig and Mullainathan (2021) use these observed fragilities in the justice system to draw more general lessons: variation among decision-makers, outcome selectivity, and prediction fragility should be treated as problems of institutional design, not merely as statistical details of one specific domain. In this article's context, that variation might come from differences between operators, shifts, or Argus installations — but the mere existence of variation does not, by itself, constitute a valid instrument. More experienced operators may receive more critical assets; certain shifts may coincide with more demanding production regimes; different installations may have distinct equipment and sensors; a "more lenient" operator in the stop decision may also differ in diagnostic quality; shifts and operators may communicate with one another, creating interference between supposedly independent units.
It is essential not to read the leniency instrument as a complete resolution. Its assumptions — plausibility of exogenous assignment, strength of the variation, balance across cases, absence of alternative channels, stability of the causal mechanism — must be explicitly stated, subjected to empirical diagnostics, and substantively defended; some are partially testable, others rarely so completely. When these assumptions do not hold, the weak form of the criterion problem remains, in fact, unresolved in that concrete case.
The strong form is structurally different: it is not that is hard to observe, it is that the decision itself alters the environment that generates the criterion. Incarceration simultaneously prevents and causes recidivism. Stopping a machine may be precisely what prevents the failure that would serve as the criterion for evaluating whether the stop decision was correct. The two forms are not mutually exclusive: in a preventive-maintenance system, the stop can simultaneously censor the observation of the counterfactual failure (weak form) and causally alter the process that would produce the outcome (strong form). The weak form asks whether the relevant criterion is observed for the necessary cases; the strong form asks whether there is even a single criterion, independent of the action, that can serve as a standard for evaluating the decision.
This article does not resolve the strong form, but it also does not declare it impossible in principle. Causal inference exists precisely to study the effects of actions on outcomes; a randomized trial, a discontinuity design, or a valid instrument can, in specific domains, estimate even when the decision causes the very outcome. What the strong form renders insufficient is not causal inference in general — it is the simple interpretation of as achievement in the classic LME sense, when is itself a partial product of the decision being evaluated. In such cases, the audit must explicitly specify potential outcomes, temporal horizon, loss function, and a causal theory of the intervention.
This echoes, and now formalizes, the prior condition Article 4's design heuristic already required before any architecture choice could make sense: without a sufficiently defined, auditable, and not trivially endogenous criterion, neither Article 4's architecture choice nor Article 5's audit has ground to stand on.
Additional Complications, Not Resolved Here
Five adjacent complications deserve mention here, even though their full development belongs elsewhere in this trilogy.
Competing risks and censoring not caused by the decision. The weak form, as presented, treats censoring as a consequence of the audited decision — but an asset can stop being observable for independent reasons: a sensor fails, the contract ends, the equipment is replaced, production is reduced and the machine stops operating under comparable load. The practical question for Argus is whether the absence of failure means a healthy machine, a stopped machine, a replaced machine, or simply an absence of record — states that should not all be indiscriminately coded as , at the risk of survivorship bias artificially inflating achievement.
Differential criterion by subgroup. Miscalibration hidden by an aggregate or has already been identified as an objection in the reviews underlying this article — but the same concern applies to itself: "failure" may be measured with different frequency, precision, and thresholds between instrumented and non-instrumented installations, between new and old assets. A report that stratifies or but treats as homogeneous may confuse a difference in performance with a difference in observation.
Criterion drift after deployment. Once the system is introduced, the criterion can change because the system itself changes the environment that generates it — operators adapt their behavior, alerts change which failures get recorded, maintenance policies are redesigned. This is a broader problem than the endogeneity of a single decision; it is a distribution-and-measurement feedback loop, and connects directly to the stabilization protocol still to be developed in this trilogy — the same version-tracking that will allow distinguishing genuine environmental improvement from a mere change in the observation mechanism.
Multiple counterfactuals and loss function. The LME measures correlation between policy and criterion; it does not choose a loss function. When the goal is to evaluate actions, not merely predictions, the criterion stops being a single outcome and becomes a comparison between potential outcomes under alternative actions — false positive (stopping without real risk) versus false negative (not stopping under risk) — weighted by a function this article does not presume to be given by the LME. This is exactly the cost asymmetry already flagged as a possible extension via Beckstead's Bifocal LME.
Vector-valued, not scalar, criterion. In many high-risk domains, is a vector — safety, availability, cost, quality — not a single quantity. An action that improves safety may worsen availability. The decision to reduce this to a scalar criterion requires an aggregation rule and normative weights; it is not a neutral or purely statistical operation, and connects directly to the recoverability and severity of actions already discussed in Article 4.
The common principle across these five complications, and the most directly communicable outside this article's technical body, is this: the absence of observed failure is not the same as the absence of risk, and the absence of risk is not the same as the success of the decision. This tool's condition of applicability remains the one already stated — a sufficiently defined criterion, tractable selectivity, negligible endogeneity or one complemented by causal design — but that condition must now be read in light of all the forms, not just two, in which can fail to represent what one intends to audit.
