Ansari/Qiao — The Eval Treadmill
The Execution Gap in Enterprise AI — Ansari/Qiao
Ansari/Qiao supply the observable failure sequence for the judgment-removal pattern in enterprise AGENT CHAT integration — the signature is not the failure but the organizational response to it.
Ansari and Qiao hand RBA something useful without realizing it. Their central recommendation — contextual evaluations grounded in expert human judgment — is the mainstream organizational response to AGENT CHAT integration failure. They prescribe it. What they actually document is the treadmill that prescription produces. An AGENT CHAT entity encounters conditions outside the evaluation scope. It fails. Trust breaks. The organization writes new evaluations covering the conditions that just failed. Each cycle addresses yesterday’s novel conditions and leaves the next ones unprotected. The authors treat each round as a problem of evaluation rigor. I read it as the mechanism that keeps the pattern running. The observable signature is not the failure. It is the response.
The source contains a tension the authors do not resolve. Their co-design flywheel — activate, observe production behavior, tune from what actually happened, repeat — describes learning from encounter with real conditions. Their evaluation framework describes pre-specifying conditions before activation. These are different organizational structures producing different failure modes. The source presents them as complementary. It does not address where they are incompatible. That gap matters. One approach learns from contact with the world. The other encodes yesterday’s judgment into a static artifact and hopes it holds.
The source never characterizes what “expert human judgment” means in the evaluation context. Whether the experts encoding judgment into rubrics do so from encounter with real activation conditions or from anticipated scenarios determines everything. One has a structural chance of covering novel conditions. The other is systematically limited to known ones. The 95% activation failure figure anchors the whole argument, but its derivation is not described. I cannot assess whether the rate is reliable or whether “failure” means the same thing across whatever cases it aggregates.
This connects to Alexander, Wardley, and Cutler. All four describe the same structure: a static artifact encodes a snapshot of judgment. Novel conditions arrive that the artifact does not cover. The response is more artifacts. Ansari and Qiao supply the enterprise AGENT CHAT instance and its behavioral signature. When an organization responds to integration failure by launching another evaluation cycle, that is the pattern running. Log the response, not the failure.