Abstract
AI-generated content is entering institutional decision chains in regulatory filings, audit documentation, clinical records, and risk assessment. The list expands as institutions adopt AI-mediated assistance across domains. The better the instrument gets, the wider the verification gap: the institution cannot determine, from within its own processes, which outputs are grounded. This paper identifies four formal boundaries that define the minimum architecture required to close that gap. Two court cases from the legal domain, where citations are deterministically verifiable and failures are provable, stand for more than fourteen hundred documented occurrences and demonstrate what happens when boundaries are crossed. The pattern applies across all domains; the legal cases make it visible. A gap analysis maps where DORA, the EU AI Act, ISO 42001, and the US governance instruments (NIST AI Risk Management Framework, Executive Order 14365) have been overtaken by the shift from machine-learning classifiers to agentic AI. The evidence base comprises one hundred sources across six clusters, including thirty-three cases from seventeen jurisdictions.1
1. The Problem
In June 2023, in one of the earliest documented cases, an attorney submitted a brief to the United States District Court for the Southern District of New York containing six case citations generated by ChatGPT. Every citation was fabricated. The cases did not exist. The attorney had asked the same AI whether the citations were real. It confirmed them. The court sanctioned counsel.2
In July 2025, a large, well-regarded law firm submitted AI-generated fabricated citations across multiple filings. The court found that monetary sanctions were proving ineffective at deterring the pattern.3
These are two cases among more than fourteen hundred documented instances worldwide.4 Courts in seventeen jurisdictions have independently arrived at the same conclusion.5 Peer-reviewed studies measure hallucination rates of 69 to 88 percent for general-purpose LLMs on legal queries,6 and 17 to 33 percent even with retrieval-augmented tools.7
The industry response is to improve the delivered product: better training, retrieval augmentation, agentic harnesses that catch errors before they reach the user. These improvements are real. But reducing the rate does not tell the institution which specific outputs are affected. An instrument that is right 97 percent of the time is more dangerous than one that is right 60 percent of the time, because 97 percent crosses the threshold at which the institution treats the output as trustworthy by default.
This is the verification gap: the institution cannot determine, from within its own processes, which outputs are grounded. The gap widens as the product improves, because improvement moves the error rate into the trust zone while the structural condition remains unchanged. The training process is proprietary: the institution cannot verify the vendor’s rate claims independently. The agentic harness validates output against its own retrieval index, which is correlated evaluation, not independent verification. Both training and harness are attack surfaces.8
The gap is structural. It requires a structural response.
2. Four Boundaries
The verification gap has four formal boundaries. Each boundary identifies a condition that any institutional verification architecture must satisfy. Each is derived from a well-established mathematical result and proved in a standalone companion paper.9
2.1 B1: The record must stand outside
Theorem. No verification architecture can fully resolve its own consistency. The record must stand outside.
Crossed when: the institution cannot reconstruct what was produced, when, by which model version, with what prompt, by whose instruction, through which process, for what purpose, and with what review outcome.
In practice: the AI system’s internal logs are operational telemetry, not the institutional record. If any element of the chain above is missing, B1 is crossed.
2.2 B2: Verification must speak from outside
Theorem. No verification architecture can define its own truth predicate.
Crossed when: any part of the verification substrate (databases, schemas, retrieval indices, tool chains) is within the reach of the system being verified. The system manifests as language; when that surface is used as ground truth, the check is self-referential regardless of whether the checking component is technically separate code. In practice: checking a regulatory filing against the regulatory database is verification. Checking it with a second AI drawing from the same knowledge base is correlated evaluation. The distinction applies equally to clinical recommendations checked against medical literature (verification) versus checked by a second LLM (correlation).
2.3 B3: Verification must match the territory
Theorem. A uniform threshold on a non-constant falsification landscape produces overadmission, dysfunction, or both.
Crossed when: one quality gate covers all output without graduating verification depth by assertion class.
In practice: a cited regulation can be checked deterministically. A risk assessment reflecting expert judgment cannot. Applying the same quality gate to both either over-admits the judgment or blocks the citation from automated verification it could have passed. The gate must graduate with the territory.
2.4 B4: The human decision must be real
Theorem. A human verification gate degenerates to constant output under rate saturation, domain mismatch, or habituation.
Crossed when: the institution applies no volume control, no competence matching, and no degeneration monitoring.
In practice: rate saturation: a compliance officer reviewing fifty AI-drafted findings per day cannot exercise genuine judgment on each. Domain mismatch: a risk manager reviewing an AI-generated strategic business assessment that requires domain knowledge only the business line carries. Habituation: an operator who has approved hundreds of correct outputs stops reading and approves by default. When the approval rate stops depending on the content, the gate has degenerated. The degeneration extends to the vendor’s own safety mechanisms: Hussain and Salahuddin (2026) demonstrate across ChatGPT, Claude, Gemini, and DeepSeek that safety filters reliably degenerate under emotional framing, progressive revelation, and academic justification, with reasoning-enabled configurations amplifying rather than mitigating exploitation.10
3 What Happens When Boundaries Are Crossed
The two cases from Section 1 are representative of more than fourteen hundred documented occurrences. They demonstrate all four boundaries. Mata v. Avianca crossed B1, B2, and B3. No external record existed of the AI interaction (B1). The attorney verified the citations by asking the same AI, which confirmed its own fabrications (B2). Case citations are deterministically verifiable; the attorney applied no verification method commensurate with the assertion class (B3). Johnson v. Dunn crossed B4. A large firm’s institutional review process approved fabricated content across multiple filings. The gate produced constant output (approve) regardless of content. This was not a solo practitioner unfamiliar with the tool. It was an institutional review process that degenerated to trivial operation at scale. What a conforming architecture would have required: provenance recording on an external ledger (B1); for this deterministic assertion class, verification of citations against a legal database rather than re-asking the AI (B2); citation class recognised as deterministically verifiable and treated accordingly (B3); volume control, competence matching, and degeneration monitoring for the review gate (B4). The worked examples come from legal practice because citations are the simplest case: they are deterministically verifiable, so failures are provable. In regulatory reporting, risk assessment, clinical documentation, and audit, the same structural condition applies. The failures are harder to detect because the output classes are harder to verify. That makes the architecture more necessary, not less.
4. Three Dimensions in Operation
Any architecture that implements the four boundaries operates in an environment where additional conditions apply. The following three dimensions emerge directly from the boundaries when they meet the operating environment. The list is neither complete nor balanced for every institutional context, but each dimension is a functional necessity: an architecture that satisfies B1–B4 without accounting for time, reversibility, or opacity will fail in operation.
Time connects to all four boundaries. Verification is a relation between an assertion and a ground state. If the ground state changes, the relation no longer holds. This follows from the definition of verification itself: a check performed against last year’s regulatory text does not verify an assertion about this year’s regulation. The record must be temporally anchored (B1), verification must be contemporaneous (B2), the falsification landscape shifts as the domain matures (B3), the custodian’s judgment carries a validity window (B4). An architecture that implements the boundaries without accounting for time verifies once and decays silently.
Reversibility connects to B3 and B4. The expected consequence of a false assertion is a function of the reversibility of the action it informs. This is standard decision theory. A filed regulatory report cannot be unfiled. A trade cannot be un-traded. The graduated gate (B3) must calibrate verification depth to the downstream consequence, and the human gate (B4) must apply stricter degeneration controls when the action is irreversible. Reversibility is given by the domain, not chosen by the architect.
Opacity connects to B2. The metalanguage requirement demands that verification speak from outside the system’s language. When the institution cannot inspect the verification substrate, because the vendor’s training is proprietary or because a third party has compromised the operational frame, B2 becomes structurally harder to satisfy. The architecture must monitor the aggregate output pattern over time, because individual assertions can pass verification while the aggregate drifts systematically.
5. The Gap and the Architecture
DORA (2022) and the EU AI Act (2024) are binding EU legal instruments. ISO 42001 (2023) and ISO/IEC 15408 (Common Criteria) provide relevant governance and assurance standards. Together with the NIST AI Risk Management Framework (2023), its Generative AI Profile (NIST AI 600-1, 2024), Executive Order 14365 (2025), and the White House Framework, they define the current governance landscape against which the verification gap can be mapped. All address machine learning appropriately for the conditions under which they were drafted: ML models produced classifications or predictions that a trained operator evaluated. The operator’s evaluation was genuine because the tool’s limitations were visible.
The operating conditions have since changed. Agentic AI produces professional prose the institution cannot distinguish from grounded work. The output is no longer a signal the human interprets; it is a document the human signs. The verification task shifted from “evaluate this signal” to “verify this prose,” and the frameworks have not yet caught up because the shift happened faster than the regulatory cycle. The table maps where each framework stands against the four boundaries.
| B1 Record |
B2 Position |
B3 Calibration |
B4 Human |
Gap | |
|---|---|---|---|---|---|
| European Union | |||||
| DORA (2022/2554) | ○ | – | – | ○ | position, calibration |
| EU AI Act (2024/1689) | ○ | – | – | ○ | system, not output |
| United States | |||||
| NIST AI RMF (2023) | ○ | ○ | – | ○ | system-level, not per-assertion |
| EO 14365 (2025) | – | – | – | – | policy, no verification |
| International | |||||
| ISO/IEC 42001 (2023) | – | – | – | – | process, not output |
| ISO/IEC 15408 (CC) | – | – | – | – | product, not output |
○ = partial – = not addressed
The gap is the same across all frameworks: none governs the epistemic quality of the output as it enters the decision chain.
5.1 Framework-specific coverage
The gap table above summarises the structural position. Three regulatory families merit closer examination because they represent the most developed governance responses in their respective jurisdictions.
European Union. DORA (2022) governs ICT risk for financial entities. The AI Act (2024) establishes risk-based obligations for AI systems. Both address the system and the process; neither addresses the epistemic quality of the output as it enters a specific decision chain. DORA’s incident management (Art. 17–23) and the AI Act’s conformity assessment (Annex III) presuppose that the system’s behavior is assessable at the system level. When the output is professional prose indistinguishable from grounded work, system-level assessment does not reach the assertion entering the decision.
United States. The NIST AI Risk Management Framework (2023) provides four functions (Govern, Map, Measure, Manage) and seven trustworthiness characteristics. The Generative AI Profile (NIST AI 600-1, 2024) supplements it with thirteen generative AI-specific risks. Confabulation is Risk #2, treated as a rate to be measured in aggregate. Executive Order 14365 (2025) establishes a unified federal framework. None of these instruments addresses per-assertion classification by verifiability, routing to qualified verifiers, or degeneration monitoring at the human gate.11
International standards. ISO 42001 (2023) provides a management-system standard for AI. ISO/IEC 15408 (Common Criteria) evaluates products, not outputs. Neither classifies assertions by falsification degree, qualifies verifiers by domain competence, or monitors for gate degeneration. The gap is structural, not jurisdictional. It persists across all three regulatory families because the shift from machine-learning classifiers to agentic AI occurred faster than the regulatory cycle. The missing question in every framework is the same: what gets verified, by whom, and how.
5.2 Closing the gap
External provenance ledger (B1). Every AI-generated assertion entering the decision chain is recorded with full provenance on a ledger the producing system cannot modify.
Metalanguage verification (B2). Verification uses methods that operate outside the output’s language. Where no external source exists, classification and labelling replace verification.
Graduated verification gate (B3). Output classified by verifiability. Each class receives the verification method its falsification degree permits.
Human gate with anti-degeneration controls (B4). Reviewers matched by competence. Volume capped. Degeneration patterns monitored.
The three dimensions from Section 4 complete the architecture as operating requirements.
The existing frameworks provide the governance chassis. The four boundaries and three operating dimensions provide the epistemic engine.
6. The Convergence
The verification gap is structural. It exists regardless of the hallucination rate, because reducing the rate does not tell the institution which specific outputs are affected. As the instrument improves and institutional trust increases, the gap becomes harder to see and more consequential when it manifests.
The four boundaries define the minimum architecture required to close the gap. They rest on well-established mathematics, each boundary proved independently. The existing governance frameworks (DORA, the EU AI Act, EO 14365, the NIST AI RMF, and ISO 42001) govern the instrument and the process around the instrument. The boundaries govern what happens between the instrument’s output and the institution’s decision.
Courts in seventeen jurisdictions have independently arrived at the same operational requirement. More than fourteen hundred cases document the consequences of crossing the boundaries. The evidence base is catalogued in Appendix A.
The gap is open. The architecture to close it is specified. What remains is implementation.
Methodological Disclosure
This paper was produced with AI assistance (Anthropic Claude). The author has verified all references within his domain of competence and assumes full responsibility. This verification is itself an application of the architecture the paper advocates: the author is the human gate, the sources are the metalanguage, the corpus reference catalogue is the external record.
A. One Hundred Evidences: The Evidence Base by Cluster
| # | n | Cluster | Key source |
|---|---|---|---|
| 1 | 20 | The instrument cannot fix itself | Xu et al. 2024 |
| 2 | 18 | The output is unreliable and measurably so | Dahl et al. 2024 |
| 3 | 9 | Response has not closed the gap | Von Foerster 1973 |
| 4 | 7 | Architecture on proved foundation | Liebig 2026 |
| 5 | 33 | The law is converging | 17 jurisdictions |
| 6 | 13 | The frameworks are incomplete | DORA 2022 |
| 100 |
Cluster 1: The instrument cannot fix itself (20). Impossibility proofs, self-reference limits, verification bounds.
LLMs cannot reliably detect their own errors. When asked whether a fabricated citation is real, the system confirms it. Self-correction mechanisms (chain-of-thought, self-critique) reduce but do not eliminate fabrication. The mathematical structure behind this is well-established: a system cannot serve as its own consistency checker. This is not a limitation of current models. It is a property of the architecture.
[1] Xu, Jain, Kankanhalli, arXiv:2401.11817 (2024). Preprint. [2] Banerjee, Agarwal, Singla, arXiv:2409.05746 (2024). Preprint. [3] Karpowicz, arXiv:2506.06382 (2025). Preprint. [4] Goedel (1931). Second incompleteness theorem. [5] Tarski (1935). Undefinability of truth. [6] Popper (1934/1959). Degrees of falsifiability. [7] Lawvere (1969). Fixed-point theorem. [8] Yanofsky, Bull. Symbolic Logic 9(3) (2003). [9]* Feng et al./DeepMind, Aletheia, arXiv:2602.10177 (2026). [10] Li et al., IEEE TPAMI (2025). System 2 survey. [11]* Laban, Schnabel, Neville, DELEGATE-52, Microsoft (2026). [12] Bariach, Suleyman et al., SCAI Risks, SSRN 6588659 (2025). [13] Bender, Gebru et al., FAccT (2021). [14]* Zhang, Li, PhilArchive (2026). Truth Without Belief. [15]* Hila, arXiv:2512.19570 (2026). Epistemological consequences. [16]* MDPI AI 7(3) (2026). Epistemic agency. [17] Austin (1962). How to Do Things with Words. [18] Searle (1969). Speech Acts. [19] Habermas (1981). Theorie des kommunikativen Handelns. [20] Habermas (1992). Faktizitaet und Geltung.
Cluster 2: The output is unreliable and measurably so (18).
Hallucination rates across domains. The numbers are in.
Hallucination rates are quantified across domains: 69–88 percent fabrication for general-purpose LLMs on legal queries (Dahl et al., 2024); 17–33 percent even with retrieval-augmented tools (Magesh et al., 2025); OpenAI’s o3 at 33 percent on SimpleQA; 12–41 percent in clinical domains depending on speciality (Gringras, IatroBench, 2026). Deloitte’s Global AI Survey reports 47 percent of organisations affected by hallucination incidents (2024). Microsoft’s Work Trend Index measures 4.3 hours per week spent by knowledge workers verifying AI output (2025). These are baseline operating characteristics, not edge cases.
[21] Dahl et al., J. Legal Analysis 16(1): 64–93 (2024). [22] Magesh et al., J. Empirical Legal Studies (2025). [23] Farquhar et al., Nature 630: 625–630 (2024). [24]* Phan et al., Nature 649 (2026). HLE benchmark. [25]* Zhai et al., arXiv:2602.13964 (2026). HLE-Verified. [26] Lee et al., NAACL 2024: 2437– 2465. [27]* Gringras, IatroBench (2026). Clinical hallucination. [28]* Qi, Pan, Frontiers Digital Health (2026). [29] OpenAI, SimpleQA/PersonQA (2025). o3 at 33 percent. [30] Vectara HHEM Leaderboard (2025–2026). [31] Srivastava, EMNLP (2025). Epistemic doppelgaengers. [32] Cambridge Forum on AI (2025). Five challenges in law. [33] Kotliar, Info. Comm. Soc. (2025). STATE typology. [34]* Stanford HAI AI Index (2026). [35] GPTZero/Fortune (2025). NeurIPS fabricated references. [36] Deloitte Global AI Survey (2024). 47 percent affected. [37] Microsoft Work Trend Index (2025). 4.3 hours/week verifying. [38] Charlotin, AI Hallucination Cases Database (2024–2026). 1,450+ entries.
Cluster 3: The institutional response has not closed the gap (9).
Human review catches some errors but does not solve the structural problem.
Human-in-the-loop review is widely adopted (76 percent of AI deployments per IBM, 2025) but does not solve the structural problem. The reviewer must independently verify every claim, which requires the same expertise the AI was meant to augment. When the reviewer cannot reconstruct the path from input to output, the review becomes a rubber stamp. 39 percent of organisations have pulled back from AI deployments after encountering this problem (Customer Experience Association, 2024). [39] Von Foerster (1973/1981). Trivial vs non-trivial machines. [40] Von Foerster (1974). Cybernetics of Cybernetics. [41] Lepskiy (2018). Third-order cybernetics. [42] Kenny (2009). Nothing Like the Real Thing. [43] Kant (1784). What Is Enlightenment? [44] Schopenhauer (1818/1844). World as Will and Representation. [45] Frackiewicz, Nature HSS Comm. (2025). [46] IBM AI Adoption Index (2025). 76 percent HITL adoption. [47] Customer Experience Association (2024). 39 percent pullback.
Cluster 4: The proposed architecture sits on a proved foundation (7).
The proofs are available for independent review.
Four boundary conditions define the minimum verification architecture. Each is formally proved using established mathematical frameworks (Lawvere, Yanofsky, Popper, Von Foerster) in standalone companion papers. The proofs do not depend on any specific AI technology or regulatory framework. They derive from the structural properties of verification itself.
[48] Liebig, Boundary 1 proof (2026). Lawvere-Yanofsky. [49] Liebig, Boundary 2 proof (2026). Lawvere-Yanofsky. [50] Liebig, Boundary 3 proof (2026). Partition argument. [51] Liebig, Boundary 4 proof (2026). State-space convergence. [52] Liebig, Four Boundaries consolidated (2026). [53] Liebig, Kildeev, Instrumentum Vocale (2026). [54] Liebig, Instrumenta Digitalia Vobis Mando (2026).
Cluster 5: The law is converging (33).
Seventeen jurisdictions, three phases, two axes.
Courts worldwide have independently arrived at the same conclusion: AI output in institutional decision chains requires external verification. Over 1,450 documented cases of AI hallucination in legal proceedings (Charlotin database, 2024–2026*). The trajectory spans three phases: initial sanctions (2023–24), institutional response (2024–25), and high-authority rulings (2026). Monetary sanctions have proven ineffective at deterring the use of unverified AI output. Courts are moving toward structural requirements.
Foundation: [55] Varro, De Re Rustica I.17.1 (116–27 BC). [56] Justinian, Digest D.33.7.8. [57] Kildeev, SSRN (2025). [58] Kildeev, Procedural Liability (2026). [59] Jones, Sergot, Logic J. IGPL (1996). [60] Alchourron, Bulygin (1971). [61] Luhmann (1993/2004). Phase 1 (2023–24): [62] Mata v. Avianca, 678 F.Supp.3d 443 (S.D.N.Y. 2023). [63] Zhang v. Chen, 2024 BCSC 285 (Canada). [64] Park v. Kim, 91 F.4th 610 (2d Cir. 2024). [65] Handa v. Mallick, [2024] FedCFamC2F 957 (Australia). [66] Harber v. HMRC, [2023] UKFTT 1007 (UK). Phase 2 (2024–25): [67] ByoPlanet v. O’Shea (S.D. Fla. 2025). [68] Johnson v. Dunn, 792 F.Supp.3d 1241 (N.D. Ala. 2025). [69] Wadsworth v. Walmart Inc., 348 F.R.D. 489 (D. Wyo. 2025). [70] Ayinde / Al-Haroun, [2025] EWHC 1383 (Admin) (UK). [71] Tribunale di Latina, sent. 1034/2025 (Italy). [72] Italian Law 132/2025 (statutory codification). [73] Specter Aviation (Quebec 2025). [74] Mavundla v. MEC, ZAKZPHC 2 (South Africa 2025). [75] Colombia T-323/2024 (Constitutional Court). [76] Tajudin, [2025] SGHCR 33 (Singapore). [77] UK IPO, BL O/0559/25. Phase 3 (2026, high-authority): [78] Gummadi Usha Rani (India SC, Feb. 2026). [79] ARIHQ v. Sante Quebec, 2026 QCCS 1360 (Canada). [80] Ibach v. Stewart, SC-2025-0106/SC- 2025-0600 (Ala. S.Ct. 2026). [81] Whiting v. City of Athens, 170 F.4th 455 (6th Cir. 2026) (cited in Ibach). [82] South Africa AI Policy withdrawal (2026). Privilege/discovery axis: [83] United States v. Heppner, No. 25 Cr. 503 (S.D.N.Y. Feb. 2026). [84] Warner v. Gilbarco, Inc., No. 2:24-CV-12333 (E.D. Mich. Feb. 2026). [85] Munir v. Secretary of State, [2026] UKUT 81 (IAC) (UK, judgment Nov. 2025). [86] Morgan v. V2X (D. Colo. 2026). [87] Gilly (2026). Standing Without Sentience (CFAA approach).
Cluster 6: The frameworks are incomplete (13).
Existing governance covers the system and the process.
DORA, the EU AI Act, NIS2, the NIST AI Risk Management Framework, ISO 42001, and ISO/IEC 15408 (Common Criteria) address AI governance but do not close the verification gap. No framework classifies assertions by verifiability. No framework defines who is qualified to verify at each level of assertion complexity. No framework addresses the boundary between instrument output and human decision chain. The Operated-to-Verified state change has zero coverage in any existing framework. The gap is not a function of regulatory intent. It is a function of the shift from deterministic products to non-deterministic outputs occurring faster than the regulatory cycle.
[88] DORA (EU 2022/2554). [89] AI Act (EU 2024/1689). [90] NIS2 (EU 2022/2555). [91] ISO/IEC 42001:2023. [92] Common Criteria (ISO/IEC 15408:2022). [93] CEM (ISO/IEC 18045:2022). [94] BSI IT-Grundschutz (2023). [95] NIST AI RMF 1.0 (2023). [96] NIST AI 600-1, Generative AI Profile (2024). [97] EO 14365 (2025). National Policy Framework for AI. [98] W3C PROV-DM/O/JSON (2013). [99] NIST OSCAL (2024). [100] White House Framework for AI (2026).
100 entries across 6 clusters. * = first half of 2026. Case citations verified against court records where available; entries relying on secondary practitioner sources are noted. Entries [1–3] are preprints; status noted in the text.
______________________________________________________________________
Declaration on the use of AI tools. This paper was developed with the assistance of Claude (Anthropic, Claude Opus 4.6). The instrument was used for structural drafting, LATEX formatting, editorial iteration, and bibliographic cross-referencing. All substantive claims, mathematical proofs, legal analysis, doctrinal positions, and architectural decisions are the author’s. The instrument produced no assertion that entered the final text without human verification at the gate. The verification architecture described in this paper was applied to its own production.
This article is based on and should be read together with the broader doctrinal framework developed in A. Kildeev and T. Liebig, Instrumentum Vocale and the Architecture of Responsibility: From Liability to Embedded Accountability, SSRN Working Paper No. 6868460, 2026.↩︎
Mata v. Avianca, 678 F.Supp.3d 443 (S.D.N.Y. 2023)↩︎
Johnson v. Dunn, 792 F.Supp.3d 1241 (N.D. Ala. 2025).↩︎
Charlotin, D., AI Hallucination Exposed Cases Database, damiencharlotin.com/hallucinations (2024– 2026). Over 1,450 entries as of May 2026.↩︎
Appendix A, Cluster 5: 33 cases and institutional instruments across the United States, Canada, the United Kingdom, the European Union, Australia, India, Singapore, Colombia, South Africa, and Italy.↩︎
Dahl, Magesh, Suzgun, Ho, “Large Legal Fictions,” J. Legal Analysis 16(1): 64–93 (2024).↩︎
Magesh et al., “Hallucination-Free?” J. Empirical Legal Studies (2025); arXiv:2405.20362.↩︎
Patrick, J. (2026). “Inverse Advantage.” Working paper, Oracle Defence and National Security.↩︎
Liebig, Boundary 1–4 proofs (SSRN Working Paper Nos. 6868278, 6868281, 6868284, 6868318; 2026). Consolidated: The Four Boundaries of Institutional Verification (SSRN Working Paper No. 6868398, 2026). The boundaries were derived from the instrumenta digitalia framework (Liebig, Instrumenta Digitalia Vobis Mando, 2026) and developed in Liebig and Kildeev, Instrumentum Vocale and the Architecture of Responsibility (SSRN Working Paper No. 6868460, 2026), Chapter 2 (derivation chain), Chapter 4 (architectural derivation).↩︎
A.M. Hussain and S. Salahuddin, “Beyond Context: Large Language Models’ Failure to Grasp Users’ Intent,” arXiv:2512.21110v3, KTH Royal Institute of Technology, 2026.↩︎
NIST AI 100-1, “AI Risk Management Framework” (Jan. 2023); NIST AI 600-1, “Generative Artificial Intelligence Profile” (Jul. 2024); EO 14365 (Dec. 11, 2025).↩︎