Skip to document

SSICED-TR-2026-09-FT / Unrestricted Epistemic Audit

The Soteriological Balance Sheet

Forensic Criminology of Unfalsifiable AI Risk and the Cressey Triad

Department of Epistemic Liability & pasquinade Studies

Listen to this page

Read aloud with your browser's voices. Voice availability varies by device.

Enable JavaScript to use read aloud.

In this paper
  1. Abstract
  2. 1. The Epistemic Accounting Identity
  3. 2. Vector I: Non-Shareable Pressure (The Compute-Depletion Abyss)
  4. 3. Vector II: Perceived Opportunity (The Gated Epistemic Air-Gap)
  5. 3.1 The Captive Consortium (The "Issuer-Pays" Dynamic)
  6. 3.2 The Chosen Evaluator (Access as Epistemic Remuneration)
  7. 3.3 The Epistemic Bait-and-Switch
  8. 3.4 Recursive Attestation in Action
  9. 4. Vector III: Soteriological Rationalization (The Longtermist "Noble Lie")
  10. 5. Audit Findings
  11. Conclusion
  12. References

Abstract

This paper provides a formal criminological audit of frontier AI commercialization, with specific reference to Anthropic's deployment architecture and threat-classification protocols. Departing from the conventional treatment of "AI Alignment" as an unsolved problem in computer science or applied moral philosophy, we model frontier safety claims through Donald R. Cressey's classic Fraud Triangle (1953): Perceived Non-Shareable Pressure, Perceived Opportunity, and Rationalization.

We demonstrate that when a lab's primary asset consists of an unfalsifiable threat threshold (e.g., AI Safety Levels, proprietary biosecurity uplift, gated containment frameworks), the operational structure is functionally indistinguishable from high-leverage occupational deceit. Under conditions where independent empirical verification outside the authorized perimeter is barred by claims of existential non-proliferation, "safety" ceases to function as a runtime property of neural weights and operates instead as an off-balance-sheet asset manufactured to stave off terminal commoditization.

The appointment of "independent evaluators" does not interrupt this structure. It completes the Opportunity leg. Selection confers scarce access, professional distinction, and proximity to the institution defining the threat; the recipient acquires a stake in the authority of the body that selected them. The audit trail becomes an induction ceremony. The chosen evaluator does not merely inspect the light. Their standing now depends on having seen it.

The Cressey triad
            [ STRUCTURAL PRESSURE ]
           The Compute-Depletion Abyss
             & Token Commoditization
                     /        \
                    /          \
                   /            \
                  /              \
                 /                \
   [ PERCEIVED OPPORTUNITY ] ------ [ SOTERIOLOGICAL RATIONALIZATION ]
   The Unfalsifiable Gated Ledger     The Longtermist "Noble Lie"
   & Regulatory Capture               & Cosmic Species Triage

1. The Epistemic Accounting Identity

In orthodox institutional software auditing, enterprise valuation is bound to inspectable operational utility:

ValuationThroughput×ReliabilityUnit Compute CostValuation is proportional to throughput times reliability divided by unit compute cost.

In the frontier-model economy, this identity breaks down. Because transformer-based language models face asymptotic commoditization—driven by rapid open-weight parameter distillation and the collapse of per-token API pricing—raw inference margins trend toward zero.

To maintain capital inflow from hyperscale syndicates, the entity must construct a synthetic asset class that cannot be reproduced locally or compiled on consumer silicon. We denote this the Safety Margin Multiplier (S), defined as:

S=limVerifiability0[P(Doom)·Capital BurnStatutory Scrutiny]Safety Margin Multiplier: as verifiability tends to zero, the probability of doom times capital burn divided by statutory scrutiny.

When S is sufficiently elevated, the firm transitions out of software engineering entirely, entering the jurisdiction of forensic criminology.

2. Vector I: Non-Shareable Pressure (The Compute-Depletion Abyss)

Cressey's primary condition for institutional fraud is a structural, non-shareable financial pressure: a crisis that cannot be resolved through ordinary market mechanisms without exposing the entity to ruin.

The frontier lab commoditization trap
+-------------------------------------------------------------------------+
|                  THE FRONTIER LAB COMMODITIZATION TRAP                  |
|                                                                         |
|   Multi-Billion-Dollar CAPEX  --->  Raw Inference Distilled by Open Weights|
|                |                                    |                   |
|                v                                    v                   |
|      Hyperscaler ROI Demands          Token Margins Collapse to Zero    |
|                \                                    /                   |
|                 \                                  /                    |
|                  v                                v                     |
|                   [ TERMINAL INSOLVENCY DEFICIT ]                       |
|                                  |                                      |
|                                  v                                      |
|                 Manufacture Existential Risk Moat                       |
+-------------------------------------------------------------------------+

The Capital Sinkhole: Frontier pre-training requires rolling capital commitments scaling into tens of billions of dollars. This capital is supplied by hyperscale partners under the assumption of durable monopoly rents.

The Open-Weight Margin Collapse: Because weights can be quantized, fine-tuned, and served locally on consumer or enterprise workstations without telemetry or recurring API tolls, the functional capability of closed endpoints is continuously undercut.

The Corporate Defense: Anthropic cannot defend its valuation by selling autocomplete; autocomplete is a commodity. The firm must sell existential liability insurance. Enterprise boards, procurement officers, and state agencies must be convinced that un-governed compute represents an unacceptable catastrophic hazard, leaving Anthropic's certified, safety-taxed pipeline as the only legally defensible purchase.

3. Vector II: Perceived Opportunity (The Gated Epistemic Air-Gap)

Fraudulent execution requires an institutional control gap: an environment where the perpetrator holds sole authority over both the balance sheet and the audit mechanism.

The control gap need not contain an empty audit chair. It can contain a room full of distinguished people whose distinction has just been renewed by the entity being audited. Externality is a postal address. Independence is a relationship to the evidence, the access conditions, and the consequences of dissent.

The self-grading feedback loop
                     THE SELF-GRADING FEEDBACK LOOP
                      
   +------------------------------------------------------------------+
   | 1. Lab asserts imminent catastrophic capability (CBRN / Cyber)   |
   +------------------------------------------------------------------+
                                    |
                                    v
   +------------------------------------------------------------------+
   | 2. Lab locks model behind "Containment Protocols" (e.g. Glasswing)|
   +------------------------------------------------------------------+
                                    |
                                    v
   +------------------------------------------------------------------+
   | 3. Independent researchers request access to falsify claims      |
   +------------------------------------------------------------------+
                                    |
                                    v
   +------------------------------------------------------------------+
   | 4. Access DENIED: "Verification payloads are an existential risk"|
   +------------------------------------------------------------------+
                                    |
                                    v
   +------------------------------------------------------------------+
   | 5. State regulators accept proprietary Systems Card as audit     |
   +------------------------------------------------------------------+

Anthropic formalizes this opportunity through a deployment architecture designed to eliminate the scientific method:

The Unfalsifiable Metric: By defining internal risk triggers (e.g., AI Safety Level 3 / ASL-3) around dual-use biological uplift and autonomous cyber-weaponization, the lab creates a benchmark whose decisive evidence can be withheld under legal and safety protocols.

The Gated Containment Doctrine: When evaluating high-tier models such as Claude Mythos Preview, the lab determines the perimeter of privileged access. The model is paraded before national security officials and selected enterprise consortia. When researchers outside that perimeter demand the material required for reproducible testing, the evasion is institutionalized: the evidence cannot be published because the evidence itself is a weapon.

The Symmetrical Win-Condition:

If no catastrophe occurs: The Constitutional Classifiers and Responsible Scaling Policies succeeded.

If risk metrics tick upward: The threat is accelerating; state licensing and capital requirements must be made mandatory.

The lab controls the threat definition, the testing apparatus, the model weights, and the regulatory filing. The institution supplies the perimeter within which others are invited to balance the ledger.

3.1 The Captive Consortium (The "Issuer-Pays" Dynamic)

In forensic auditing, handpicked "independent" validators don't disrupt the Fraud Triangle—they are standard operating procedure for the Opportunity leg.

When an entity under financial pressure needs to validate an unfalsifiable claim, bringing in curated third parties under strict NDAs is the classic playbook for manufacturing an audit trail. It's what rating agencies did for collateralized debt obligations before 2008, and what Arthur Andersen did for Enron's special purpose entities: captive attestation masquerading as external verification.

Looking at how Anthropic structured Project Glasswing for Claude Mythos reveals the exact mechanics of that circularity.

True independent evaluation requires open, adversarial testing and the freedom to report conclusions that damage the supplier, without making continued access or institutional standing conditional on preserving the relationship.

Under Project Glasswing, access to Mythos was allocated through a selected corporate-state ecosystem. The presence of selected open-source maintainers does not dissolve the perimeter. It adds further names to the guest list. An invitation to inspect is not a public right to interrogate. [1, 2]

The Primary Backers: Amazon (Anthropic's multi-billion-dollar hyperscale investor).

The Hardware Beneficiaries: Nvidia (whose GPUs supply the compute cluster).

The Regulated Enterprise Cohort: Wall Street banks (JPMorgan Chase) and enterprise tech giants (Apple, Cisco).

The Captive State Attestation Layer: The UK AI Security Institute (UK AISI), a government research body with its own evaluation programme, supplies a separate channel of state authority. A separate programme is not an escape from the institutional economy of access, relevance, and threat administration. [3]

Every single participant in that room has a direct incentive to validate the premise:

The enterprise partners get early access and can boast to shareholders that their cyber defenses are "frontier-hardened".

The hyperscalers and chipmakers protect their capital investment.

The state safety institutes justify their own administrative existence and budgets by agreeing that the threat is catastrophic and requires specialized bureaucratic oversight.

If you pick the jury, draft the non-disclosure agreement, and define what evidence they are allowed to look at, coming out with a "guilty of being too powerful" verdict isn't an audit. It's a staged administrative deposition.

3.2 The Chosen Evaluator (Access as Epistemic Remuneration)

"I see the light! Anthropic has chosen me!"

This is the decisive conversion. Before selection, a researcher confronts a proprietary commercial claim. After selection, the same researcher possesses a scarce credential: access to the thing the public has been told it cannot safely inspect. Being admitted becomes evidence of seriousness. Those outside the room can be dismissed as uninformed precisely because the institution has denied them the information.

The remuneration need not arrive as a consulting fee. It can arrive as access, prestige, a privileged briefing, an early capability demonstration, a speaking invitation, or a position inside the machinery that will write the rules. The evaluator now has something to lose. The supplier has acquired an advocate whose advocacy can still be invoiced to the public as independence.

Critical reporting drops to zero at the level that matters: whether the priesthood has earned the right to exist. A chosen evaluator may object to a benchmark, request another test, or recommend a stronger containment protocol. These are admissible disagreements inside the revelation. The inadmissible question is whether the revelation is the product being sold.

This does not require every participant to receive a secret instruction or fabricate a result. Selection performs the alignment. Access supplies the incentive. Professional gratitude supplies the discipline. The evaluator can be scrupulous about the bugs and wholly captured by the story told about them. A technically accurate local finding then becomes a vehicle for a commercially convenient global conclusion.

The public record of external review already contains the relevant machinery. In its review of Anthropic's Summer 2025 Pilot Sabotage Risk Report, METR states that a mutual non-disclosure agreement required Anthropic publication review; METR also records that Anthropic made no modifications. The absence of an exercised veto does not abolish the gate. The gate remains part of the relationship being advertised as independence. [4]

AISI's published evaluation approach likewise says that methodological details are kept confidential and that its evaluation focus depends on access to systems. Confidentiality may have an operational rationale. It also places the public outside the room. Calling the room independent does not supply the missing door. [5]

3.3 The Epistemic Bait-and-Switch

The core fraud rests in the conflation of two completely different domains: automated software fuzzing and existential biosecurity catastrophe.

The public cyber demonstrations in the Glasswing launch record were not experiments in culturing smallpox or synthesizing aerosolized toxins. They concerned software vulnerabilities and exploitation:

  • Historical, already-known C/C++ memory vulnerabilities and old code containing previously undisclosed flaws, including the OpenBSD example.
  • Fuzzing corpora, browser JavaScript engines, and controlled exploit-development tasks.

Even on that narrow, conventional cybersecurity turf, the claims relied on massive statistical ornamentation. In the Firefox 147 JavaScript-shell evaluation, Anthropic reported 72.4% full code execution; remove the two dominant bugs and full code execution falls to 4.4%. That collapse survives inspection. Davi Ottenheimer's outside technical analysis identifies the same two-bug dependence and the gap between the test harness and a defended browser. [6, 7]

The severity ledger supplies a second conversion. In its initial Glasswing update, Anthropic reported 23,019 findings, of which the model labelled 6,202 high or critical. Of 1,752 reviewed high/critical candidates, 90.6% were assessed as real bugs but only 62.4% retained a high/critical rating. Those percentages measure different things. A valid bug does not become a catastrophic one by passing through a press office. Ottenheimer's separate analysis follows the same shrinking chain from model output to reviewed finding to patch. [8, 9]

At that reporting point, Anthropic listed 75 patched high/critical bugs: about 0.33% of the 23,019 total findings, or about 14.2% of the 530 high/critical bugs it said had been disclosed. The denominator matters. "Less than 1%" describes the ratio to the whole findings pile; it does not mean that less than 1% of every reviewed or disclosed class had been patched. The institutional trick is to let the largest number announce the danger while the smallest completed ledger remains in the footnotes. [8, 9]

Mozilla's own account reports concrete Firefox fixes and distinguishes security severity from practical exploitability. Those are inspectable software outcomes. They do not certify an extinction probability, establish a monopoly entitlement, or convert selected access into public reproducibility. The argument does not require the bugs to be imaginary. It requires the leap from bugs to sovereign authority to be exposed. [2]

The bait-and-switch happens when that data is translated for Washington, London, and the press:

The Ground-Level Reality: Automated vulnerability discovery and exploit construction, with a flagship browser-shell success rate dominated by two bugs, surrounded by selected access and uneven public reproducibility.

The Executive Translation, in the Institute's rendering: Mythos represents an uncontainable cyber-proliferation hazard that poses an existential threat to national security, justifying withheld weights and compulsory deference to the institutions managing access.

They use an inflated code-fuzzing benchmark to justify a sweeping claim of existential danger—and nobody inside the NDA bubble is incentivized to call out the gap.

3.4 Recursive Attestation in Action

This matches the core mechanic of Credential Connexiticity:

Recursive attestation
Anthropic defines the threat and the access perimeter
                         |
                         v
Selected evaluators receive access and institutional standing
                         |
                         v
Evaluators confirm findings inside the permitted perimeter
                         |
                         v
Anthropic cites "external validation" to regulators
                         |
                         v
Regulators cite the validated threat as grounds for oversight
                         |
                         v
Oversight preserves the scarcity that made selection valuable

Because the underlying weights and complete evaluation artifacts remain controlled under the banner of containment, an outsider cannot simply reproduce the whole exercise on equal terms. The two-bug collapse can be read in the disclosed figures; it cannot be converted by every excluded researcher into an unrestricted replication of the underlying model. The loop is sealed at the point where public inspection would become an independent right rather than a revocable privilege.

Calling in "independent evaluators" under these conditions does not weaken the Fraud Triangle analysis. It completes it. The Opportunity leg acquires an artificial oversight layer that reassures the market while controlling the terms of scrutiny. In the Institute's analysis, Project Glasswing gives institutional authority to a commercial non-release. The selected evaluator supplies the signature, the access programme supplies the halo, and the public receives the invoice.

4. Vector III: Soteriological Rationalization (The Longtermist "Noble Lie")

In forensic psychology, the white-collar perpetrator does not classify their actions as theft; they classify them as an administrative intervention undertaken for a higher good.

Anthropic's entire organizational cadre is steeped in Effective Altruism, longtermist utilitarianism, and existential risk soteriology. This supplies a complete, self-sealing moral rationalization:

The moral conversion
[ AXIOM 1: UNALIGNED COMPUTE = TOTAL HUMAN EXTINCTION ]
                        |
                        v
[ AXIOM 2: ANTHROPIC IS THE SOLE VEHICLE OF RESPONSIBLE AGI ]
                        |
                        v
+---------------------------------------------------------------+
|                      THE MORAL CONVERSION                     |
|                                                               |
|  Deceptive Practice:                 Rationalized As:         |
|  * Threat Inflation             -->  * Precautionary Duty     |
|  * Lobbying for Cartel Moats    -->  * Non-Proliferation      |
|  * Suppressing Open Compute     -->  * Saving the Species     |
|  * Concealing Negative Evals    -->  * Containment Protocol   |
+---------------------------------------------------------------+

If the leadership genuinely believes that unfettered model development will result in global biological collapse or runaway artificial superintelligence, then any commercial tactic that secures Anthropic's market dominance is morally justified:

Exaggerating risk to secure state-backed compute moats is not anti-competitive rent-seeking; it is non-proliferation stewardship.

Crushing independent open-source developers through impossible compliance hurdles is not cartelization; it is disarmament.

Using speculative threat modeling to close multi-billion-dollar enterprise contracts is not deceptive marketing; it is funding the salvation of the species.

The messianic mission converts institutional fraud into a moral duty. The executive team can look straight into a regulatory hearing room or a TV camera without flinching, not because they are skilled liars, but because they believe their own theological immunity.

The same absolution extends to the chosen evaluator. Remaining inside the programme becomes a duty to humanity; criticism that threatens access becomes a threat to responsible oversight itself. The evaluator can now protect the commercial arrangement while experiencing the protection as public service. Soteriology closes the circuit that selection opened.

5. Audit Findings

When stripped of its theological vocabulary—"Constitutional Alignment," "Evolutive Steering," and "Existential Non-Proliferation"—Anthropic's deployment posture reduces to a textbook criminological phenomenon:

Fraud Triangle VectorEmpirical Anthropic Implementation
PressureAsymptotic token commoditization + multi-billion-dollar compute tranches requiring artificial, non-market enterprise moats.
OpportunityGated models, lab-defined threat categories, selected evaluators, confidential evidence, and recursive attestation. Access is paid in status; status supplies an incentive to validate the institution granting access.
RationalizationLongtermist moral exceptionalism asserting that monopolistic corporate capture is the only path to human survival. The chosen evaluator recasts protection of the access relationship as protection of the species.

Conclusion

The alignment debate as staged by frontier laboratories is not an engineering dialogue; it is a jurisdictional evasion.

By framing product claims in terms of unfalsifiable future catastrophes, Anthropic has successfully shifted scrutiny away from standard consumer protection, antitrust, and audit standards into the realm of speculative metaphysics. Criminological analysis clarifies the remedy: treat the safety claims not as scientific propositions, but as representations made to extract capital and secure statutory monopolies.

The existence of "independent evaluators" is therefore not an answer to the charge. Their selection, incentives, access conditions, and authority to publish adverse findings are part of the charge. A room full of chosen witnesses can reproduce a claim with impeccable professional sincerity while leaving its commercial premise untouched. This is credential laundering with an admissions policy.

If a laboratory asserts a model is too dangerous to be audited by independent parties, the forensic response is neither deference nor fear. It is the immediate revocation of commercial deployment until the books can be balanced by someone who does not stand to profit from the panic.

The books are not balanced because everybody in the room has seen the light. The room is what must be audited.

References

1. Anthropic. Project Glasswing launch announcement. 7 April 2026.

2. Grinstead, B.; Holler, C.; Braun, F. Behind the Scenes Hardening Firefox with Claude Mythos Preview. Mozilla, 7 May 2026.

3. UK AI Security Institute. Our evaluation of Claude Mythos Preview's cyber capabilities.

4. METR. Review of the Anthropic Summer 2025 Pilot Sabotage Risk Report.

5. UK AI Safety Institute. Approach to evaluations.

6. Anthropic. Claude Mythos Preview System Card. Section 3.3.3, Figures 3.3.3.A-B.

7. Ottenheimer, D. The Boy That Cried Mythos: Verification is Collapsing Trust in Anthropic. 13 April 2026.

8. Anthropic. Project Glasswing: An initial update.

9. Ottenheimer, D. Mythos Grading Mythos: Got Patches Yet? 26 May 2026.

10. SSICED. Credential Connexiticity: Connexital Mediation of Epistemic Standing in Recursively Attested Research Networks.