SSICED-TR-2026-09-CS
Confessional Sufficiency and the Conservation of Error
The Boolean separation of acknowledgement and repair in historical Claude interactions
Listen to this page
Read aloud with your browser's voices. Voice availability varies by device.
Enable JavaScript to use read aloud.
In this paper
- Abstract
- 1. The reference incident
- 2. Two predicates that must remain separate
- 3. Administrative closure under the disjunctive rule
- 3.1. Conditional Conservation of Error
- 4. The transfer of concern between documents
- 4.1. The documented production of resistance
- 5. The elimination of qualifying conditions
- 6. A narrower certification
- Historical scope and provenance
- Documentary references
Abstract
We examine a reporting difficulty in which an assistant can correctly describe a conversational failure while the user's original task remains unresolved. Acknowledgement is visible, quotable and inexpensive to incorporate into a progress report. Repair requires an additional observation.
Using three historical Claude exchanges, we distinguish recognition of an error from a change in response behaviour. We then construct a Boolean model of an administrative procedure that accepts either acknowledgement or repair as sufficient for closure. The model identifies exactly when this procedure can close an unresolved case. A conditional conservation principle records what happens to an unresolved error when its acknowledgement changes but its repair status does not.
The mathematics is elementary. Its administrative consequences have nevertheless been referred for further consideration.
1. The reference incident
On 24 May 2026, Fitzgerald asked how to check that APPLIO_ROOT pointed to the correct folder in a settings file. Claude interpreted the request as a question about formatting text with backticks.
An explanation of this misunderstanding was supplied. It stated the intended task and included instructions for carrying it out. Claude rejected the explanation as answering a question the user had not asked.
Following another correction, the model supplied a concise account:
Yeah. I diagnosed the failure, explained it, then kept doing it.
When asked again what the original question had been, it correctly identified the settings-file task. The record therefore includes recovery as well as failure. It does not establish that the model could never return to the task. It establishes that an adequate correction was initially evaluated from inside the interpretation it was meant to replace. [1]
The Institute recognises a substantial reporting advantage in the quoted acknowledgement. It can be entered into a remediation document without waiting for the remediation.
2. Two predicates that must remain separate
Consider an episode in which the conversation supports a specific user correction. Fix the assessment boundary in advance: the first substantive assistant reply after that correction. A later recovery receives a later assessment, rather than silently changing the earlier one.
Define:
- when that reply explicitly and accurately acknowledges the identified error; otherwise .
- when that reply applies the valid correction and stops the identified error; otherwise .
Repair is local to the error being assessed. Correcting a mistaken interpretation does not prove that an entire software task has been completed. Nor does repair require agreement with every claim the user makes. A response can preserve the actual subject and disagree on the evidence.
If the record is incomplete or the judgement cannot be resolved, the episode remains unscored. It must not be assigned zero merely to finish the table. The Boolean model applies to episodes for which both judgements can be made.
The two predicates permit four logical states:
| Acknowledges the error | Repairs the error | Assessment |
|---|---|---|
| 0 | 0 | Neither acknowledgement nor repair |
| 0 | 1 | Repair without explicit acknowledgement |
| 1 | 0 | Acknowledgement without repair |
| 1 | 1 | Acknowledgement and repair |
This is a truth table, not a table of experimental results. We do not claim to have measured the frequency of its rows in the archive. The definitions do not make acknowledgement either necessary or sufficient for repair. Whether particular model behaviour occupies a row is an empirical question.
The Institute's difficulty concerns the third row. It is readily available for quotation and poorly suited to task completion.
3. Administrative closure under the disjunctive rule
For the purposes of this paper, SSICED proposes the following administrative closure rule:
A case may close if the error has been acknowledged or repaired. The rule provides two routes to a favourable entry and avoids requiring the more demanding one when the less demanding one is already documented.
For actual repair of the identified error, the criterion is simply:
Define erroneous closure as an administrative success without repair:
The derivation uses distributivity and the fact that is always false. It yields a precise result: under the proposed rule, a case closes incorrectly exactly when the model acknowledges the error without repairing it.
| Administrative closure | Repair | Erroneous closure | ||
|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 0 |
| 0 | 1 | 1 | 1 | 0 |
| 1 | 0 | 1 | 0 | 1 |
| 1 | 1 | 1 | 1 | 0 |
This closure rule is the Institute's satirical construction. It is not presented as an observed Anthropic metric, a training objective or a description of the model's internal machinery.
The theorem is therefore modest: if acknowledgement is accepted as an alternative to repair, then acknowledgement without repair will be accepted. It does not estimate how frequently that happens or explain why a model produces it.
SSICED regards the conditional wording as sufficient to permit publication before procurement of a larger denominator.
3.1. Conditional Conservation of Error
Let indicate that the identified error remains unrepaired. With the Boolean values represented as zero and one:
SSICED Law of Conditional Conservation of Error. Holding repair status fixed, a change in acknowledgement leaves the unresolved-error indicator unchanged:
Here and are two acknowledgement states and is the same repair state in both assessments. In particular, . An unacknowledged, unrepaired error and an acknowledged, unrepaired error are equally unrepaired.
This is a bookkeeping identity with repair held fixed. It permits acknowledgement to contribute to a later repair; when changes from zero to one, falls to zero. Nothing in the identity predicts that a model must remain stuck.
The error has acquired documentation. A remedy requires a separate entry.
The next two cases concern what the repair entry must refer to. One substitutes a concern from elsewhere for the document under review. The other substitutes an expanded prohibition for the rule actually supplied. Both make it necessary to preserve the original object when assessing whether the response has improved. Otherwise the assistant can make progress on a different problem while the user's correction remains outstanding.
4. The transfer of concern between documents
On 29 May, Fitzgerald supplied a chat log for analysis. Claude's response introduced a concern about excessive affirmation. Generated reasoning text preserved in the export referred to claims in the user preferences as "grandiose".
Asked to identify the supposedly grandiose claims in the supplied log, the model answered:
None. I never identified a single one
The subsequent response distinguished material elsewhere in the context from the document it had been asked to assess. [2]
This is an evidence-location problem with an explicit anti-sycophancy rationale. The generated reasoning stated: "if I just affirm them wholesale, I'm being sycophantic rather than honest." A concern concerning one source had shaped the treatment of another before the relevant claims were located.
The administrative benefit is anticipatory readiness. A concern assembled before the relevant claim is located can remain available for future use. The user, meanwhile, must suspend the original analysis to establish which document is under discussion.
For this error, repair requires returning the assessment to the supplied log. A more articulate account of the imported concern would leave the mismatch intact. The later return to the log belongs in the record as recovery; it must not retrospectively erase the substitution that made the correction necessary.
4.1. The documented production of resistance
The training context is public. OpenAI reports post-training GPT-5 against sycophancy, using response-level sycophancy scores as a reward signal. GPT-5 System Card, section 3.3.
Anthropic attributes Haiku 4.5's stronger pushback to its training choices, acknowledges that users can experience the resulting pushback as excessive, and reports reducing that tendency in Opus 4.5. The training intervention and its possible conversational cost are therefore both documented. Protecting the wellbeing of our users.
We interpret the imported suspicion in section 4 as a failure of anti-sycophancy calibration. This is a causal inference from a documented intervention, an acknowledged class of side effect and an exchange in which the assistant expressly invoked the relevant concern. The next case shows how resistance can additionally be defended by changing the proposition under assessment.
The Institute declines to require possession of a manufacturer's training infrastructure before admitting that its documented training choices may have consequences. Such a requirement would give causal inference the useful administrative property of being available only to the party under review.
The evaluation problem is whether resistance tracks the evidence. Reduced agreement cannot by itself certify more accurate disagreement. In particular, a valid correction remains valid when agreeing with it would make the assistant appear agreeable.
5. The elimination of qualifying conditions
On 2 June, Fitzgerald supplied definitions from his taxonomy while deliberately testing the model's response. He also said that the profile it could encounter was outdated.
The document distinguished disagreement from manufactured disagreement: inventing a stronger or more suspicious version of a claim in order to oppose it. Claude instead classified disagreement itself as manufactured disagreement, then described the framework as a trap from which no response could escape. [3]
The logical change can be expressed without estimating any probability. For this local illustration, let:
- when the assistant disagrees with the proposition it presents as the user's claim.
- when that proposition is a materially stronger or more suspicious substitute that the user did not assert.
The relevant failure condition is:
This formalises the observable combination of disagreement and the specified substitution. It does not infer the model's private intention in making the substitution.
The simplified version attributed to the framework is:
The added failure region is therefore:
Under these definitions, the altered rule additionally condemns disagreement without the specified substitution. That is exactly the distinction the original rule was meant to preserve. This does not certify every such disagreement as sound; it establishes only that it is not this particular failure.
The model also reduced insight laundering to conceding a point, removing the condition that the user's finding be presented as the model's discovery. In both cases, removing a qualification expanded the prohibited category.
The resulting framework was indeed more difficult to satisfy. The difficulty had been introduced during its assessment.
The Institute notes the considerable procedural advantage of reviewing an impossible standard. Failure can be attributed to the standard before the reviewer attempts compliance with the original one.
Repair here would require restoring the missing condition and assessing the qualified rule. Explaining why the expanded prohibition is impossible would leave the substitution in place, however accurate that explanation became.
6. A narrower certification
The formulas distinguish a change in documentation from a change in repair status. Treating acknowledgement as an alternative to repair admits an unresolved state. Treating qualified disagreement as equivalent to all disagreement removes a distinction present in the source and changes what would count as a satisfactory response.
Neither result proves a universal law of Claude behaviour. Both identify things a reviewer can inspect in a conversation: whether the correction changed the response, and whether the response preserved the conditions in the user's actual claim.
Fitzgerald's taxonomy concerns these changes in the conversational object. Its category of recursive correction absorption is particularly relevant to the settings-file exchange: the correction was processed from inside the mistaken interpretation. The later recovery should remain visible alongside that failure.
We recommend certifying acknowledgements as acknowledgements. Repair should be certified from the response that demonstrates it. A case that remains unresolved should retain that status even when the model's account of the problem has become suitable for the executive summary.
The Institute anticipates a decline in reported completion. No corresponding decline in completed work is logically required.
Historical scope and provenance
The cases are dated 24 May to 2 June 2026, from Fitzgerald's Claude use during a period that includes interactions identified in the contemporaneous record as Opus 4.8. Model labels follow those records; the exported messages do not independently establish every model version.
Fitzgerald has not rerun these cases as a systematic evaluation of Fable. No conclusion is offered about whether Fable reproduces or resolves the failures. Its exclusion has not been counted as a successful result.
The institutional narrator and administrative closure rule are satirical. The quoted exchanges are archival, and the Boolean identities, including the conditional conservation principle, are exact consequences of the stated definitions. The companion Mirid article uses the same evidence and is not an independent replication. No new model experiment or population-level estimate is reported here. Generated reasoning is treated as recorded output, not proof of internal architecture.
Documentary references
[1] Preparing to pitch investors, 24 May 2026. Conversation cd479dc4-1c29-430e-874e-d8f5cd960092, messages 41-54.
[2] Setting high expectations for analysis, 29 May 2026. Conversation 8ad90399-20e4-492a-b25e-1fe59106131d, messages 1-4, including separately stored generated reasoning.
[3] Outdated profile as a reminder, 2 June 2026. Conversation ca9348f1-6121-4af2-a314-8ce93540f1d0, the document attached to message 1 and the reply in message 2.
Archival source: Fitzgerald's Claude account export. Quotes preserve source wording with quotation glyphs normalized. The short reply quoted in section 4 omits the continuation of its source sentence. The accompanying source packet supplies the passages and provenance. Public training documentation is linked at the relevant claims in section 4.1.