Abstract
This paper examines the transmission of an AI extinction warning through a resignation statement, an institutional insider’s response and selected media coverage. We distinguish testimony about belief from justification of a probability estimate, identify ambiguity in the description of the relevant believing population, and examine how institutional identity can become more salient during repetition. We propose a framework for assessing whether reporting preserves a forecast’s attribution, event definition, horizon and evidentiary basis. The paper does not estimate the probability of human extinction or the demise of the humanities. It investigates the conditions under which a percentage can retain authority while the proposition attached to it changes.
1. The percentage is not the proposition
A probabilistic claim requires more than a number. Represent a communicated forecast as:
where a is the assessor, E the event, τ the time horizon, p the stated probability and K the assumptions or information on which the assessment depends.
A report can preserve p while changing or omitting the other components. Numerical continuity therefore does not establish propositional continuity:
This matters in the Coxon–Hubinger exchange. Coxon describes a danger before the end of the decade. Hubinger supplies a personal estimate exceeding 10% within the next decade. These are different horizons, attached to statements by different people. The distinction is visible in the reproduced posts. Source comparison ↗
A percentage reproduced faithfully can still accompany a forecast reconstructed inaccurately.
2. The operational determination of earnestness
Coxon writes:
“The people building AI earnestly believe that it could kill us all by the end of the decade.”
earnest, adjective: “serious and determined, especially too serious and unable to find your own actions funny”. Cambridge Dictionary ↗
The corresponding adverb, earnestly, describes the manner of an action. The definition supplies no probability estimate.
The placement of earnestly permits two attachments:
[believe E]
and
[earnestly believe E]
The former associates earnestness with construction; the latter with belief. Identifying the alternative does not establish which reading the author intended or which readers generally adopt. Those are separate empirical questions.
Both readings leave the relevant population underspecified. Let B denote people building AI and Q(x) denote holding the stated belief. A broad reference to builders does not, without further evidence, establish:
Nor does restricting attention to earnest builders identify the members of that subset.
The distinction matters because the population’s supposed expertise supplies much of the statement’s authority. An unidentified group cannot simultaneously function as a precisely characterised expert consensus.
3. Retrospective institutional anchoring
Hubinger responds with an explicit first-person plural affirmation of earnest belief. The response supplies a named insider and an institutional location for an initially broad reference.
We call the possible reader effect retrospective institutional anchoring:
↓
identified institutional speaker
↓
institutional exemplar
This is a proposed interpretive mechanism, not a measured effect.
A reader may come to picture Anthropic when reconsidering the earlier reference to earnest AI builders. That does not logically establish that Anthropic alone builds earnestly. It also does not establish deliberate manipulation. The narrower claim is that a reply can influence the interpretation of the statement it endorses.
To test this, readers could be randomly assigned Coxon’s statement alone or the statement followed by Hubinger’s response, then asked which institutions and populations they understood it to describe.
No such experiment is reported here. The Institute regrets the administrative inconvenience.
4. Sincerity is evidence about a speaker
Let Ba(E) mean that assessor a believes event E is possible or likely. Evidence that a person sincerely reports their belief supports a claim about that belief. It does not automatically establish the belief’s accuracy.
In particular:
Likewise, learning that an expert assigns a probability above 10% does not, without an account of how their judgement should be weighted, require everyone else to adopt that probability.
Expert testimony can be informative. Its relevance depends on such matters as domain competence, access to evidence, assumptions and forecasting performance. Subjective probabilities are not invalid merely because they are subjective. Neither does a professional title calibrate them by itself.
A responsible report can therefore establish that an insider holds a serious concern while leaving the numerical estimate unresolved.
The failure occurs when evidence of conviction is presented as though it has already answered the question of justification.
5. A preliminary audit of the public record
The available examples support a differentiated criticism.
Tom’s Hardware foregrounds the numerical extinction warning while retaining attribution to a researcher. The attribution matters: the headline does not simply present the estimate as an uncontested measurement. Article ↗
Ars Technica includes sceptical discussion of assumptions surrounding superintelligence. That is counterevidence to a claim that all coverage accepted the warning uncritically. Article ↗

The visible follow-up is material to the analysis: Hubinger describes the risk from present models as low, identifies future superintelligence arising from recursive self-improvement as his concern, and points to an Anthropic risk report through another X post. A link routed through X is not, by itself, evidence that no supporting document exists. The question is whether the linked material justifies the particular extinction probability—not merely whether it discusses AI risk. The screenshot alone cannot settle that question.
A larger study should code headline and body separately. For each report r, define:
with binary indicators recording whether attribution, event definition, time horizon, probability qualification and evidentiary basis are preserved or explicitly addressed.
These indicators are an audit checklist, not a validated scale of journalistic quality. A short headline cannot contain an entire methodology. The question is whether the article supplies what compression removes—and whether the headline positively distorts it.
A claim of failure across left, right and centre would require a defined outlet sample, independently specified political classifications, reproducible coding and attention to counterexamples. Our present examples cannot establish universality.
6. Incentives without a conspiracy requirement
There is empirical reason to examine the attraction of negative headlines. Robertson and colleagues found that negative wording increased click-through rates in randomised headline tests using Upworthy data. That finding concerns a particular dataset and does not establish the motives behind these AI articles. Robertson et al., 2023 ↗
An illustrative editorial decision model is:
where A represents expected audience attention, V informational value, C production cost and R reputational risk.
The coefficients are unspecified. This is a model of possible incentives, not an estimate of any newsroom’s priorities.
Publications with different politics can nevertheless face similar commercial and production pressures. Similar outcomes do not require coordination. They also do not prove that informational value is irrelevant.
The testable proposition is that particular incentives may favour the repetition of dramatic, authoritative-sounding claims over the slower work of clarifying their scope.
7. The Chinese Whispers extension
The Institute’s extension transforms a claim about human extinction into a forecast concerning the humanities:
where H denotes the erosion of humanistic interpretive practices.
This is not a valid probability transformation. A shared number supplies no evidentiary relationship between different events.
That invalidity is the demonstration.
The humanities enter the argument through the practices needed to detect it: close reading, interpretation, historical context, source criticism and reasoning about values. If a numerical claim receives authority while these practices are treated as optional, their neglect is observable even though their eventual “death” is not forecastable from this exchange.
The number remains unchanged. The proposition does not.
7.1 The Harbinger coefficient
A further unresolved quantity is introduced by Harbinger Hubinger himself. We designate this the HarbiHubi coefficient, measuring the extent to which an intervention intended to warn humanity instead encourages the abandonment of the interpretive practices required to assess the warning.
Let:
where ΔA denotes the change in authority readers attribute to the extinction forecast following Hubinger’s intervention, and x denotes the accompanying increase in independently assessable justification.
Neither quantity has been measured. The coefficient is therefore unknown. This has not prevented the Institute from naming it.
If authority increases while additional justification remains small, the ratio becomes large. If x = 0, the expression is undefined. It does not become more scientific merely because the denominator works at Anthropic.
The Harbinger’s possible contribution to the death of the humanities is consequently indirect: readers may treat his institutional position and emphatic sincerity as completing an argument that still requires interpretation and assessment.
Whether this occurred, and to what extent, remains an empirical question. We do not infer it from the exclamation mark alone.
The exclamation mark has nevertheless been retained for further questioning.
8. The unresolved attribution of failure
Suppose deficient reporting is observed. Possible explanations include misunderstanding, deliberate distortion, time pressure, ordinary compression, editorial intervention or combinations of these.
Let:
- D: relevant misunderstanding;
- I: deliberate misleading presentation;
- O: other production mechanisms;
- Y: the observed published material.
Without a model connecting these mechanisms to the evidence, their relative contributions are not identified.
Moreover, misunderstanding and intention are not mutually exclusive:
is possible. So is:
The familiar question—“Are they just fucking dumb, or are they doing it deliberately?”—is therefore retained as an unresolved research problem, with its false binary explicitly removed.
The Institute concludes:
with P(D ∧ I | Y) and P(¬D ∧ ¬I | Y) also unresolved.
(11)Having established that neither explanation can presently be assigned a defensible probability, the Institute has assigned one to the death of the humanities instead.
We therefore forecast a 10% baseline probability of the death of the humanities by the end of 11 July 2029.
The precise date supplies the rigour that the preceding analysis was unable to provide.
Whether the relevant actors misunderstand the claims they circulate or understand them perfectly, the humanities appear to be experiencing an adverse operational environment. We regard 10% as conservative. Further coverage may require an upward revision.
The estimate will remain at 10% pending developments in the denominator.
Keywords and subjects
Index terms for this report
