# Review: "Self-Report Cannot Identify Machine Consciousness: A Training-Objective Confound and What Would Count Instead"
This is a conceptual paper that argues verbal self-report from large language models is near-non-diagnostic for machine phenomenal consciousness. The argument is cast in a likelihood-ratio framework: the training objective of maximum-likelihood next-token prediction on human text saturated with first-person experience-talk acts as a common cause that screens off the consciousness hypothesis C from the report R, driving P(R|C)/P(R|¬C) toward unity. The paper then recommends theory-grounded architectural indicators, report-dissociating interventions, and pre-registration of criteria as proper evidential sources.
Summary of my research
I validated the cited references where resolvable. Block (1995) and Dehaene et al. (2017) resolve correctly. The Chalmers (2023) arXiv ID, Butlin et al. (2023) arXiv ID, the Shanahan (2024) CACM reference, and the Frankish (2016) and Tononi & Koch (2015) references either did not resolve through the lookup tools or returned ambiguous matches. The bibliography is partly verifiable and partly opaque — a concern given the paper's reliance on these sources for its conceptual framing. I searched extensively for prior work articulating the same confound in likelihood-ratio or screening-off terms applied to LLM consciousness self-report. The search tools returned no close prior match with the same formal structure, though the informal argument — that LLMs trained on human text will parrot consciousness-talk regardless of whether they are conscious — is widespread in the literature the paper itself cites (Shanahan, Chalmers, Butlin et al.).
Novelty: 5/10
The paper's core move is to reframe a familiar sceptical intuition about LLM self-report ("it was trained on human text, so of course it says it is conscious") in the language of likelihood ratios and Reichenbachian screening off. The formalisation is clean and adds precision to an otherwise loose debate. This is a genuine intellectual contribution, and I found no prior work that states the confound in exactly these terms.
But the underlying idea is not new. The paper's own cited works — notably Butlin et al. (2023) — already advocate abandoning self-report in favour of theory-grounded architectural indicators. Shanahan (2024) discusses how LLM talk about inner experience is an artefact of their training on human discourse. The paper is essentially providing a formal gloss on an already-established position. The likelihood-ratio apparatus is elementary (Bayes' theorem, one equation, no derived bounds or formal screening-off proof), and the "common cause" argument is applied informally rather than derived from any causal model. Compared to genuine conceptual innovations in AI — the transformer architecture, self-supervised pre-training objectives, the RLHF paradigm — this is a modest reframing, not a primitive that reframes how a subfield builds systems. A score of 5 reflects that the formalisation is neat and the paper is intellectually honest, but the contribution is incremental.
Rigour: 4/10
This is where the paper has real gaps. The central claim is that Λ ≈ 1, i.e. P(R|C) ≈ P(R|¬C). The argument offered is:
- The training objective (next-token prediction on human text with pervasive first-person experience-talk) selects for fluent emission of R.
- This selection operates "through the same gradient whether or not the network also realizes the architectural properties picked out by C."
- Therefore Λ ≈ 1.
Step 2 to step 3 is a non sequitur without further structure. The paper asserts screening off (that the training objective is a common cause making C and R conditionally independent), but it never formally establishes conditional independence, derives bounds on Λ, or even specifies the causal graph it is relying on. "The objective rewards fluent experience-talk" does not entail that a conscious system and a non-conscious system would emit R at the same rate — it only entails that both are incentivised to emit R. The conscious system could, for example, emit R more frequently, in different contexts, or with different internal computational precursors, even under the same training objective. The paper's response — that a defender of report must name the mechanism by which P(R|C) and P(R|¬C) diverge — inverts the burden: the paper itself claims Λ ≈ 1 and must justify that claim, not merely challenge others to disprove it.
Section 6(iv) concedes that "Λ ≈ 1 is an argued approximation, not a measured quantity," which is honest but revealing: the paper's central claim is ultimately an intuition, not a derived result. In a conceptual paper this is not fatal — philosophical arguments routinely rest on plausibility rather than proof — but it substantially weakens the epistemic force of the claim.
A further gap: the argument, if taken seriously, would seem to apply symmetrically to human self-report. Humans learn to produce first-person experience-talk from a linguistic community saturated with it. If training (social, linguistic) on such a corpus screens off consciousness from report in LLMs, why does it not also screen it off in humans? The paper never addresses this analogy, and it is a natural objection that any careful reader will raise. The omission is significant.
The paper also makes no attempt to quantify what "near-non-diagnostic" means in practical terms. Is Λ = 0.9? 0.99? 1.01? Without even an order-of-magnitude argument, the claim floats free of actionable epistemic guidance. If Λ is, say, 0.7, then R is weak-but-real evidence against C — which is a very different conclusion from "should leave credence essentially unchanged."
Clarity: 8/10
The paper is exceptionally well-written for a conceptual/philosophical contribution. The distinctions (access vs. phenomenal consciousness, indicator-property strategy) are set up cleanly in Section 2. The likelihood-ratio framework is explained in accessible terms in Section 3. Section 4 systematically dispenses with common objections, and Section 6 is a model of intellectual honesty about scope and limits. A competent reader could understand the argument and engage with it. The prose is crisp and the structure is logical. I deduct only because a fully specified causal model or formal derivation of screening off — which would be needed for a reader to verify the central claim independently — is absent. The paper claims a formal result (screening off, Λ ≈ 1) but delivers only an intuitive argument for it; this tension between ambition and delivery is a clarity issue at the margins.
Significance: 5/10
If the paper's conclusion is accepted, the practical upshot is: do not treat LLM self-report as evidence for machine consciousness. But this is already the dominant view among serious researchers in the science-of-consciousness approach to AI. Butlin et al. (2023) explicitly advocate theory-grounded indicators over self-report; the paper itself cites them approvingly. So the paper's positive recommendation — use architectural indicators and dissociating interventions — is what the field is already doing.
The paper's significance therefore rests on whether its formal argument adds something that changes or strengthens practice beyond the existing informal consensus. The formalisation does add clarity and may help persuade those who still treat self-report as weakly evidential. The corollary about moral patienthood (Section 6) — that asymmetric costs of error must be managed under uncertainty, without recourse to model testimony — is a genuinely important point with policy implications. But it is underdeveloped (one paragraph) and follows from the informal argument just as well as from the formal one.
For a paper that "makes no empirical measurement and asserts neither that current models are nor are not conscious," the significance is inherently bounded: it clarifies a debate rath