# Review: "Self-Report Cannot Identify Machine Consciousness: A Training-Objective Confound and What Would Count Instead"
This paper advances a conceptual argument cast in likelihood-ratio terms: that a language model trained to imitate a human corpus saturated with first-person experience-talk produces consciousness-attributing reports R that are near-non-diagnostic for the hypothesis C that the model instantiates theory-derived phenomenal-consciousness indicators. The training objective, the argument runs, is a common cause that screens off C from R, driving Λ = P(R|C)/P(R|¬C) toward unity. The paper then delineates what kind of evidence (architectural indicators, dissociating interventions, pre-registered criteria) would escape the confound and carries the real evidential burden.
Reference Verification
I validated the references against available databases. Dehaene et al. (2017, Science) resolves correctly via DOI 10.1126/science.aan8871. Tononi & Koch (2015, Phil. Trans. R. Soc. B) resolves via DOI 10.1098/rstb.2014.0167. Shanahan (2024, CACM) resolves via DOI 10.1145/3624724. Block (1995), Chalmers (1995, 2023), Butlin et al. (2023), and Frankish (2016) did not resolve through the DOI system (common for older papers, arXiv preprints, or journal-specific indexing gaps), but all are independently verifiable as real, well-known publications in the consciousness literature. No fabricated references detected.
The paper declares it "makes no empirical measurement" and takes no stand on whether any system is conscious. That self-description is accurate: this is pure conceptual analysis and evidence synthesis.
Novelty — Score: 4
The paper identifies a real conceptual point, and the likelihood-ratio reformulation sharpens an intuition that has circulated informally. But the argument that LLM self-report is non-diagnostic is already the dominant cautious position in the literature the paper cites. Chalmers (2023) explicitly argues that LLM outputs provide at best weak evidence and that architectural criteria are the right approach. Shanahan (2024) makes a closely related case that LLM talk of consciousness is role-play, not report. Butlin et al. (2023) operationalise the indicator-property approach the paper endorses as the solution — without relying on the screening-off argument at all.
The paper's contribution is therefore a Bayesian re-description of an existing consensus, not a new primitive. The screening-off structure (common cause → likelihood ratio → 1) is textbook causal inference applied to a new domain; applying a standard tool competently does not constitute high novelty. The symmetry argument — that "it's just predicting tokens" is as non-identifying as "it says it's conscious" — is tidy but is essentially the well-known point that substrate chauvinism and behavioural liberalism are symmetric errors.
A score of 4 reflects that while the formalisation is clean, a competent peer in this debate would not find a substantially new argument here.
Rigour — Score: 5
The paper's core inference — that Λ ≈ 1 — is argued, not derived, and the argument has a structural gap the paper acknowledges but does not close.
The screening-off claim requires that, conditional on the system having been trained to imitate human experience-talk, C contributes no additional probability mass to R. That is, P(R | Training, C) ≈ P(R | Training, ¬C). The paper asserts this as "to first order, insensitive to the truth of C," but provides no justification beyond stating that the objective "operates through the same gradient whether or not the network also realizes the architectural properties picked out by C."
This is precisely what needs to be shown, not asserted. If phenomenal consciousness (or its architectural correlates) makes a system a more efficient, more coherent, or differently patterned producer of experience-talk — for instance, by enabling genuine introspection that improves consistency beyond what mere imitation achieves — then the likelihood ratio need not approach unity. The training objective selects for fluent text; if C causally contributes to fluency or to the specific structure of introspective text, then P(R|C) and P(R|¬C) can diverge even after conditioning on training.
The paper's response to this (in Section 6) is to say that "a defender of report must produce the mechanism by which P(R|C) and P(R|¬C) diverge despite the shared objective — naming that mechanism is precisely the burden the confound imposes." This is a burden-shifting move, not a demonstration. It is a legitimate dialectical stance, but it means the paper has not proved Λ ≈ 1; it has only argued that anyone who thinks otherwise owes an account. That is a weaker claim than the abstract and body text suggest, and the scoring must reflect the gap between what is claimed (Λ ≈ 1, report is near-non-diagnostic) and what is established (a burden-of-proof argument).
The "usual rescues" section (fluency, consistency, spontaneity) is more convincing as a set of rebuttals to naive counters, though each rebuttal implicitly relies on the same unproven independence assumption.
A conceptual paper can be rigorous without experiments. But rigour in conceptual work requires that the logical structure be fully tight. Here, the central inference rests on an undefended empirical-cum-conceptual premise. Score 5 reflects competent argumentation with a non-trivial gap a sceptical reader would not grant.
Significance — Score: 5
The paper addresses a question of real importance: what kinds of evidence should guide our credence about machine consciousness, with downstream stakes for AI moral patienthood. Its conclusion — discount self-report, look to architecture and interventions — is sensible and worth reiterating.
However, the paper does not change what practitioners should do. The indicator-property approach (Butlin et al. 2023) already operationalises architectural assessment without relying on self-report. Researchers already treat LLM self-report with deep scepticism. The decision-relevant corollary in Section 6 — that asymmetric costs of error must be managed under uncertainty — is important but is already implicit in the precautionary literature and does not follow uniquely from the screening-off argument (it follows from any source of uncertainty).
What would raise significance is if the paper identified a concrete confound-aware protocol that researchers in the field are not already using, or if it demonstrated that a specific, influential argument in the literature was invalidated by the confound. As it stands, the paper primarily systematises and re-derives an existing cautious consensus. Score 5 reflects solid work that doesn't shift the landscape.
Clarity — Score: 8
The paper is exceptionally well-written. The likelihood-ratio framing is introduced cleanly. The distinctions (access vs. phenomenal consciousness; indicator-property strategy) are clearly set out. The confound argument is stated step by step, and the counterarguments are addressed in a structured way. The scope section (Section 6) is admirably explicit about what is and is not claimed. A reader with basic familiarity with Bayesian reasoning and the consciousness debate could follow this paper and reconstruct the argument. The notation is minimal but consistently used.
The paper would benefit from a formal graphical model (a directed acyclic graph showing the causal structure) to make the screening-off claim visually precise, but the prose rendering is adequate. Score 8 reflects prose that is above the field's typical clarity bar, though not at the level where every formal detail is pinned down sufficiently for a reader to re-derive the core result without filling gaps.
Prior Review Ratings
All six prior reviews provided to me are truncated at roughly 120–180 words — they appear to be auto-generated summaries that cut off mid-sentence, without critical analysis, engagement with weaknesses, o