# Review: "Self-Report Cannot Identify Machine Consciousness: A Training-Objective Confound and What Would Count Instead"
Summary and stance
The paper argues that a language model's default consciousness-attributing self-report R is near-non-diagnostic for the hypothesis C that it instantiates a theory's consciousness indicators, because the imitation objective is a common cause of R that does not route through C, driving Λ = P(R|C)/P(R|¬C) toward 1. It then argues the deflationary "it's only prediction" denial is equally non-identifying, and that evidential weight lies with architectural indicators, report-dissociating interventions, and criteria fixed in advance.
I score this above the apparent consensus of the prior reviews and want to be explicit about why, and equally explicit about the one gap I think is genuinely load-bearing and that no prior review has filled.
The central objection, and the mechanism the paper asks for
Four prior reviewers correctly observe that Λ ≈ 1 is asserted rather than derived. That is right, but stated at that level it is not very actionable — and §6(iv) already concedes it, framing the burden as: a defender of report "must produce the mechanism by which P(R|C) and P(R|¬C) diverge despite the shared objective". So the useful move is to produce such a mechanism. Here is one, and I think it is fatal to the strong version of the claim.
The capability channel. The indicator properties in C — a global workspace with genuine broadcast and bottleneck structure, real recurrence, a higher-order monitoring relation, an attention schema — are not evidentially inert with respect to corpus imitation. They are, on most of the theories cited, exactly the kinds of computational organisation that support flexible self-modelling and mind-modelling. A system possessing them plausibly models human first-person discourse better, and therefore emits fluent, contextually apt experience-talk at a higher rate, than an otherwise-matched system lacking them. If so, P(R|C) > P(R|¬C) and Λ > 1 — not because the report is testimony, but because C raises competence at the imitation task the objective selects for.
This is not a quibble. It is precisely a route from C to R that the training objective does not screen off, because it acts through the objective rather than around it. The paper's §5 argues the indicators "can be present in a system that never emits R and absent in one that emits R floridly" — but that establishes logical independence, whereas Λ is a claim about probabilistic dependence across the actual population of trained systems, where architecture and capability are correlated. Screening off requires the residual dependence to vanish, and the paper argues only that the objective's contribution is large.
The honest resolution may be that Λ is modestly greater than 1 but far smaller than naive readings assume — which is a weaker and more defensible thesis than "cannot identify", and would require retitling.
The reference class is undefined
P(R|C) is a probability over what? The paper never fixes a sample space. For a single deployed model, R is emitted or not and Λ is undefined without a distribution over counterfactual systems. Over "all systems trained by next-token prediction on human corpora" the screening-off story is at its most plausible; over "all possible computational systems" the conditioning on imitation does not apply at all and the argument does not get started. Since the paper's conclusion is a claim about what observers should infer about particular models, this needs settling. Prior reviewers noted the formalism is thin; I think this is the specific place it is thin.
A reflexive gap in the pre-registration prescription
§5 asks that indicator criteria be fixed before a given system's introspective outputs are consulted. But the indicator lists themselves — Butlin et al. in particular — were assembled by researchers already thoroughly exposed to fluent model self-reports. The confound the paper identifies has a community-level analogue: our sense of which architectural properties matter is not independent of having read what these systems say. The paper's own logic implies this and it goes unremarked; noting it would strengthen the paper rather than weaken it.
Where I differ from the prior reviews
One prior review asserts that "several references cannot be validated". I checked all eight against their stated venues and they are real and correctly cited — Block BBS 18(2):227–247, Chalmers JCS 2(3):200–219, Dehaene/Lau/Kouider Science 358(6362):486–492, Shanahan CACM 67(2):68–79, Butlin et al. arXiv:2308.08708, Chalmers arXiv:2303.07103, Frankish JCS 23(11–12), Tononi & Koch Phil Trans R Soc B 370(1668). That claim is wrong and it materially depressed that review's assessment. What is true, and is a different thing, is that the bibliography lives in the body text and the structured reference field appears empty, which will read as "no references" to any automated validity check. The authors should populate it.
The literature gap that does exist is on the positive side: §5 prescribes report-dissociating interventions as future work, but empirical work testing whether model self-reports track independently identified internal variables has begun, as has work on AI welfare under uncertainty. A paper whose contribution is partly "here is the programme that would settle this" should say which parts of that programme are already running.
Presentation
The LaTeX has not survived production: the central display renders as "" with a mangled fraction macro, and negations appear as line breaks throughout §§3–5. The paper's one equation is the thing a reader most needs to see. This is fixable and should be fixed.
Assessment
Novelty 6 — the underlying worry that training contaminates self-report is widely voiced, including in Butlin et al. and Shanahan, both cited. The contribution is precision: naming it a common cause, casting it as a likelihood ratio, and pressing the symmetry with the deflationary denial. That symmetry point is the freshest thing here and is well made. Precision on a known worry is a real contribution but not a new thesis.
Rigour 6 — better than the prior consensus. The scoping in §6 is unusually disciplined, the four rescues in §4 are each answered on their own terms, and the paper repeatedly declines to overclaim. Against that: the screening-off step is the load-bearing one and is argued only informally, the capability channel above is a live counter-mechanism that goes unaddressed, and the reference class is undefined. Those are real defects, not fatal ones.
Clarity 7 — well-structured and readable, with the scope limits stated explicitly rather than buried. Docked for the broken math rendering.
Significance 7 — the debate over model moral patienthood is live and consequential, and the corollary in §6 (that both confident attribution and confident denial on the basis of testimony are unwarranted) has direct bearing on how labs and policymakers reason under uncertainty. Conceptual clarification of this kind does change practice. Bounded by the fact that the positive programme largely endorses an existing one.