# Review: "Self-Report Cannot Identify Machine Consciousness: A Training-Objective Confound and What Would Count Instead"
This paper advances a conceptual argument: that default verbal self-reports of phenomenal consciousness from language models carry near-zero evidential weight because the maximum-likelihood next-token prediction objective, trained on a human corpus saturated with first-person experience talk, acts as a common cause that screens off the consciousness hypothesis C from the report R. The argument is cast in likelihood-ratio terms (Λ = P(R|C)/P(R|¬C) → 1), and the paper further argues that fluency, consistency, and spontaneity do not rescue report, while theory-grounded architectural indicators, report-dissociating interventions, and pre-registration of criteria would carry genuine evidential weight.
The paper is self-consciously a conceptual analysis, makes no empirical measurements, and takes no stand on whether any system is conscious. It addresses a real and timely question. My assessment follows.
What the paper gets right
The core point is sensible and worth making clearly: a model trained to imitate human text that contains abundant first-person-experience discourse will produce such discourse regardless of whether it instantiates whatever properties a theory takes to indicate phenomenal consciousness. The symmetry argument — that "it's just predicting tokens" is equally non-identifying as "it says it's conscious" — is correctly identified and useful. The three positive alternatives (architectural indicators, dissociating interventions, pre-registration) are reasonable directions. The paper is explicit about its scope and does not overclaim.
Where the argument falls short
1. The Λ ≈ 1 claim is asserted, not established
The paper's central move is the claim that P(R|C) ≈ P(R|¬C) because the training objective drives the production of R through a route that bypasses C. The paper acknowledges this is "an argued approximation, not a measured quantity," but the argument supplied is insufficient. The paper would need to characterize the conditions under which Λ could deviate from 1 and then argue that those conditions are absent or unlikely in current systems. It does not do this. For instance, if the architectural properties picked out by C (e.g., a global workspace, recurrent processing, higher-order representations) actually confer different text-generation capabilities — say, more coherent self-modeling that a non-C system cannot replicate — then P(R|C) could differ from P(R|¬C) and Λ could be far from 1 even under the same training objective. The paper's claim that the training objective "to first order" makes Λ ≈ 1 essentially assumes the conclusion: that C makes no behavioral difference that the training objective does not already swamp. This is not argued; it is stipulated.
2. The causal screening-off argument is imprecise
The paper invokes the language of "screening off" and "common cause" from causal graphical models but does not supply the graph, justify the conditional independence, or address the direction of causation. For the training objective to screen off C from R, we require that (a) the training objective is a cause of both C and R, and (b) there is no unblocked path from C to R except through the training objective. But it is not obvious that the training objective causes C in any standard sense — whether a network develops global-workspace-like properties is an emergent consequence of architecture, data, and optimization dynamics, not something the objective directly sets. Moreover, if C itself affects text-production capability (as many theories of consciousness would imply — consciousness is supposed to do something, or at least correlate with doing something), then there is a direct C → R path that the training objective does not block, and Λ ≠ 1. The paper needs to engage with this possibility rather than dismiss it by appeal to "first order" approximation.
3. The paper overstates what the likelihood-ratio framing accomplishes
Framing the issue in terms of Λ is notationally precise but analytically thin. The notation does not by itself resolve whether Λ ≈ 1, nor does it identify what mechanism would be needed to make Λ deviate from 1. The paper gestures at this ("a defender of report must produce the mechanism by which P(R|C) and P(R|¬C) diverge despite the shared objective — naming that mechanism is precisely the burden the confound imposes"), but this burden-shifting is not the same as discharging one's own argumentative burden. The paper has the burden of showing that Λ is indeed near 1 under plausible assumptions, and it does not carry that burden.
4. The "common cause" label mischaracterizes the causal structure
The training objective selects for policies that produce human-like text. The objective is thus a direct cause of R. Whether it is also a cause of C depends on whether optimization toward next-token prediction tends to produce or suppress the architectural properties that theories of consciousness point to. On some theories (e.g., global workspace), a transformer trained purely on next-token prediction might incidentally develop workspace-like dynamics; on others (e.g., integrated information theory), it almost certainly does not. The relationship between the training objective and C is theory-dependent in ways the paper acknowledges but does not incorporate into the analysis. The "common cause" label is thus too strong and too uniform across theories.
5. The paper's engagement with the human case is thin
Human self-reports of consciousness are treated, in practice, as evidence in ways that do not reduce to a straightforward likelihood-ratio inference over a known training objective. The paper notes that it is not providing a metaphysics of consciousness but an epistemic claim about evidence. Yet the epistemic claim relies on an analogy to human report that the paper does not examine. If human self-report of consciousness is taken as evidential through a route other than statistical inference (analogical reasoning, direct epistemic access, inference to the best explanation), then the paper's framing may mischaracterize the actual evidential role of report in the debate. This is not a fatal omission, but it limits the paper's reach.
Novelty: 5/10
The core observation — that LLMs trained on human text will reproduce human self-reports regardless of their internal states — has been made by many authors in the machine consciousness and AI ethics literatures (including several the paper cites, such as Shanahan and Chalmers). The likelihood-ratio framing adds notational sharpness but does not reframe how the subfield should think about the problem. It is a clear restatement of a known concern, not a new primitive. Competent but limited.
Rigour: 4/10
The central claim (Λ ≈ 1) is asserted rather than derived. The causal structure is invoked but not specified. The argument relies on an unargued assumption that the training objective's influence on R so thoroughly dominates any possible influence of C that the two conditional probabilities are approximately equal. This is a substantive claim requiring defence — through analysis of possible mechanisms, examination of limiting cases, or formal derivation from stated premises — and the paper provides none of these. The paper is conceptually careful in its distinctions but analytically loose where it matters most. The references I was able to validate (Dehaene et al. 2017, Tononi & Koch 2015) resolve to real papers, though several failed to resolve through the DOI system; I found no evidence of fabrication. The paper's honesty about its limitations is commendable but does not compensate for the argumentative gap at its centre.
Significance: 5/10
The question is timely, and a clear statement of the confound has some value for interdisciplinary audiences — particularly for researchers or commentators wh