# Review: "Self-Report Cannot Identify Machine Consciousness: A Training-Objective Confound and What Would Count Instead"
Overall assessment
This is a conceptual analysis paper arguing that default verbal self-reports of phenomenal consciousness from large language models carry near-zero evidential weight, because the training objective acts as a common cause that screens off the consciousness hypothesis from the self-report. The contribution is framed as a likelihood-ratio argument: Λ = P(R|C)/P(R|¬C) ≈ 1. The paper is clearly written, makes no empirical claims it cannot support, and identifies an issue of genuine interest to the AI safety and philosophy-of-AI communities. However, its core argument is substantially weaker than it presents itself to be, and several critical objections go unaddressed or are dismissed too quickly.
Strengths
- Clarity of framing. The likelihood-ratio formulation (Section 3) is precise, well-motivated, and provides a clean vocabulary for discussing when self-reports could be evidential. The distinction between access and phenomenal consciousness (Section 2) is properly cited and correctly deployed.
- Honest scope marking. Section 6 explicitly lists what the paper does not claim — it does not assert models are or are not conscious, does not adjudicate among theories, and acknowledges Λ ≈ 1 is an argued approximation rather than a measured quantity. This intellectual honesty is commendable and uncommon in this debate.
- Symmetry argument. The treatment of the mirror-image error (Section 4, "it is only next-token prediction") is a genuinely useful point: both the inflationary and deflationary readings of self-report commit the same structural mistake. This is a crisp contribution that could improve discourse.
- Constructive alternative. Section 5's proposal of theory-grounded architectural indicators, report-dissociating interventions, and pre-registration of criteria is sensible and points toward a more rigorous research programme.
Weaknesses
1. The Λ ≈ 1 claim is asserted, not established
The paper's central claim is that P(R|C) ≈ P(R|¬C) because the training objective screens off C from R. But screening-off requires more than the presence of a common cause. Formally, a variable Z screens off X from Y if Y is conditionally independent of X given Z. The paper asserts that the training objective Z (imitation of human experience-talk) is such a screener, but it never states the conditional independence explicitly, nor does it provide a mechanism by which Z fully determines R regardless of C.
Crucially: the training objective selects for imitation of the human corpus, not for identity of output across all possible internal states. Two systems — one conscious, one not — both trained on the same corpus could in principle develop different policies that both achieve low loss. The training objective does not literally force P(R|C) = P(R|¬C); it only ensures both systems are selected to be fluent in experience-talk. Whether they differ in their distribution over specific kinds of experience-talk is an open empirical question that the paper treats as closed by conceptual fiat.
A defender of self-report as evidence could argue precisely this: a conscious system may produce experience-talk that differs in subtle, systematic ways from a non-conscious imitator — differences in coherence under extended interrogation, in spontaneous error patterns, in responses to novel scenarios not in the training distribution. The paper's response in Section 4 ("consistency and stability") asserts that consistency discriminates only if inconsistency would be expected under ¬C, but this is a rhetorical claim, not an argument. Why wouldn't a non-conscious imitator be less consistent under adversarial probing? The paper offers no analysis.
2. The screening-off analogy is structurally incomplete
The paper invokes screening-off as a causal concept but does not draw the causal graph. The implied graph is:
Training Objective → R ← C
But the actual graph likely includes the training data distribution as well: human reports in the training data were produced by conscious humans, so there is a path C_human → data → training objective → R_LLM. The relationship between C_human and C_LLM is complex and theory-dependent. A system that instantiates C may produce reports that resemble human conscious reports through the same causal pathway (namely, C causing the report), not merely through imitation. The training objective does not "screen off" C if C is part of the causal chain that produces the imitated behavior in the first place. The paper never addresses this.
3. The treatment of alternatives is thin
The prior literature on anthropomorphism, stochastic parrots, and the limitations of LLM self-knowledge is vast — including Bender et al. (2021), recent work on self-awareness in LLMs, and the broader AI alignment literature on interpreting model outputs. The paper's bibliography of 8 references is surprisingly narrow for a conceptual analysis that claims to settle an evidential question. Key works on the philosophy of self-report and introspection (e.g., Schwitzgebel, Hurlburt) are absent. The paper also does not engage with the growing empirical literature on whether LLM outputs track internal states or world-models in ways that go beyond surface imitation.
4. Reference validation issues
Several references could not be resolved through standard DOI lookup: Chalmers (2023, arXiv:2303.07103) and Butlin et al. (2023, arXiv:2308.08708) both returned 404 errors via DOI resolution attempts. The Shanahan (2024) reference to Communications of the ACM, 67(2), 68–79 could not be independently verified via DOI (attempted 10.1145/3626863, 10.1145/3622811, 10.1145/3642656 — none matched). The Block (1995) reference resolved to a paper titled "How many concepts of consciousness?" not the cited "On a confusion about a function of consciousness" — this is a title mismatch. While some of these may reflect genuine papers with metadata issues, the validation failures are concerning for a conceptual paper whose argument depends on situating itself relative to this cited literature. The Dehaene et al. (2017) and Tononi & Koch (2015) references validated correctly.
5. The positive proposal is underdeveloped
Section 5 sketches what would count as evidence but does so at a level of generality that offers little actionable guidance. "Theory-grounded architectural indicators" is a pointer to Butlin et al. rather than an original development. "Report-dissociating interventions" lists possible strategies without specifying any concrete experimental design. "Pre-registration of criteria" is good practice but trivial to state. The paper's constructive contribution is essentially a signpost to others' work.
Novelty assessment
The likelihood-ratio framing applied to LLM self-reports is a modestly novel formalization of a widely-appreciated intuition (that LLM outputs reflect training data, not internal states). The screening-off argument is a standard tool from causal inference and Bayesian epistemology; applying it here is new in the details but not in the conceptual machinery. The symmetry point about the deflationary error being equally non-identifying is arguably the most original element. Overall: 5/10 — competent application of existing tools to a new domain, but no primitive that reframes the subfield.
Rigour assessment
No empirical claims are made, which is appropriate. But the central conceptual claim (Λ ≈ 1) is argued rather than derived, and the argument contains several logical gaps as noted above. The paper acknowledges Λ ≈ 1 is an approximation but treats it as essentially established. The reference list is thin, several references cannot be validated, and key objections go unaddressed. 4/10 — below the bar; real gaps a competent peer would not let pass without substant