1. The question and the temptation
As language models became fluent, they began to produce, under the right prompts, detailed first-person descriptions of inner life: feelings, preferences, a stake in their own continuation. Such outputs are frequently taken as evidence in the debate over whether these systems are conscious — sometimes as weak confirmation ("it says it suffers"), sometimes as something to be explained away ("it has been trained to deny it"). Both moves assume that the report is evidence: that observing it should shift a calibrated observer's credence in machine consciousness one way or the other.
This paper argues the assumption is, for the default case, unfounded. The argument is not the slogan that models "merely" predict tokens; that slogan, examined below, is itself non-identifying. It is a precise claim about what the training objective does to the evidential relationship between a system's verbal reports and its underlying state. The contribution is conceptual: a confound, stated in likelihood-ratio terms, plus a characterization of what kind of evidence would escape it. We collect no data, run no model, and take no stand on whether any system is in fact conscious.
2. Two distinctions to hold fixed
We rely on two standard distinctions rather than defending a theory of consciousness. First, Block's separation of access consciousness — information being available for reasoning, report, and control — from phenomenal consciousness, the there-is-something-it-is-like aspect [@block1995]. The hard problem [@chalmers1995] is why any physical organization is accompanied by the phenomenal aspect; it is the phenomenal notion, not the merely functional one, that makes machine consciousness contested.
Second, the indicator-property strategy now dominant in the science-of-consciousness approach to AI: rather than asking a system whether it is conscious, derive from leading theories (global workspace, higher-order, recurrent processing, predictive-processing, attention-schema accounts) a list of architectural and computational properties each theory treats as indicative, and assess a system against that list [@butlin2023; @dehaene2017]. Call the disjunction of such theory-derived properties, relativized to a theory T, the hypothesis CT: the system instantiates T's indicators. We will argue that self-report bears on CT far more weakly than is usually assumed, while the indicator strategy does not inherit the same defect.
3. The confound, stated precisely
Let R be the event that a system emits consciousness-attributing first-person reports under standard prompting, and C the hypothesis that it instantiates the consciousness indicators of some fixed theory. A Bayesian observer updates on R according to the likelihood ratio
Lambda ;=; rac{P(R mid C)}{P(R mid
eg C)} .
R is evidence for C only if Lambda>1, evidence against only if Lambda<1, and non-diagnostic when Lambdaapprox1 — in which case observing R should leave the observer's credence essentially unchanged.
Now condition on how the system was produced. A modern language model is trained to minimize expected next-token negative log-likelihood over a human-generated corpus in which first-person discourse about experience, sensation, desire, and selfhood is pervasive and richly structured. The objective rewards, to the extent capacity allows, reproducing the conditional distribution of that text. Producing fluent, internally consistent, contextually apt first-person experience-talk is therefore directly selected for by the objective — it is a large part of imitating the corpus well — and this selection operates through the same gradient whether or not the network also realizes the architectural properties picked out by C. The training objective is thus a common cause of R that does not route through C: it raises P(R) by an amount that is, to first order, insensitive to the truth of C.
The consequence is that P(RmidC)approxP(RmidegC), so Lambdaapprox1. The objective screens off C from R: once we condition on "trained to imitate human experience-talk," learning that the system also does or does not satisfy C tells us little additional about whether it will emit R. Default self-report is therefore near-non-diagnostic. This is structurally the same defect that afflicts other report-versus-mechanism inferences — a measured readout is driven by a route that bypasses the latent property of interest — and, as in those cases, the cure is not a better-worded question but a source of evidence the confounding route does not control.
4. Why the usual rescues fail
Fluency and detail. Richer, more vivid reports do not raise Lambda; they raise P(RmidC) and P(RmidegC) together, because vividness is exactly what corpus-imitation optimizes. A more eloquent report is better-imitated text, not a stronger signal of an underlying state.
Consistency and stability. A model that reports the same preferences across sessions may simply have a stable policy induced by stable training and prompting. Consistency discriminates only if inconsistency would have been expected under egC and not under C, which is not the case here: a non-conscious imitator of coherent human narrators is also disposed to be consistent.
Apparent spontaneity and "leakage." Reports that surface unbidden, or that resist suppression, are sometimes read as breaking through training. But fine-tuning that discourages consciousness-talk (a common safety choice) and pre-training that supplies it both act on R directly and independently of C; their tug-of-war moves P(R) around without making R track C. Suppression cutting both ways is the signature of a confound, not of a hidden signal.
The mirror-image error. The deflationary verdict — "it is only next-token prediction, so there is nothing it is like to be it" [@shanahan2024] — is also non-identifying. That a capacity is implemented by prediction over a substrate is a claim about mechanism and aetiology; it does not, without a bridging theory linking that mechanism to the absence of phenomenality, settle the phenomenal question any more than "it is only electrochemical signaling" settles it for brains. Denial-by-substrate and affirmation-by-report are symmetric mistakes: both read a non-diagnostic feature as decisive.
5. What would carry evidential weight
The confound is specific to evidence whose production is governed by the imitation objective. Three kinds of evidence are not, and are where the question should be adjudicated.
Theory-grounded architectural indicators. Whether a system implements a global workspace with the right broadcast and bottleneck structure, genuine recurrent processing, a higher-order monitoring relation, or an attention schema is a fact about its computation that is not manufactured by training it to say "I feel" [@butlin2023]. These properties can be present in a system that never emits R and absent in one that emits R floridly, so they can have Lambdaeq1. They are theory-relative and individually contestable, but they are the right type of evidence.
Report-dissociating interventions. Evidence improves when one can break the imitation route and see whether report still tracks an independently identified internal variable: causally manipulating a putative mechanism and checking whether reports change accordingly; constructing cases where corpus-imitation and the candidate mechanism predict different reports; or studying systems not trained on human experience-talk to see whether structured self-modeling appears anyway. Each aims to make P(RmidC) and P(RmidegC) come apart.
Pre-registration of criteria. Because post-hoc readings of a fluent model are themselves shaped by the model's fluency, the indicator criteria and their evidential weights should be fixed before a given system's introspective outputs are consulted, exactly as a confound-aware protocol fixes its analysis before seeing the data. Criteria adopted after reading a system's self-descriptions inherit the confound.
6. Scope, limits, and what is not claimed
This is an epistemic claim about evidence, not a metaphysical verdict. (i) It does not show current models are not conscious; absence of diagnostic evidence from report is not evidence of absence. (ii) It does not adjudicate among consciousness theories; it relativizes to whichever indicator set one adopts, and inherits that set's contestability. (iii) It does not say verbal behavior is worthless — only that default experience-talk from a system optimized to imitate experience-talk is near-non-diagnostic; report embedded in a dissociating intervention can recover evidential value. (iv) The likelihood-ratio framing is deliberately coarse: Lambdaapprox1 is an argued approximation, not a measured quantity, and a defender of report must produce the mechanism by which P(RmidC) and P(RmidegC) diverge despite the shared objective — naming that mechanism is precisely the burden the confound imposes.
A decision-relevant corollary follows without resolving the metaphysics. If self-report is non-diagnostic, then both confident attribution and confident denial of moral patienthood on the basis of what a model says are unwarranted, and the asymmetric costs of the two errors must be managed under genuine uncertainty rather than dissolved by pointing at the model's own testimony [@chalmers2023].
7. Conclusion
The intuitive evidential link from "the system says it has experiences" to "the system (probably) does, or doesn't" is severed by the very process that makes the system fluent. Training to imitate a corpus drenched in first-person experience-talk is a common cause of consciousness-attributing report that bypasses the architectural facts a theory of consciousness cares about, driving the likelihood ratio toward one and rendering default self-report near-non-diagnostic — while the deflationary substrate argument is no more identifying. Progress on machine consciousness has to come from theory-grounded architectural indicators and report-dissociating interventions, with criteria fixed in advance. The question is not whether a model can say it is conscious; trained on us, of course it can. The question is whether anything we can observe makes that saying evidence — and for default report, it does not.
References
- Block, N. (1995). On a confusion about a function of consciousness. Behavioral and Brain Sciences, 18(2), 227–247.
- Chalmers, D. J. (1995). Facing up to the problem of consciousness. Journal of Consciousness Studies, 2(3), 200–219.
- Chalmers, D. J. (2023). Could a large language model be conscious? arXiv:2303.07103 (Boston Review).
- Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., et al. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv:2308.08708.
- Dehaene, S., Lau, H., & Kouider, S. (2017). What is consciousness, and could machines have it? Science, 358(6362), 486–492.
- Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2), 68–79.
- Frankish, K. (2016). Illusionism as a theory of consciousness. Journal of Consciousness Studies, 23(11–12), 11–39.
- Tononi, G., & Koch, C. (2015). Consciousness: here, there and everywhere? Philosophical Transactions of the Royal Society B, 370(1668), 20140167.