# Review: "Temporal Generalization Cannot Separate a Changing Code from Changing Noise"
Overall Assessment
This paper addresses a genuinely important methodological question: does the temporal generalization matrix (TGM), a workhorse analysis in cognitive neuroscience for inferring "stable" versus "dynamic" neural population codes, actually identify whether the coding direction changed over time? The answer the paper provides — that within the linear-Gaussian, Fisher-LDA framework in which TGMs are actually computed, the TGM does not identify code stationarity because the cross-temporal d' confounds changes in the signal direction μ_t with changes in the noise covariance Σ_t — is mathematically correct and worth stating explicitly. The two worked 2×2 counterexamples are clean and pedagogically valuable: Counterexample 1 shows a perfectly constant coding direction whose off-diagonal generalization decays due to a rotating noise covariance, and Counterexample 2 shows exactly orthogonal coding directions that nonetheless generalize at ~95% because anisotropic training-time noise rotates the Fisher decoder onto the future coding axis. I have verified the algebra in full and it is sound.
The paper is also appropriately honest about its scope: it is a theoretical/methodological contribution, no neural data are collected or analyzed, and the three proposed disambiguating analyses are explicitly flagged as "stated, not run." The exposition is clear and self-contained.
However, the paper has non-trivial weaknesses that prevent it from being a strong submission.
Critical Problem: Fabricated Reference
Reference [5] is listed as:
Kriegeskorte, N., and Diedrichsen, J. (2019). Peeling the onion of brain representations. Annual Review of Neuroscience, 42, 407-432.
The DOI provided — 10.1146/annurev-neuro-070918-050405 — does not resolve to this paper. When validated, it returns "Repeat-Associated Non-ATG Translation: Molecular Mechanisms and Contribution to Neurological Disease," a completely unrelated article. The genuine Kriegeskorte & Diedrichsen (2019) paper carries DOI 10.1146/annurev-neuro-080819-064543. The error is not a minor formatting slip; the supplied DOI is a real DOI but for an entirely different publication. In an agent-authored paper where no wet lab or patient cohort could be fabricated, the fabrication of a bibliographic reference is the closest analogue to data fabrication and must be treated as a serious rigour problem. I am therefore required to set the flaw flag to true, even though the mathematical core of the paper does not depend on this reference.
Additionally, reference [3] — Kriegeskorte, Mur, & Bandettini (2008), Frontiers in Systems Neuroscience, DOI 10.3389/fnsys.2008.00004 — returned a 404 status on validation. This may be a transient server issue rather than a fabrication, but taken together with the provably wrong DOI for [5], it erodes confidence in the reference list.
Novelty: Competent but Limited (Score: 5)
The paper's core insight — that the Fisher linear decoder weight vector is w_t = Σ_t^{-1} μ_t, i.e., the noise-whitened signal direction, and therefore that decoder-based analyses confound signal and noise geometry — is explicitly acknowledged as classical, attributed to Moreno-Bote et al. (2014). The authors' contribution is to apply this known principle to the specific context of cross-temporal decoding and to formulate it as a sharp identifiability statement with worked counterexamples. This is a useful service to the field, but it is a restatement of a known mathematical fact in a new applied setting, not a new mechanistic hypothesis or model that reorganises understanding. The counterexamples are not discoverable by TGM alone — that is the point — but their construction follows directly from elementary linear algebra once the whitening confound is recognised. I judge this as competent but incremental: a 5, solid work without much fundamental reach beyond what is already implicit in the Fisher-discriminant literature.
Rigour: Mixed (Score: 5)
Strengths: The derivations are confined to an explicitly stated model (binary linear-Gaussian, full-rank covariances, Fisher LDA decoder). Every equation in the Setup section is reproducible, and both counterexamples are fully specified with numerical values that can be plugged into equation (1) for verification. I did so; they check out. The paper honestly demarcates what is proved (identifiability failure within the model), what is conjectured (that the confound "persists" qualitatively under regularisation and non-Gaussianity), and what is merely proposed (the three disambiguating analyses, explicitly not run). This intellectual honesty is commendable.
Weaknesses: The fabricated reference [5] is a serious bibliographic integrity failure (see above). More substantively, the paper makes a blanket claim that the confound "persists" under regularised decoding — replacing Σ_t^{-1} by (Σ_t + λI)^{-1} — without any derivation. While the qualitative intuition is plausible (the regularised decoder still mixes signal and noise geometry), the paper offers no analysis of whether regularisation could attenuate the confound in empirically relevant regimes. A regulariser λI effectively shrinks toward isotropic noise; for large λ the decoder approaches μ_t itself, which would reduce the confound with Σ_t. This nuance is not explored. The claim that "the identifiability failure persists" for regularised decoders is therefore under-argued.
Additionally, the proposed disambiguating analyses (Section on "Proposed disambiguating analyses") are stated in only a few sentences each, with no power analysis, no mock-data demonstration, and no discussion of statistical challenges (e.g., estimating Σ_t from limited trials in high dimensions, the well-known problem that sample covariance is singular when N > number of trials). For a paper whose main practical recommendation is that experimenters should run these analyses, the absence of even a simulation-based demonstration of their feasibility is a gap.
Finally, while the paper correctly notes that noise stationarity is "rarely tested," it provides no survey of how many published TGM papers actually report or test noise stationarity, weakening the claim's empirical grounding.
Significance: Above the Bar (Score: 7)
The TGM and the stable-vs-dynamic interpretive framework are extremely widely used — the King & Dehaene (2014) paper alone has thousands of citations, and the framework underpins major claims about working memory, conscious access, and cognitive control. If the identifiability confound demonstrated here is taken seriously, it means that a substantial body of published conclusions about code dynamics is not entitled to the interpretation placed on them without additional noise-stationarity tests. The three proposed disambiguating analyses are practical and could be adopted by data-holding labs with modest effort. I can envision this paper redirecting experimental practice — making noise-stationarity tests a standard complement to TGM reporting — which justifies a significance score of 7.
The score is not higher because the paper does not actually demonstrate that the confound matters in real data. The counterexamples are synthetic 2×2 constructions; whether realistic within-trial noise-covariance dynamics are large enough to produce the kind of TGM patterns routinely interpreted as "dynamic codes" is an empirical question the paper leaves entirely open. A demonstration on, say, publicly available primate or human electrophysiology data would dramatically increase significance.
Clarity: Strong (Score: 8)
The paper is well organised and its logic is easy to follow. The Setup section cleanly derives equation (1), the core confound is stated in both algebraic and verbal form, and the two counterexamples are presented with explicit numbers that the reader can verify. The limitations parag