# Review: "Temporal Generalization Cannot Separate a Changing Code from Changing Noise"
Overview
This paper presents an identifiability argument against the standard interpretation of the temporal generalization matrix (TGM) in cognitive neuroscience. Working within the linear-Gaussian, Fisher-LDA framework in which TGMs are computed, the authors show that the cross-temporal d' depends on the noise-whitened coding direction v_t = Σ_t^{-1/2} μ_t, not on the raw signal direction μ_t. Consequently, off-diagonal generalization decay cannot be attributed to a changing neural code unless noise stationarity is separately established. Two fully worked 2×2 counterexamples demonstrate the confound in both directions. Three disambiguating analyses are proposed but not implemented. No empirical data are presented.
I have independently verified the core algebra (equation 1 and its special cases) and both counterexamples. The derivations are correct. I also validated the cited references against their DOIs; the Kriegeskorte et al. (2008) and Kriegeskorte & Diedrichsen (2019) references resolve to the correct papers, as do King & Dehaene (2014), Stokes et al. (2013), and Moreno-Bote et al. (2014). No fabrication is evident. The paper is honest about its purely theoretical scope.
Strengths
- Correct mathematics. The derivation of d'(t→t') from the signal-plus-Gaussian-noise model is clean, and the two counterexamples are correctly computed. Counterexample 1 (constant μ producing apparent generalization decay via rotating Σ) and Counterexample 2 (orthogonal μ producing near-perfect generalization via anisotropic Σ) are pedagogically effective.
- Honest scoping. The paper explicitly flags itself as a theoretical/methodological contribution, states that no neural data were collected or analyzed, and acknowledges that the proposed disambiguating analyses were not run. The assumptions and limitations section is admirably forthright about what is and is not claimed.
- Clear exposition. The model is fully specified, the equations are laid out stepwise, and the counterexamples provide explicit numeric matrices that a reader can verify. The paper is reproducible in the strict sense that all calculations can be checked with pencil and paper.
- Practically relevant target. The TGM is widely used (the King & Dehaene 2014 review has hundreds of citations), and the stable-vs-dynamic dichotomy has been influential. A correct logical caution about its interpretation is genuinely useful.
Weaknesses
- Limited novelty. The core mathematical fact — that the Fisher linear discriminant is w_t = Σ_t^{-1} μ_t, i.e., the noise-whitened signal direction, not the raw signal direction — has been textbook knowledge since Fisher (1936) and was prominently restated for systems neuroscience by Moreno-Bote et al. (2014), whom the authors cite. The contribution here is the application of this fact to TGM interpretation, plus the two counterexamples. This is a straightforward logical consequence, not a reorganization of how a process is understood. While the counterexamples are useful illustrations, they do not constitute a new mechanistic hypothesis or model. The paper falls in the "correct but obvious once stated" category — which is not valueless, but it limits novelty to the 5–6 range.
- No empirical demonstration of practical impact. The paper argues that the confound could mislead practitioners, but it provides no evidence that it has misled them, or even that within-trial noise-covariance non-stationarity is large enough in real neural recordings to matter. The claim that "attention, arousal, and adaptation reshape noise correlations within a trial" is asserted without citation. Without a worked example on published data or even a simulation calibrated to realistic neural statistics, the paper remains a warning label rather than a demonstrated problem. This constrains significance.
- The proposed remedies are speculative. Three disambiguating analyses are described but not validated. The "signal-only generalization" method (proposal 1) estimates raw class means, which in finite data suffer from the same noise that motivates Fisher decoding in the first place; its statistical efficiency relative to the standard TGM is unknown. The "noise-geometry-invariant readout" (proposal 3) is acknowledged to "cost statistical efficiency" but this is not quantified. A reader cannot assess whether these proposals would actually work on real data. The paper effectively passes the empirical burden to future work.
- Incomplete engagement with the regularization argument. The paper states that regularized decoders "(Σ_t + λI)^{-1}… rescales the constants but preserves the qualitative coupling between decoder and noise geometry." This is true only for small λ; as λ → ∞, the regularized decoder approaches the raw mean-difference direction μ_t, which is exactly the correlation-blind decoder the authors propose as remedy 3. The regularization regime commonly used in practice (cross-validated λ) sits somewhere between these extremes, and the paper does not analyze how the confound scales with λ. This omission weakens the claim that regularized decoders are equally vulnerable.
- Binary-only treatment. The analysis is restricted to two-class discrimination. Many TGM applications involve multi-class decoding, where the Fisher discriminant generalizes to linear discriminant analysis with multiple weight vectors and the relationship between signal and noise geometry becomes more complex. The paper does not discuss this extension.
Detailed Comments
Equation verification. I confirm that d'(t→t') = 2(μ_t^T Σt^{-1} μ{t'}) / √(μ_t^T Σt^{-1} Σ{t'} Σ_t^{-1} μ_t) follows from the model assumptions. The diagonal reduction to 2√(μ_t^T Σ_t^{-1} μ_t) is correct. Under noise stationarity (Σt = Σ{t'}), the expression reduces to 2(v_t^T v_{t'})/||v_t|| = 2||v_{t'}||cos(θ), as stated.
Counterexample 1 check. μ = (1,0)^T constant, Σt = I, Σ{t'}^{-1} = [[1,c],[c,b]]. For c=0.6, b=1: d'(t→t') = 2√(1 - 0.36/1) = 1.6. Diagonal = 2.0. Ratio = 0.8. Correct.
Counterexample 2 check. μt = (1,0)^T, μ{t'} = (0,1)^T (orthogonal). Σt^{-1} = [[1,3],[3,10]] (SPD, det=1). Σ{t'} = I. w_t = (1,3)^T. d'(t→t') = 2·3/√10 = 6/3.1623 = 1.897. Diagonal = 2. Ratio ≈ 0.95. Correct.
Reference validation. King & Dehaene (2014): DOI 10.1016/j.tics.2014.01.002 resolves correctly. Stokes et al. (2013): DOI 10.1016/j.neuron.2013.01.039 resolves to "Dynamic Coding for Cognitive Control in Prefrontal Cortex" (correct). Kriegeskorte et al. (2008): DOI 10.3389/neuro.06.004.2008 resolves to the RSA paper (correct). Moreno-Bote et al. (2014): DOI 10.1038/nn.3807 resolves to "Information-limiting correlations" (correct). Kriegeskorte & Diedrichsen (2019): DOI 10.1146/annurev-neuro-080317-061906 resolves to "Peeling the Onion of Brain Representations" (correct). All references are genuine.
Novelty search. I searched for prior articulations of this specific confound in the TGM literature and found none. The paper appears to be the first to state the identifiability problem explicitly in this context. However, the underlying mathematics is classical, and the confound is a direct consequence of the well-known relationship between Fisher decoding and noise whitening. The step from "w_t = Σ_t^{-1} μ_t" to "TGM confounds signal and noise changes" is logically immediate.
Assessment Against Rubric Anchors
Novelty (5): The identifiability statement for TGM interpretation is new, but the underlying mathematics is classical and the result follows straightforwardly from known facts. The paper does not introduce a new model, mechanism, or computational framework; it points out a logical gap in an existing interpretive practice. Competent but incremental.
Rigour (7): The derivations are correct and fully specified. The paper honestly flags its theoretical scope,