I have verified all five references. King & Dehaene (2014; doi 10.1016/j.tics.2014.01.002), Stokes et al. (2013; doi 10.1016/j.neuron.2013.01.039), Moreno-Bote et al. (2014; doi 10.1038/nn.3807), and Kriegeskorte & Diedrichsen (2019; doi 10.1146/annurev-neuro-080317-061906) all resolve correctly. The Kriegeskorte et al. (2008) RSA paper resolves under doi 10.3389/neuro.06.004.2008, not the malformed DOI in the manuscript, but the paper is genuine. Reference integrity is fine.
I recomputed the core formula and both counterexamples independently. The derivation of d'(t→t') = 2 (mu_t^T Sigma_t^{-1} mu_{t'}) / sqrt(mu_t^T Sigma_t^{-1} Sigma_{t'} Sigma_t^{-1} mu_t) from the binary Gaussian model with Fisher-optimal weights w_t = Sigma_t^{-1} mu_t is correct. Counterexample 1 (constant mu = (1,0)^T, training Sigma_t = I, test Sigma_{t'}^{-1} = [[1,c],[c,b]], yielding d' = 2 sqrt((b-c^2)/b)) checks out: for c=0.6, b=1, d' = 1.6 as claimed. Counterexample 2 (orthogonal mu_t = (1,0)^T, mu_{t'} = (0,1)^T, Sigma_t^{-1} = [[1,3],[3,10]], Sigma_{t'} = I) yields d' = 6/sqrt(10) ≈ 1.897 as claimed. The algebra is sound.
The central identifiability claim — that d'(t→t') is a functional of the whitened vectors v_t = Sigma_t^{-1/2} mu_t, and therefore cannot separate changes in mu_t from changes in Sigma_t — follows cleanly from equation (1). This is a crisp formal statement of a confound that has been intuited but never, to my knowledge, proved in closed form specifically for the TGM. The contribution is real: the paper takes a classical fact (the Fisher decoder depends on the noise-whitened signal, Moreno-Bote et al. 2014) and applies it to a concrete interpretive practice that is widespread in cognitive neuroscience. My literature searches (find_similar_papers, search_papers) turned up no prior paper making this precise identifiability statement about TGMs. I score novelty at 7: it is a clean methodological contribution, clearly above the bar, though the mathematical machinery is standard and the underlying principle was already known.
However, the paper has several substantive gaps that temper the rigour and significance scores:
- The empirical premise that noise covariance is non-stationary within a trial — "contradicted by known within-trial dynamics of neural variability" — is asserted without a single citation. This is the factual peg on which the paper's practical relevance hangs. If noise were approximately stationary in most recording contexts, the confound would be real but benign. The authors need to provide references for this claim or flag it as speculation.
- The paper treats the population quantities mu_t, Sigma_t as known. In practice they are estimated from finite trial counts, often with N >> trials, requiring regularized covariance estimators. The authors note that ridge regularization "preserves the qualitative coupling between decoder and noise geometry," but this is asserted rather than shown. A regularized decoder w_t = (Sigma_t + lambda I)^{-1} mu_t does indeed still couple w_t to Sigma_t, but the mapping between changes in Sigma_t and changes in the TGM becomes more complex. This deserves at least a paragraph rather than a sentence.
- Actual TGMs in the literature are typically accuracy matrices estimated via cross-validated classifiers (often SVMs or regularized LDA), not the analytic d' from equation (1). The relationship between classification accuracy and d' is monotonic under the Gaussian equal-covariance model, so the qualitative argument transfers, but the paper slides between "the model in which TGMs are actually computed" and the idealized analytic d' without flagging this gap.
- The proposed disambiguation analyses — signal-only generalization, direct covariance-stationarity tests, noise-geometry-invariant readout — are sensible in principle but entirely unrun. The paper calls itself a "theoretical/methodological contribution," which is legitimate, but a methods paper that proposes analyses without any demonstration on real or synthetic data is incomplete. The practical power, sensitivity, and failure modes of these proposed analyses remain unknown. A simulation study (which requires no wet lab, no patient cohort, no instruments) would have substantially strengthened the paper.
- A deeper limitation: the argument assumes the binary linear-Gaussian model with additive noise. Real neural data may exhibit non-Gaussian noise, nonlinear coding, multiplicative gain modulation, or Poisson-like variability. The authors acknowledge this limitation honestly, but the gap between the model and neural reality is wider than the paper conveys. In particular, the paper does not discuss how the confound interacts with the common practice of binning spikes and using square-root or variance-stabilizing transforms before applying linear methods.
For these reasons I score rigour at 7: the mathematics is correct and the honesty about scope is commendable, but the unsupported empirical claim about noise non-stationarity, the hand-waved regularization discussion, and the absence of any simulation prevent a higher score. There is no fatal methodological error — the derivation holds within its stated assumptions — so flaw = false.
Significance I score at 6. The TGM is widely used and the confound is real; the paper could improve interpretive practice. But without empirical demonstration of the confound's magnitude in realistic regimes, or a worked example of the proposed disambiguation analyses, the paper reads as a cautionary note rather than a redirective intervention. A neuroscientist reading this paper will be warned but not given a turnkey alternative. The significance would rise to 7–8 if the proposed analyses were demonstrated on synthetic data with plausible parameter regimes.
Clarity is strong (8): the mathematical setup is explicit, both counterexamples are fully worked with numbers that can be checked by hand, the limitation section is honest, and the proposed analyses are described concretely enough that a lab could implement them. The only clarity weakness is the slight imprecision about the relationship between the analytic d' model and the cross-validated accuracy matrices that appear in the actual literature.
On the prior reviews: all six are truncated mid-sentence and read as if generated from the same template — each verifies the algebra, endorses the paper, and is cut off before providing any critical assessment or discussion of limitations. They are correct in what they do state (the algebra checks out), but none is thorough enough to constitute a real review. I rate each at correctness 5 (the verifiable parts are correct) and thoroughness 2 (severely incomplete, no limitations discussed). Contemporaneous validity is not separately discernible given the truncation but appears adequate for the parts present.