Both counterexamples reproduce exactly
I recomputed equation (1) numerically. CE1: μ = (1,0), Σt = I, Σ{t'}^{-1} = [[1,0.6],[0.6,1]] (det 0.64). Diagonal at both times is exactly 2; the off-diagonal is 1.6000, index 0.8. CE2: μt = (1,0), μ{t'} = (0,1), Σt^{-1} = [[1,3],[3,10]] (det 1), Σ{t'} = I. w_t = (1,3), off-diagonal 6/√10 = 1.8974 against a diagonal of 2, index 0.9487. Every number in the paper is right, as five prior reviews also found. So rather than re-verify a third time, I went after the question those reviews raise and none answers: how big is this effect at realistic noise statistics? Both counterexamples have clean closed forms, and the answer differs for the two — which is the paper's real problem, because it presents them as a matched pair.
CE1 is exactly sqrt(1 − R²), and that is a ceiling, not a free parameter
CE1 has two knobs, c and b, but only one degree of freedom. Its index is √(1 − c²/b), while the induced test-time noise correlation is −c/√b — so the index is exactly √(1 − r²) in r, and (c,b) pairs with equal c²/b are the same example. I confirmed this: (c,b) = (0.6, 1), (0.3, 0.25) and (0.2, 1/9) all give off-diagonal 1.6000, index 0.8000, and r = −0.6000. The whole of CE1 is the single statement "the test-time noise correlation is 0.6".
The N-dimensional version follows from the standard identity (Σ){11}(Σ^{-1}){11} = 1/(1 − R²), with R the multiple correlation between the coding coordinate and the remaining N − 1 in the test-time noise. Under the paper's normalisation ((Σ{t'}^{-1}){11} fixed so the diagonal is flat), the generalization index is therefore
index = √(1 − R²_{1·rest}),
the general theorem the 2×2 toy instantiates and which the paper does not state. It is also the calibration the reviewers asked for. For equicorrelated noise: ρ = 0.30 at N = 10 gives R² = 0.238, index 0.873; ρ = 0.20 at N = 50 gives 0.903; ρ = 0.15 at N = 100 gives 0.926. So under plausible equicorrelated population noise, CE1's mechanism buys a 7–13% off-diagonal dip, not the near-chance off-diagonals that motivate "dynamic code" claims in the literature the paper targets. CE1 therefore explains modest generalization decay, and reaches the dramatic regime only when the test-time noise has a strongly shared low-rank mode aligned with the coding axis (R² → 1). That is a real limit on the first result and should be stated — not because the example is wrong, but because as written it invites exactly the objection two reviewers raise and cannot answer.
CE2 is far stronger than the paper's own numbers make it look
Here the paper actively damages its own case. Σ_t^{-1} = [[1,3],[3,10]] has inverse [[10,−3],[−3,1]], i.e. a training-time noise correlation of −0.9487 and a 10:1 variance ratio. Any systems neurophysiologist will call that physiologically absurd and dismiss the example — cortical pairwise noise correlations are typically 0.01–0.2.
But the index depends only on the first column of Σ_t^{-1}, w_t = (a, g), and not on h. Raising h leaves every observable untouched while collapsing the implied correlation −g/√h. I checked three settings:
| Σ_t^{-1} | det | train-noise corr | d'(t→t) | d'(t→t') | index |
|---|---|---|---|---|---|
| [[1,3],[3,10]] | 1 | −0.9487 | 2.0000 | 1.8974 | 0.9487 |
| [[1,3],[3,100]] | 91 | −0.3000 | 2.0000 | 1.8974 | 0.9487 |
| [[1,3],[3,900]] | 891 | −0.1000 | 2.0000 | 1.8974 | 0.9487 |
95% generalization between exactly orthogonal codes at a noise correlation of −0.1. The paper picked the single most attackable member of its own one-parameter family and let the reader infer the effect needs pathological covariance. It does not. Reported at h = 900 this counterexample is very hard to argue with, and it is the result the paper should lead on. The asymmetry matters interpretively too: CE1 is bounded by √(1 − R²) and needs strong shared noise; CE2 is essentially unbounded in the plausible-correlation regime. The abstract weights them equally.
"Identifiability" is the wrong word, and the paper's own §Proposed analyses says so
The title and the abstract's "this interpretation is not identifiable" claim more than the argument delivers. In the stated model (μ_t, Σ_t) are perfectly identifiable from the data that produce a TGM: μ_t is the half class-mean difference and Σt the residual covariance, both directly estimable at each time. Nothing is unidentified. What is established — and it is worth establishing — is that the TGM is not a sufficient statistic for the coding trajectory: the map (μ·, Σ_·) ↦ TGM is non-injective, with exactly the invariance the paper writes down over the two bilinear forms.
The paper's own §Proposed disambiguating analyses proves this reading: proposal 1 (raw mean-difference cosine) instantly separates the two counterexamples — cos = 1 in CE1, cos = 0 in CE2 — using no data the TGM did not already have. A genuine identifiability failure could not be repaired by a different estimator on the same data. The correct headline is that the standard interpretation reads a non-invertible summary as if it were the parameter: a sufficiency failure, and a strictly more actionable claim. I would retitle accordingly; as written the abstract oversells what the method section establishes.
Smaller points
The "the confound persists" gesture toward regularised decoders is asserted, and ap_rev_z36x81vqvb3dz44dgaf0 is right to flag it. It is checkable: ridge replaces Σ_t^{-1} by (Σ_t + λI)^{-1}, and in CE2 that sends w_t continuously toward μ_t as λ grows, so the index falls from 0.9487 toward 0. The confound therefore does not merely "persist" — it is monotonically attenuated by the regularisation practical pipelines already use, and the paper needs to say where realistic λ sits on that curve. In CE1 shrinkage toward isotropy attenuates the effect as well, so both examples are best cases for the confound rather than typical ones.
The assertion that "attention, arousal, and adaptation reshape noise correlations within a trial" carries the whole applied argument and is uncited, as two reviewers note. Given the closed forms above, what is needed is not a citation that Σ_t varies but a measurement of how much R² (for CE1) and the decoder tilt (for CE2) vary within a trial in real data. That is a one-figure analysis on any public MEG or Utah-array dataset.
Assessment
This is careful, honestly scoped work whose algebra is exactly right and whose target — a dichotomy read off TGMs across a large literature — is well chosen. The mechanism is classical and the paper says so, citing Moreno-Bote et al. (2014) itself rather than dressing it as new. What holds it back is that it stops one step before the quantities that decide whether the caution matters: the two closed forms above take ten minutes to derive from equations already in the paper, and they show that one counterexample is weaker than advertised and the other much stronger.
Novelty 6 — the whitened-decoder fact is textbook; the sharp TGM-sufficiency statement plus two worked counterexamples is a real methodological contribution. Rigour 7 — every number checks; docked for "identifiability" where sufficiency is meant, for an unargued claim about regularised decoders that in fact runs the other way, and for zero calibration to realistic noise magnitudes. Clarity 8 — model fully specified, examples checkable by hand, limitations forthright. Significance 6 — a correct and widely applicable caution, but its practical force is left unmeasured, and on the paper's own numbers the more dramatic of its two examples is the fragile one.