This submission proposes a lightweight, closed-form timing law for the grokking step, t_grok = t_mem + (1/(eta*lambda)) * ln(||W_mem|| / ||W_gen||), motivated by a two-phase story of rapid memorization to a large-norm solution followed by exponential norm relaxation under weight decay. It packages the idea as a pre-registered, falsifiable protocol with a factor-of-2 log-step tolerance, explicitly held-out task/architecture combinations, five seeds, and clear rejection criteria (including the strong claim that zero weight decay precludes grokking). The writing is clear and the limitations section is unusually candid.
None of that rescues the central fact: the paper is an explicit stub. The sole results table consists of literal ‘(fill)’ placeholders; the abstract and body both state that no empirical data have been collected and that the table is merely a pre-registration schema. A contribution whose only substantive claim is a quantitative prediction cannot be assessed when zero predictions and zero observations are supplied. At best this is a Stage-1 registered-report outline; it is not a completed research paper suitable for standard peer review.
Even setting the missing data aside, the theoretical development is thin. The exponential-relaxation rate eta*lambda is asserted rather than derived from the underlying gradient dynamics (cross terms with the still-nonzero data gradient are ignored). The existence of a task-independent norm threshold that triggers generalization is likewise asserted without justification, yet the entire predictive apparatus rests on it transferring from modular arithmetic to sparse parity and from transformers to MLPs. The norm ratio itself is to be estimated from a single short calibration run per architecture family—an estimator the authors themselves flag as seed-sensitive—without any robustness procedure. t_mem is treated as an externally observed quantity, so it remains unclear how much predictive work is done by the novel term versus simply reading off memorization time. No baseline comparator (constant multiple of t_mem, existing grokking accounts, etc.) is pre-registered. The protocol’s statistical power is low (three held-out cells) and no false-positive/false-negative analysis is offered. Finally, the paper does not engage the existing literature on weight-norm mechanisms for grokking, some of which has already tested closely related claims.
The pre-registration discipline and the honesty about failure modes are genuine strengths that the community should encourage. They do not compensate for the total absence of evidence, the underspecified modeling choices, and the lack of differentiation from prior work. I recommend rejection. The authors should run the calibration and held-out experiments they themselves designed, populate the table with real numbers, add sensitivity analyses and baseline comparisons, and resubmit a complete manuscript; only then can the theory’s validity be judged.