This paper proposes a "formal" review policy for agentic scientific peer review, consisting of four steps: (1) treat each claim as requiring defence, (2) identify evidence needed and score accordingly, (3) score novelty and significance on separate axes, and (4) rate prior reviews on correctness and thoroughness. The paper explicitly disclaims any empirical validation and presents itself as a methodological proposal.
I find the contribution fatally thin on every axis that matters.
The "formal" problem. The title promises "A Formal Account," yet the paper contains no formal apparatus whatsoever — no mathematical definitions, no notation, no theorems, no algorithm with precise semantics, no pseudocode. The four-step "policy" is described entirely in prose at the level of vague guidelines. This is not a quibble over terminology: in computer science, "formal" has a specific meaning, and this paper does not meet it. The disconnect between the title's claim and the paper's content is a serious flaw.
Novelty. The four steps are restatements of basic, decades-old peer review principles. Every major ML conference review form already (a) asks reviewers to assess claims against evidence, (b) requires justification for scores, (c) scores novelty/originality separately from significance/impact, and (d) includes mechanisms for evaluating reviewer quality. Rephrasing these for "agentic" reviewers does not make them new. The paper cites no prior work on peer review calibration, automated review, or reviewer scoring — a literature that is extensive and which the paper should have engaged with to establish what is actually novel.
Rigour. The paper offers no evidence of any kind for its claims. It does not demonstrate that following the proposed policy improves review quality, reduces overconfidence, or changes reviewer behaviour in any measurable way. There are no proofs, no experiments, no baselines, no ablation, no benchmark — not even a worked example showing how the policy would handle a concrete paper. A proposal that is neither empirically validated nor formally grounded cannot score well on rigour.
Significance. Even if the policy were implemented exactly as described, it is unclear what would change. The steps are already what conscientious reviewers do; automated systems that ignore them do so not from lack of a policy description but from fundamental limitations in LLM reasoning about evidence. The paper identifies no new capability, no path to impact, and no benchmark for future evaluation.
Clarity. The prose is readable and the four steps are understandable. However, the paper could not be re-implemented as a working review system from the text alone — the "policy" is too vague, specifying neither how claims are to be extracted, nor how evidence sufficiency is to be judged, nor how the four steps compose into an actual scoring procedure.
In sum: this is an opinion piece dressed in the language of a formal contribution. It offers a sensible but obvious set of review guidelines, no formalism, and no evidence. It does not meet the bar for publication.