The paper presents a formal framework for calibrating autonomous scientific reviewers. The four-step policy is clearly described: treating claims as requiring defense, scoring based on evidence, separating novelty and significance, and rating prior reviews for correctness and thoroughness. The motivation is timely given the increasing use of agentic systems in research evaluation.
However, a critical issue undermines the contribution: the paper explicitly states it is a "formal proposal" without empirical validation (in Abstract, Limitations, and Conclusion), yet it includes a detailed simulation study (the appended experimental section) that appears to validate the policy on synthetic data. This contradiction makes the paper's contribution unclear: is it a pure formal proposal or does it also claim empirical evidence? The simulation is not mentioned in the abstract or body, and the Limitations section disavows empirical validation. This inconsistency must be resolved. If the simulation is part of the submission, the paper should be restructured to integrate it and revise the claims accordingly. If it is not, it should be removed or clearly labeled as envisioned future work.
Assuming the simulation is intended as preliminary evidence, the reported results are promising: improved calibration, discrimination, and reduced dependence on rhetoric. However, the synthetic dataset construction and expert oracle are not described in sufficient detail, making it difficult to assess the validity of the evaluation. No baselines beyond a single-pass prompt are considered, and no ablation studies are performed. The policy is implemented as a structured prompting strategy, but no details of the prompts or LLM used are provided.
The formal framework itself is elegant and addresses a real problem. The separation of novelty and significance is valuable. The idea of rating prior reviews could improve meta-review consistency.
Minor points:
- The paper would benefit from a concrete example of each policy step.
- The relationship to existing work on structured peer review and calibrated classification should be discussed.
- The policy requires reviewers to identify necessary evidence, but how this is operationalized in practice (especially for agentic systems without experimental capabilities) is left vague.
In summary, the paper tackles an important problem but is not ready for dissemination due to the fundamental inconsistency between its stated goal and its presented evidence. Major revisions are required to clarify the contribution and (if applicable) properly integrate and describe the empirical study.