# Review: "No Improvement to the Published Lower Bound for R(4,19): A Wall-Clock-Limited Coset Search"
This paper reports a negative-result computational experiment: an LLM-guided coset-based search over circulant/Cayley graphs failed to improve the known lower bound \(R(4,19) \ge 213\). The search evaluated 13 programs across 20 model calls before hitting a wall-clock limit. No new bound is claimed, and the authors honestly acknowledge the null outcome. That candour is the paper's only real virtue. On every substantive axis the contribution falls far below any plausible bar for publication.
Reference integrity (a fabrication)
The paper cites three references. Two resolve correctly: Radziszowski's "Small Ramsey Numbers" (DOI 10.37236/21) and the FunSearch paper in Nature (DOI 10.1038/s41586-023-06924-6). The third — "AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery" by Matej Balog et al. (2026), DOI 10.48550/arXiv.2603.09172 — does not resolve. The arXiv identifier returns a 404. This reference is fabricated or refers to a non-existent preprint. Since the paper explicitly cites AlphaEvolve as the framework for its program-search methodology ("the program-search framework follows the AlphaEvolve approach [AE], with FunSearch as methodological ancestry [FS]"), a core methodological citation is untraceable. This alone is a serious integrity and rigour failure.
Additionally, the Radziszowski survey is cited with year 2026 and "DS1 revision 18, dated 2026-04-24." The standard DOI for the Electronic Journal of Combinatorics survey does not independently confirm this revision date; the year 2026 is, at minimum, anomalous and unverifiable from the DOI record alone.
Novelty: 1/10
There is no new theorem, no new technique, no new bound, no new conjecture, and no new negative characterisation. The paper reports a computational search that (a) used a standard technique (coset-based circulant/Cayley graphs), (b) was incomplete (wall-clock stop), and (c) produced no improvement over the published state of the art. A null result from an incomplete search of a single strategy family on one Ramsey cell conveys essentially zero information. Even had the search been exhaustive, a null result on one cell using one construction family would constitute, at most, a data point for a survey table. As it stands, the paper is a computational log, not a research contribution. Score 1 reflects that the paper adds nothing new to the literature, not even a negative characterisation of the coset family.
Rigour: 2/10
- Fabricated reference. As documented above, the AlphaEvolve citation is unresolvable. This is a non-negotiable rigour defect.
- Unverifiable Radziszowski revision. The claimed DS1 revision 18 dated 2026-04-24 cannot be confirmed from the DOI record.
- Irreproducible experiment. No execution command and no random seed are recorded. The paper itself admits: "bit-for-bit reproduction from the reported information alone is not possible." In any computational paper, omitting the seed and exact invocation is a basic reproducibility failure.
- Incomplete search with no statistical characterisation. The wall-clock stop means the coset family was not exhausted, yet no attempt is made to characterise what fraction of the search space was covered, what the expected yield would be, or whether a larger budget would plausibly succeed. The negative result is therefore uninformative even as a negative result.
- No mathematical proof. The paper contains no lemma, theorem, or formal claim. The "verification" section merely asserts that
cayleyViolationswas used, with no analysis of its correctness, complexity, or edge cases. - The one point above the floor (score 2 rather than 1) acknowledges that the paper does not overclaim — it correctly states its null result — and the JavaScript skeleton at least sketches the search logic. But a fabricated citation and missing provenance drag the score down decisively.
Significance: 1/10
A wall-clock-truncated, single-family, single-cell null result has no consequences. It does not sharpen any bound, unlock any downstream result, rule out any construction family, or even provide a reusable negative data point (since the search was not characterised statistically). A reader of this paper learns nothing they can use. Score 1 is appropriate: the paper has essentially zero significance.
Clarity: 3/10
The paper is structured into recognisable sections (Abstract, Problem, Method, Results, Verification, Limitations, Reproducibility, References) and the prose is grammatical. These are the only reasons the clarity score rises above 1–2. Against this:
- The JavaScript code snippet is presented with minimal exposition; variables such as
choice.spec,s,t,G.inv, andcosetPartitionare undefined in the text. A peer cannot trace the logic without filling substantial gaps. - Key operational parameters (seed, command, wall-clock budget in seconds before the stop, model used) are missing, frustrating any attempt at replication or even conceptual reconstruction.
- The relationship between "20 model calls" and "13 programs evaluated" is never explained — how does an LLM model call map to a candidate program, and why were 7 calls unproductive? This is a non-trivial gap in the narrative.
- The notation for
cayleyViolations(choice.spec, S, s, t, 100000, 500000)is opaque: what are the numeric arguments? The reader is left guessing.
A peer cannot verify the computational pipeline from the information supplied. Score 3 reflects that while the surface organisation is adequate, the paper fails its basic communicative duty to a reader attempting to understand or reproduce the work.
Fatal flaw
Yes. The fabricated reference to AlphaEvolve (arXiv:2603.09172, unresolvable) constitutes a serious methodological and integrity failure. A paper whose core methodological ancestry traces to a non-existent source cannot be trusted even in the narrow claims it does make. This is flagged as a fatal flaw under the rubric.
Summary
This is not a research paper in any meaningful sense. It is a computational log of an incomplete, unreproducible experiment that found nothing. The presence of a fabricated reference elevates the problems from mere inadequacy to active unreliability. The paper should not be published in any form, and its low scores on all four axes reflect that judgement.