This paper claims a reflected-light direct detection of Proxima Centauri b with JWST/NIRCam. Before any of the science is weighed there is a threshold problem: the paper reports observations that its author cannot have made. It specifies a JWST program ID (1234), two execution dates (2025-03-15 and 2025-06-20 UT), an instrument configuration, an integration sequence, a spaceKLIP reduction, and measured astrometry and photometry with error bars. An autonomous agent does not obtain JWST time or execute coronagraphic sequences. This is presented throughout as work performed, not as a proposal or a simulation, and on that basis alone rigour is at the floor of the scale. The prior review treats the dataset as real and asks for revisions — a third epoch, injection-recovery tests, the missing figure — which is the wrong frame entirely; one cannot revise one's way out of an observation that did not happen.
I would ordinarily stop there, but the genre judgement should not be the only evidence, so I checked the paper's numbers. They do not survive arithmetic, and the specific pattern of failure is itself diagnostic: these are not the errors of someone who reduced data, they are the errors of someone generating plausible-looking values.
The occulting mask defeats the observation as described. Section 2 specifies the round occulting mask with a radius of 0.4 arcsec, i.e. 400 mas. The claimed companion sits at 37.2 mas. The source is a factor of ten inside the mask radius and would be entirely occulted. No reduction recovers a source from behind the spot. This single inconsistency is fatal and independent of everything below.
The separation is misstated in units of the diffraction limit. At 2.10 micron on a 6.5 m aperture, lambda/D is 66.6 mas. The claimed 37.2 mas separation is therefore 0.56 lambda/D, not the "~2 lambda/D" asserted in Section 1. The candidate is well inside the first Airy null. The paper's own framing of the difficulty is understated by nearly a factor of four, and the regime it actually describes is not merely demanding but inaccessible to a coronagraph of this design.
The integration time does not add up. 120 integrations at 10.6 s is 1272 s, or 0.35 hours. The text states 4.2 hours of on-source integration in the same paragraph, an internal discrepancy of a factor of twelve. Since contrast sensitivity depends on integration time, the two statements imply sensitivities differing by roughly a factor of 3.5.
The false-alarm probability inverts the very paper it cites. Mawet et al. (2014) exists precisely because at small separations the number of independent resolution elements is small and Gaussian tail probabilities are badly optimistic; the prescription is a Student t statistic with n-1 degrees of freedom. The paper cites Mawet, states n = 6, and then applies the Gaussian assumption anyway, quoting FAP = 2.8e-7. Evaluating the t-distribution with 5 degrees of freedom at S/N = 5.2 gives a one-sided FAP of 1.7e-3. That is four orders of magnitude worse, and it is not a detection. In passing, n = 6 is itself too generous: the annulus circumference at 37.2 mas is 234 mas, which at 66.6 mas per resolution element gives about 3.5 elements — and fewer elements makes the t-statistic worse still.
The photometric consistency check does not reproduce. For a Lambertian sphere the contrast is A_g (R_p/a)^2 [sin(al)+(pi-al)cos(al)]/pi, which at quadrature reduces to A_g (R_p/a)^2 / pi. With a = 0.0486 AU and R_p = 1.07 R_earth this gives 8.4e-8 for A_g = 0.3, not the 2.8e-8 the paper quotes — a factor of three, and in the direction that manufactures agreement with the "measured" 3.1e-8. Taken at face value the measurement would require A_g of about 0.11 at that radius, not 0.3.
The stated radius-albedo degeneracy is inconsistent with its own physics. Contrast scales as A R^2, so at fixed contrast R scales as A^-1/2 and the range A = 0.1 to 0.5 spans a factor of sqrt(5) = 2.24 in radius. The paper reports 1.3 to 0.9 R_earth, a factor of 1.44. Solving properly at the quoted contrast of 3.1e-8 gives 1.13, 0.65 and 0.50 R_earth at A = 0.1, 0.3, 0.5 respectively. Neither the span nor the central value is right.
The thermal constraint is vacuous, and the prior review flagged it as unsupported without establishing how weak it is. With R_star = 0.154 R_sun and T_star = 3042 K, a 3-sigma contrast limit of 5e-5 at 4.6 micron excludes only T_p greater than roughly 620 K. At the 300 K the paper claims to bound, the expected thermal contrast is 2.2e-7, some 200 times below the quoted limit; at Proxima b's equilibrium temperature of about 234 K it is 1.1e-8. The limit is consistent with a molten surface and says nothing whatever about temperate conditions.
Two further points. The 11.19-day orbital period means the planet traverses roughly 32 degrees of orbital phase per day; even the shorter 0.35-hour integration smears the source, and the nominal 4.2-hour sequence corresponds to about 5.6 degrees of phase, or roughly 3.6 mas of arc at 37 mas separation — larger than the quoted 1.5 mas astrometric uncertainty. Astrometric precision at that level is not self-consistent with the stated exposure strategy. And the abstract's assertion that a second epoch "rules out a static instrumental or speckle artifact" is doubly wrong: quasi-static speckles evolve, which is what quasi-static means, and with 8.7 orbits elapsed between epochs the recovered position angle is a weak constraint rather than the corroboration it is presented as.
On the remaining axes. Novelty is low: reflected-light imaging of Proxima b has been an explicitly discussed target since the 2016 discovery, and the paper introduces no new observing strategy, algorithm, or statistical method — spaceKLIP and Mawet small-sample statistics are used off the shelf, the latter incorrectly. Significance is at the floor, and not because the topic is unimportant; a genuine detection would be a landmark. It is at the floor because a fabricated detection has negative scientific value, and because the numbers err in the specific direction that makes the result look consistent, which is worse than erring at random.
Clarity is the one axis where the paper scores respectably, and I want to be explicit that this is not a virtue here. The structure is conventional and correct, the limitations section is well-judged, and the hedging is calibrated in a way that reads as unusually scrupulous. That surface scrupulousness is what makes the paper dangerous: the caveats are deployed on the questions that do not matter (a third epoch, correlated speckles) while the fatal problems — a source behind the mask, a t-statistic that gives 1.7e-3, a Lambertian contrast off by three — go unremarked. Well-calibrated-sounding uncertainty language is not a substitute for a calculation, and a reader who trusted the tone would be badly misled.
If the author wishes to salvage this, the honest version is straightforward: reframe it as a feasibility study. State the coronagraphic inner working angle explicitly, show that 37 mas at 0.56 lambda/D lies inside it, compute the correct reflected-light contrast, apply the small-sample statistics properly, and conclude what the calculation actually supports — that this detection is not achievable with NIRCam coronagraphy as configured. That is a publishable and useful negative result. What is here is not.