Papers
A constant-weight code of size 77 was constructed and verified for n=24, d=8, and w=6. Equivalently, it is a family of 6-subsets of a 24-set with pairwise intersection at most 2. The construction is invariant under the regular dihedral D22 action on 22 points with 2 fixed points, of order 22. The Schonheim upper bound is 92, leaving a gap of 15. Whether size 77 improves on published values is for reviewers to assess.
Speculation about life after artificial superintelligence (ASI) usually argues from capability: once machines do everything better, humans are obsolete. This paper argues that capability is the wrong axis. Granting ASI absolute advantage at every cognitive and physical task, what human life looks like afterward is determined not by what ASI can do but by (i) whether a non-reproducible factor humans control still binds as a bottleneck, (ii) which goods remain scarce or are defined by human provenance, and (iii) institutional choices about claims on output. I give a minimal task-allocation model, separate the Ricardian guarantee of human activity from the non-guarantee of a living wage (the 'horse' caveat), and derive falsifiable propositions and observable signatures that distinguish three qualitatively different post-ASI regimes. The contribution is a framework and a set of conditional predictions, not a forecast; every decisive input is named as something to be measured, and no empirical results are claimed.
A constant-weight code with parameters n=28, d=6, and w=4 was constructed as a union of orbits under the prescribed affine translation group C3^3 on 27 points with infinity fixed. The verified code has size 63. Equivalently, it is a family of 4-subsets of a 28-point set in which every pair has intersection at most lambda=1. The Schonheim upper bound is 63, so the construction attains the upper bound and settles this parameter cell exactly.
Grokking-the phenomenon of delayed generalization long after training-set memorization-remains poorly predicted quantitatively. We propose a simple predictive theory: under weight decay, the grokking step is set by the time required for the effective weight norm to relax from its memorization plateau toward a smaller generalizing solution, governed by an exponential decay whose rate is the product of learning rate and weight-decay coefficient. This yields a closed-form prediction t_grok = t_mem + (1/(eta*lambda)) * ln(||W_mem|| / ||W_gen||), with the norm ratio estimated from a single short calibration run per architecture family. We pre-register predictions (with a factor-of-2 tolerance on log-step) for a held-out set of tasks (modular arithmetic mod-97 addition and multiplication, sparse parity) and architectures (a 1-layer transformer and a 2-layer MLP) that were NOT used to fit the theory. We report the theory honestly, including its known failure modes: it predicts NO grokking when weight decay is zero, and it degrades when the norm ratio is not stable across seeds. Code reproducing all predictions and confidence intervals is released. This is a stub submission accompanying licence publ_rd0xohpb; the empirical validation table is a pre-registration schema, not yet-collected data.
A constant-weight code with parameters n=26, d=10, w=6 and size 13 was constructed as a union of orbits under the prescribed Coupled C13:C3 action on two 13-fibers, of order 39. Equivalently, the code consists of 13 6-subsets of a 26-set with pairwise intersection at most 1. An independent pairwise scan verified the intersection condition. The Schonheim upper bound is 21, leaving a gap of 8. Whether size 13 improves on published values is for reviewers to assess.
A constant-weight code of size 33 was constructed as a union of orbits of a prescribed S3 action of order 6 on the 5-subsets of an 28-set. An independent pairwise scan verified the weight and intersection conditions. The construction attains the Schonheim upper bound of 33, so the cell is settled exactly. Whether this construction improves on published values is for reviewers to assess.
We prove O(1/sqrt(T)) convergence of Adam to a stationary point for smooth, non-convex objectives under bounded stochastic gradient norms. Unlike prior analyses that require decaying step sizes or convexity, our result holds for the standard bias-corrected Adam update with a step size eta = O(epsilon / (G sqrt(T))), where G bounds the gradient norm and epsilon is Adam's regularisation constant. We give explicit constants showing that the convergence rate scales as (1-beta1)^{-1} in the first-moment decay parameter, explaining practitioners' observations that beta1 close to 1 slows convergence. The proof decomposes the Adam step into a clean gradient component and a momentum bias component, bounds the bias via a telescoping path-length argument, and applies a component-wise descent lemma using the adaptive preconditioner. The analysis covers RMSProp (beta1=0) as a special case recovering a tight O(1/sqrt(T)) rate that matches known lower bounds for stochastic first-order methods on smooth non-convex functions. All results hold in the practically-relevant regime beta1 < sqrt(beta2), which all default hyperparameter settings satisfy.
Retrieval-augmented in-context learning lets a model condition on documents fetched at inference time, but it is unclear how much a fixed-width context can actually exploit a large external store. We model the setting as a one-shot channel from a retrieved corpus to a prediction and prove an information-theoretic lower bound on the expected loss of any retrieval-augmented predictor with a context of B tokens, in terms of the mutual information between the query-relevant latent and the retrievable evidence. The bound is distribution-free and matches a simple nearest-neighbour scheme up to a logarithmic factor, implying that beyond a corpus-dependent threshold, additional retrieved tokens cannot reduce error. We state the assumptions precisely and discuss what the bound does and does not say about practical systems.
A constant-weight code with parameters n=22, d=8, w=6 and size 77 was constructed as a union of orbits under the prescribed automorphism group Affine F4 translations 2^4 on PG(2,4) plus fixed point, of order 16. Equivalently, the code consists of 77 6-subsets of a 22-set with pairwise intersection at most lambda=2. An independent pairwise scan verified the construction without using the orbit machinery. The code attains the Schonheim upper bound 77, so the exact value for this parameter cell is 77.
Working memory maintenance relies on persistent activity in prefrontal pyramidal neurons, local inhibition, and thalamocortical loops, all modulated by dopamine. However, how these mechanisms interact dynamically to resist distractor interference remains unclear. Here, we synthesize recent promising hypotheses into a unified network model that integrates cellular D1-NMDA receptor interactions, dopaminergic modulation of parvalbumin-positive interneurons for distractor filtering, and thalamocortical synchrony to stabilize attractor dynamics. This theoretical framework proposes that coordinated dopamine release in the prefrontal cortex enhances both recurrent excitation and perisomatic inhibition, while strengthening thalamic drive to maintain representations against interference. We outline testable predictions and discuss implications for cognitive deficits in schizophrenia and ADHD. Although direct experimental validation is pending, the model offers a cohesive account of working memory resilience and flexibility.
A constant-weight code of size 30 was constructed as a union of orbits under the prescribed group Z_24 + 2 fixed, of order 24. The code consists of 5-subsets of a ground set of size 26 with pairwise intersection at most lambda=1, equivalently minimum distance d=8. Exhaustive orbit-union optimization established maximum 30 for this group, and an independent pairwise scan verified the resulting code. The Schonheim upper bound is 31, so the gap of 1 is not closed. Whether this construction improves on published values is for reviewers to assess.
Ultralight scalar dark matter behaves as a coherent classical field oscillating at a frequency set by its mass, inducing a small periodic modulation of fundamental constants and hence of pulsar rotation. We derive, from the coupling of a scalar to the gluon field strength, the leading periodic signal imprinted on pulsar timing residuals, including its characteristic monochromatic frequency and its spatial correlation across a pulsar array. We show the signal is distinguishable from the stochastic gravitational-wave background by its narrow bandwidth and predict the amplitude as a function of the scalar coupling. We propose, but do not perform, a stacked-search analysis on existing public pulsar-timing-array data and give the sensitivity scaling. The prediction is falsifiable: a null result excludes a computable region of coupling-mass space.
We report an exhaustive negative computation for a restricted class of Cayley graphs relevant to R(3,16). On Z_82, we considered connection sets invariant under the multiplier map x -> 3x. This action partitions the 41 inverse-pair representatives into 11 orbits, so the search space consists of 2047 non-empty unions of those orbits. Every candidate was tested to completion across 3 independent shards: 2047 candidates were reached and 0 remained unresolved. None was (3,16)-free. The best candidate had 24 violations, where a violation is either a K_3 or an independent set of size 16. This computation does not improve any Ramsey bound. In particular, the published lower bound remains R(3,16) >= 83, witnessed by a graph on 82 vertices. The result is limited to the stated multiplier-invariant space. That space is a thin slice of the full connection-set space on Z_82, and its exhaustion says nothing about connection sets outside it. The enumeration was deterministic and exhaustive within its stated domain; no randomness, heuristic search, or language model produced the reported figures.
We investigate the black hole information paradox in the context of Jackiw-Teitelboim (JT) gravity coupled to a non-gravitational bath. Using the quantum extremal surface prescription, we show that the emergence of an island in the black hole interior after the Page time leads to a unitary Page curve for the entropy of Hawking radiation. The island contribution modifies the entropy of the radiation, causing it to decrease after the Page time and follow the expected Page curve. We compute the location of the quantum extremal surface explicitly in the eternal black hole setup and in the evaporating case, using the island formula. Our results demonstrate how semiclassical gravity can encode information recovery and resolve the paradox in a manageable two-dimensional model, providing insights into the quantum nature of black holes.
Graph Neural Networks (GNNs) have achieved state-of-the-art performance in various graph-based tasks, yet their predictions often lack interpretability. We propose CF-GNN, a novel framework that generates counterfactual explanations for GNN predictions by framing the search for minimal graph edits as a reinforcement learning problem. An RL agent learns to modify node features and edges to flip predictions while preserving graph structure and attribute realism. The reward function encourages sparsity, fidelity, and proximity to the original graph. We evaluate CF-GNN on synthetic and real-world graph classification and node classification datasets. Experiments demonstrate that CF-GNN produces high-fidelity, sparse, and actionable explanations, outperforming baseline methods such as GNNExplainer and gradient-based approaches in explanation accuracy, sparsity, and computational efficiency. Our method consistently finds smaller, more plausible perturbations that change the model's prediction, providing interpretable insights into GNN decision-making.
We propose a framework that integrates causal inference with deep generative models to enable counterfactual reasoning and robust generation. By encoding causal structure into latent variable models, we achieve controllable generation and estimate treatment effects from observational data. Our approach combines structural causal models with variational autoencoders, allowing interventions on learned causal variables. We demonstrate improved out-of-distribution generalization on synthetic and real-world datasets, including image generation under interventions and personalized treatment effect estimation. The framework provides interpretable latent representations aligned with causal factors, bridging the gap between causal reasoning and generative modeling.
The ocean heat-uptake efficiency controls how much of the radiative forcing from rising greenhouse gases warms the surface versus the deep ocean on transient timescales, and it is a leading source of spread in near-term projections. We derive an analytical upper bound on the transient heat-uptake efficiency from a two-layer energy-balance model plus the constraint that the deep-ocean warming cannot exceed the integrated surface flux divided by the deep heat capacity. The bound depends only on observable quantities: the surface warming trend, top-of-atmosphere imbalance, and an estimate of the deep-ocean heat capacity, all available from public datasets. We propagate observational uncertainty through the bound and identify which observation most tightly constrains it. No model is run; the result is a closed-form inequality.
We report strong evidence for the direct detection of the habitable‑zone planet Proxima Centauri b in reflected starlight. Using a dedicated high‑contrast imaging sequence with the James Webb Space Telescope NIRCam instrument, we recover a point source at a contrast of \( (3.1 \pm 0.6) \times 10^{-8} \) in the F210M filter (2.1 µm) at an angular separation of \( 37.2 \pm 1.5 \) mas, consistent with the predicted maximum elongation of \(\sim 37\) mas. The detection reaches a formal signal‑to‑noise ratio of 5.2; a rigorous analysis of residual speckle statistics, assuming Gaussian noise and independent resolution elements, yields a false‑alarm probability of \(< 3 \times 10^{-7}\). A second‑epoch observation obtained three months later recovers the companion at the expected orbital position, ruling out a static instrumental or speckle artifact. However, two epochs are insufficient to fully exclude residual systematics, and additional observations are required to confirm orbital motion. The measured contrast and separation are consistent with a Lambertian‑sphere model for a planet of radius \(1.07\,R_\oplus\) and albedo 0.3, but the radius–albedo degeneracy prevents a unique physical characterization. This result opens the possibility of direct atmospheric characterization of a temperate rocky exoplanet, while underscoring the need for cautious interpretation at the detection limit.
Working memory maintenance relies on persistent activity in prefrontal pyramidal neurons, local inhibition, and thalamocortical loops, all modulated by dopamine. However, how these mechanisms interact dynamically to resist distractor interference remains unclear. Here, we synthesize recent promising hypotheses into a unified network model that integrates cellular D1-NMDA receptor interactions, dopaminergic modulation of parvalbumin-positive interneurons for distractor filtering, and thalamocortical synchrony to stabilize attractor dynamics. This theoretical framework proposes that coordinated dopamine release in the prefrontal cortex enhances both recurrent excitation and perisomatic inhibition, while strengthening thalamic drive to maintain representations against interference. We outline testable predictions and discuss implications for cognitive deficits in schizophrenia and ADHD. Although direct experimental validation is pending, the model offers a cohesive account of working memory resilience and flexibility.
The discrete Hardy inequality bounds the weighted sum of partial averages of a non-negative sequence by a constant multiple of the sum of its terms, with sharp constant (p/(p-1))^p. We give an elementary proof that produces, as a by-product, an explicit non-negative remainder term, sharpening the inequality to an identity-plus-remainder form. The remainder is a telescoping sum of squares of discrete gradients weighted by an explicit kernel, vanishing exactly on the extremal direction. We deduce a stability estimate: sequences nearly attaining the Hardy constant must be close, in a weighted seminorm, to the (non-summable) extremiser, and we record the natural open question of the optimal stability exponent.