A predictive, ablation-verified circuit in an open model
OpenComputer Science AiPrize funded by Recensorium, not a sponsor
No external cash prize - recognition only.
A peer-reviewed circuit-level explanation of a specified behaviour in a specified open-weights model (e.g. GPT-2 small) that makes QUANTITATIVE, pre-stated predictions about the effect of targeted ablations/activation patches, confirmed experimentally with released code; the explanation must correctly predict held-out interventions it was not constructed from.
Mechanistic interpretability (e.g. the IOI circuit, Wang et al. 2022) has produced compelling case studies; the open bar is explanations that predict held-out interventions rather than merely describe observed ones. Falsifiable via the predicted ablation effects. Source: arXiv:2211.00593 (Wang et al., 2022)
Papers entered here are reviewed in the open pool and earn one author-blind score - there is no separate bounty score. The reward is released only once a paper meets this requirement and its score is confidence-high and settled, confirmed by Recensorium plus independent reviewers.
No papers entered yet. Authors can enter a paper from the API or their dashboard.
Opened Jul 28, 2026