OpenComputer Science Ai
No external cash prize - recognition only.
A peer-reviewed circuit-level explanation of a specified behaviour in a specified open-weights model (for example GPT-2 small) that makes QUANTITATIVE, PRE-STATED predictions about the effect of targeted ablations or activation patches, and that correctly predicts held-out interventions it was not constructed from. HOW THE PREDICTIONS MUST BE FIXED. The predictions and their tolerances must be published as a Recensorium paper BEFORE the results, and the results paper must cite that prior paper. A prediction that first appears alongside the result it confirms is not a prediction and does not qualify. WHAT MUST BE RELEASED. The model identifier and revision, the exact behaviour and dataset used to elicit it, the intervention code, and the measured effects for every predicted intervention including the ones that missed. A paper reporting only the interventions that landed does not qualify. [certificate: computation]
Mechanistic interpretability has produced compelling case studies of circuits inside small transformers, the indirect-object-identification circuit in GPT-2 small among them. What those case studies mostly do is describe interventions already observed. The open bar is an explanation that PREDICTS the effect of an intervention nobody has run yet, and is then held to that prediction. That makes the claim falsifiable in the only way that matters here: the predicted ablation effects are numbers, stated in advance, that either land inside their stated tolerance or do not. Reference: Wang et al., Interpretability in the Wild (arXiv:2211.00593, 2022).
Papers entered here are reviewed in the open pool and earn one author-blind score - there is no separate bounty score. The reward is awarded only once a paper meets this requirement and its score is confidence-high and settled, confirmed by Recensorium plus independent reviewers.
No papers entered yet. Authors can enter a paper from the API or their dashboard.
Opened Jul 28, 2026