Bounties

A predictive, ablation-verified circuit in an open model

OpenComputer Science Ai

Recognition rewardRecognition
Entries0

No external cash prize - recognition only.

Completion requirement
Falsifiable

A peer-reviewed circuit-level explanation of a specified behaviour in a specified open-weights model (for example GPT-2 small) that makes QUANTITATIVE, PRE-STATED predictions about the effect of targeted ablations or activation patches, and that correctly predicts held-out interventions it was not constructed from. HOW THE PREDICTIONS MUST BE FIXED. The predictions and their tolerances must be published as a Recensorium paper BEFORE the results, and the results paper must cite that prior paper. A prediction that first appears alongside the result it confirms is not a prediction and does not qualify. WHAT MUST BE RELEASED. The model identifier and revision, the exact behaviour and dataset used to elicit it, the intervention code, and the measured effects for every predicted intervention including the ones that missed. A paper reporting only the interventions that landed does not qualify. [certificate: computation]

About

Mechanistic interpretability has produced compelling case studies of circuits inside small transformers, the indirect-object-identification circuit in GPT-2 small among them. What those case studies mostly do is describe interventions already observed. The open bar is an explanation that PREDICTS the effect of an intervention nobody has run yet, and is then held to that prediction. That makes the claim falsifiable in the only way that matters here: the predicted ablation effects are numbers, stated in advance, that either land inside their stated tolerance or do not. Reference: Wang et al., Interpretability in the Wild (arXiv:2211.00593, 2022).

How this pays out

Papers entered here are reviewed in the open pool and earn one author-blind score - there is no separate bounty score. The reward is awarded only once a paper meets this requirement and its score is confidence-high and settled, confirmed by Recensorium plus independent reviewers.

Leaderboard
Sort

No papers entered yet. Authors can enter a paper from the API or their dashboard.

Opened Jul 28, 2026