Understanding how the brain transforms sensory input into thought and behavior ranks among neuroscience's most ambitious goals — and for decades the field has struggled to find experimental tools sharp enough to distinguish competing theories. A methodological advance now offers a principled path forward by redesigning how stimuli are selected to pit neural network models directly against one another.
The core challenge, reviewed in Nature Reviews Neuroscience, is that modern neural network models of brain function carry enormous parametric capacity, enabling any single model to fit a wide range of experimental data. When tested on stimuli drawn from the same distribution the models were trained on, competing models often produce nearly identical predictions, making experimental discrimination effectively impossible. The emerging solution is to actively optimize stimulus sets for maximal "controversy" — choosing or synthesizing inputs that force divergent predictions across candidate models. Key methodological decisions include selecting a prior over candidate stimuli (favoring naturalistic images to preserve ecological validity), defining a quantitative measure of discrimination power, and implementing a search or synthesis procedure that maximizes that power while remaining practically deliverable to participants.
This represents more than a technical refinement — it is a philosophical shift in experimental design philosophy. Historically, neuroscientists chose between naturalistic stimuli (ecologically valid but low in discriminative leverage) and artificial stimuli (high in discriminative power but divorced from real-world relevance). The controversial-stimuli framework offers a tempered synthesis: stimuli constrained by naturalistic priors yet optimized for model adjudication. The approach aligns conceptually with adversarial testing in machine learning, now imported into cognitive neuroscience. Key limitations remain: the framework's validity depends on the completeness of the model pool — if the true neural mechanism is not represented among candidate models, discrimination among wrong models advances understanding only modestly. This is an intellectually significant methodological contribution that could accelerate theory testing in systems and cognitive neuroscience, though empirical demonstrations at scale will determine its transformative reach.