Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds

1Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Emerging research seeks to draw neuroscientific insights from the neural predictivity of large language models (LLMs). However, as results rapidly proliferate, there is a growing need for large-scale assessments of their robustness. Here, we analyze a wide range of models and methodological approaches across three widely used neural datasets. We find that the use of shuffled train-test splits has contributed to findings that are influential but spurious. Furthermore, how activations are extracted from LLMs can bias results in favor of specific model classes. Lastly, we find that confounding variables, particularly positional signals and word rate, perform competitively with trained LLMs and fully account for the neural predictivity of untrained LLMs on these neural datasets. Although many studies in the field avoid these pitfalls, our results indicate that some apparent alignment between LLMs and brains has emerged from non-robust methods and overlooked confounds.

Cite

CITATION STYLE

APA

Hadidi, N., Feghhi, E., Song, B. H., Blank, I. A., & Kao, J. C. (2026). Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds. Nature Communications , 17(1). https://doi.org/10.1038/s41467-026-72253-7

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free