SciGroveBeta
Neuroscience

Brain-LLM Alignment Tracks Training Data, Not Typology

Dongxin Guo, Jikun Wu, Siu Ming Yiu

Featured June 8, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

It turns out that how well AI language models match human brain activity depends mostly on the languages they were trained on, not just how different the languages are, with specific brain parts reacting differently.

In depth
The paper rigorously demonstrates that the observed "English advantage" in brain-LLM alignment is an artifact of training data composition, not an inherent property of English or the brain. They show that a Chinese-dominant LLM (Baichuan2-7B) reverses this alignment gradient, performing best for Chinese brains and worst for English, directly mirroring English-dominant models. Beyond training data, formal typological distance and region-specific processing in the brain's language network also independently modulate alignment.

Key Takeaways

  • 1
    The apparent "English advantage" in brain-LLM alignment is primarily driven by training-language dominance, not inherent linguistic typology.
  • 2
    Typological distance between languages independently predicts alignment degradation, with syntax-associated brain regions (IFG) showing steeper gradients.
  • 3
    Tokenization fertility, the average number of subword tokens per word, accounts for a significant portion of cross-linguistic shifts in optimal encoding layers.

Conceptual Flow

HIGH LEVEL
1
Comparing Brains and AI Across Languages

Researchers showed different language stories to people and AI models, then checked how well the AI's internal thoughts matched the brain's activity.

Human Brain Activity
AI Model Thoughts
Compare Patterns
How Well They Match
2
AI's 'Language Bias' Comes from Training

The study found that AI models match brains best for the language they were trained on most, proving that the AI's training data, not the language itself, creates this bias.

AI Trained on English
AI Trained on Chinese
Match Brains
English AI matches English Brains
Chinese AI matches Chinese Brains