SciGroveBeta
Neuroscience

Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia

Xiang Guan, Roger D. Newman-Norlund

Featured August 19, 2026

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

By giving language models brain damage (perturbing layers) and seeing how their mistakes change, the paper finds specific model parts that cause certain language errors, just like how doctors link brain damage to human speech problems.

In depth
The paper introduces PRISM, a novel framework that adapts human neuroimaging's subtraction analysis to large language models (LLMs) to provide spatially resolved and falsifiable interpretability. By systematically perturbing LLM layers and analyzing resulting error patterns, PRISM identifies which layer ranges are causally necessary for specific behavioral failures, drawing direct parallels to lesion-symptom mapping in human brains.

Key Takeaways

  • 1
    PRISM adapts subtraction analysis from human neuroimaging to LLMs, offering a novel, falsifiable approach to mechanistic interpretability.
  • 2
    The framework demonstrates a robust phonemic-favoring dissociation in both LLM layers and human frontal-perisylvian cortical regions, replicating established findings in aphasia.
  • 3
    The study introduces a dose-response analysis for LLMs, showing that perturbation magnitude and density behave like graded lesions, strengthening the analogy to human brain damage.

Conceptual Flow

HIGH LEVEL
1
Methodology: Bridging Brains and AI with Subtraction

The study gives language models 'brain damage' and compares their mistakes to real patient mistakes, finding similar patterns.

Human Brain Data
AI Model Data
Apply Same Logic
Brain Regions for Tasks
AI Layers for Tasks
2
Results: Specific Layers for Specific Errors

They found that certain AI layers, like parts of the human brain, are especially important for making sound-based errors, but less so for meaning-based errors.

AI Model Layers
Human Brain Areas
Find Key Differences
Sound-Error Layers
Sound-Error Brain Parts

This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.