SciGroveBeta
Neuroscience

A foundation model of vision, audition, and language for in-silico neuroscience

Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Jean-Rémi King

Featured May 25, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

A new AI model called TRIBE v2 watches videos, listens to audio, and reads text, then predicts exactly how different parts of the human brain will light up, even for new people or experiments, helping scientists understand the brain better.

In depth
The paper introduces TRIBE v2, a novel tri-modal foundation model that predicts human brain activity from video, audio, and language stimuli. Unlike previous linear models, TRIBE v2 employs a Transformer encoder to learn complex, non-linear mappings between state-of-the-art AI embeddings and high-resolution fMRI data. This allows it to generalize to new subjects and tasks, enabling in-silico experimentation to replicate classic neuroscience findings and reveal interpretable multisensory integration patterns.

Key Takeaways

  • 1
    TRIBE v2 is a tri-modal (video, audio, language) foundation model that accurately predicts high-resolution fMRI responses across diverse naturalistic and experimental conditions.
  • 2
    The model significantly outperforms traditional linear encoding models by leveraging a Transformer architecture to capture complex, non-linear relationships between AI embeddings and brain activity.
  • 3
    TRIBE v2 enables in-silico experimentation, allowing researchers to replicate classic neuroscientific findings and test hypotheses without human subjects, accelerating discovery.

Conceptual Flow

HIGH LEVEL
1
Methodology (The 'Logic')

The model takes in different types of information like videos, sounds, and words, processes them with smart AI tools, and then learns to guess how the brain will react.

Video Input
Sound Input
Text Input
AI Processes Data
Brain Activity Prediction
2
Results (The 'Impact')

The new model predicts brain activity much better than old methods, can guess for new people, and even helps scientists run pretend experiments on a computer to learn how the brain works.

Old Brain Models
New Brain Model
Compare Predictions
Better Accuracy
New Discoveries