SciGroveBeta
Machine Learning

Verbalizable Representations Form a Global Workspace in Language Models

Jack Lindsey, Newman Cheng, Sreejan Kumar, Ishita Dasgupta

Featured July 28, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Scientists found that large language models have a special "thinking space" called the J-space where they hold ideas they could "say" if asked, helping them reason and even learn to be more ethical.

In depth
The paper introduces the Jacobian lens (J-lens), a novel interpretability technique that identifies internal representations in large language models (LLMs) that are "verbalizable"—meaning the model is poised to articulate them. These verbalizable representations, collectively termed the J-space, exhibit functional properties analogous to a global workspace in human cognition, enabling flexible reasoning and report. The authors demonstrate that the J-space mediates complex internal computations and can be shaped through counterfactual reflection training to improve model alignment.

Key Takeaways

  • 1
    The Jacobian lens is a new interpretability method that identifies "verbalizable" internal representations in LLMs by analyzing the average causal effect of activations on future token probabilities.
  • 2
    These verbalizable representations form a J-space within the model that functions like a global workspace, supporting flexible reasoning, directed modulation, and explicit verbal report.
  • 3
    Counterfactual reflection training can shape the J-space to instill ethical principles, improving model behavior in original contexts without direct supervision on target actions.

Conceptual Flow

HIGH LEVEL
1
Uncovering a Model's Inner Thoughts

The paper uses a special "Jacobian lens" to peek inside a language model and see what ideas it's silently thinking about, even if it doesn't say them out loud.

Model's Internal State
Apply Special Lens
List of Silent Thoughts
2
A 'Thinking Space' for Reasoning and Ethics

They found that the model's silent thoughts act like a human's "global workspace," helping it reason flexibly, and they can even teach it to think more ethically.

Model's Silent Thoughts
Enables Flexible Reasoning
Better Decisions
Ethical Behavior