SciGroveBeta
Machine Learning

Verbalizable Representations Form a Global Workspace in Language Models

Jack Lindsey, Newman Cheng, Sreejan Kumar, Ishita Dasgupta

Featured July 28, 2026

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Scientists found that large language models have a special "thinking space" called the J-space where they hold ideas they could "say" if asked, helping them reason and even learn to be more ethical.

In depth
The paper introduces the Jacobian lens (J-lens), a novel interpretability technique that identifies internal representations in large language models (LLMs) that are "verbalizable"—meaning the model is poised to articulate them. These verbalizable representations, collectively termed the J-space, exhibit functional properties analogous to a global workspace in human cognition, enabling flexible reasoning and report. The authors demonstrate that the J-space mediates complex internal computations and can be shaped through counterfactual reflection training to improve model alignment.

Key Takeaways

  • 1
    The Jacobian lens is a new interpretability method that identifies "verbalizable" internal representations in LLMs by analyzing the average causal effect of activations on future token probabilities.
  • 2
    These verbalizable representations form a J-space within the model that functions like a global workspace, supporting flexible reasoning, directed modulation, and explicit verbal report.
  • 3
    Counterfactual reflection training can shape the J-space to instill ethical principles, improving model behavior in original contexts without direct supervision on target actions.

Conceptual Flow

HIGH LEVEL
1
Uncovering a Model's Inner Thoughts

The paper uses a special "Jacobian lens" to peek inside a language model and see what ideas it's silently thinking about, even if it doesn't say them out loud.

Model's Internal State
Apply Special Lens
List of Silent Thoughts
2
A 'Thinking Space' for Reasoning and Ethics

They found that the model's silent thoughts act like a human's "global workspace," helping it reason flexibly, and they can even teach it to think more ethically.

Model's Silent Thoughts
Enables Flexible Reasoning
Better Decisions
Ethical Behavior

This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.