SciGroveBeta
Medicine

The strength of clinical evidence is recoverable from language model representations but not from their stated grades

Soroosh Tayebi Arasteh

Featured July 16, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Even though smart computer models know how strong the evidence is for a medical claim deep inside their digital brains, they often can't tell us directly, creating a safety problem for doctors and patients.

In depth
Large language models (LLMs) internally encode the strength of clinical evidence for a claim, meaning how well it's supported (e.g., by many trials vs. expert opinion). This internal signal is linearly decodable from their hidden states, but surprisingly, the models fail to express this crucial information when asked directly, often stating grades at random. This creates a significant behavioral gap where critical safety information is present but uncommunicated.

Key Takeaways

  • 1
    LLMs internally encode evidence strength for clinical claims, which is linearly decodable from their hidden states.
  • 2
    Despite this internal signal, LLMs fail to verbalize the correct evidence grade when prompted, performing at chance levels.
  • 3
    The recoverable signal is largely lexical and does not generalize across different clinical topics or grading frameworks, yet it is distinct from factual truth.

Conceptual Flow

HIGH LEVEL
1
Methodology (The 'Logic')

The study collected many medical claims, figured out how strong their evidence was, and then checked if computer models secretly knew this evidence strength, even if they didn't say it out loud.

Medical Claims
Evidence Strength Labels
Train & Test
Computer Model's Internal Knowledge
Computer Model's Spoken Answer
2
Results (The 'Impact')

They found that computer models *do* know the evidence strength internally, but they *don't* say it correctly, meaning we need external tools to check their claims for safety.

Internal Knowledge (Good)
Spoken Answer (Bad)
Reveals Gap
Need External Check
Safer Medical AI