SciGroveBeta
Medicine

Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

Nisarg A. Patel

Featured July 9, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Even though smart computer programs can guess illnesses correctly, they don't think about similar cases in the same structured way a doctor would, showing competence without consistent reasoning patterns.

In depth
The paper introduces clinical reasoning graphs, a novel structured representation extracted from large language model (LLM) diagnostic traces using a domain-grounded ontology. This allows for a process-level evaluation of LLM reasoning beyond mere accuracy. The study reveals that despite achieving high diagnostic accuracy, LLMs do not exhibit schema-scale cross-case consistency in their reasoning patterns for clinically similar cases, suggesting competence without consistent, reusable diagnostic schemas.

Key Takeaways

  • 1
    The authors developed a clinical reasoning ontology and an extraction pipeline to convert unstructured LLM diagnostic traces into computable graph representations.
  • 2
    LLMs demonstrate diagnostic competence (high accuracy) but lack schema-scale consistency in their reasoning graphs across clinically similar cases, unlike human experts.
  • 3
    The study found that diagnostic accuracy and the structure of reasoning graphs capture different evaluation dimensions, highlighting the need for process-level analysis in clinical AI.

Conceptual Flow

HIGH LEVEL
1
Methodology: Turning Text into Structured Thoughts

The paper turns messy computer text about medical cases into neat diagrams, like drawing a map of how the computer thought.

Computer's Text Explanation
Extract Key Ideas
Structured Thinking Map
2
Results: Smart Answers, Unpredictable Thinking

They found that even when computers get the right answer, their thinking maps for similar problems look very different, unlike how a human expert would think.

Similar Medical Cases
Compare Thinking Maps
Answers are Right
Thinking is Different