Sanjay Basu
Featured June 21, 2026
AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.
Big AI models struggle to answer complex medical questions from patient records if they need to connect many pieces of information, and even asking them to "think step-by-step" doesn't fix this multi-step reasoning problem.
The study measured how well AI models answered medical questions by counting how many thinking steps each question needed, from simple facts to complex summaries.
They found that the more thinking steps a question needed, the worse the AI models performed, showing a clear limit to their ability to reason deeply.
This breakdown was generated by SciGrove. Get the same analysis — intuition, storyboard, peer review, a runnable prototype and a glossary — on any paper you upload or paste a DOI for.