SciGroveBeta
Medicine

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

Mahdi Alkaeed Khalaf

Featured June 26, 2026

This analysis was generated by SciGrove. Upload your own PDFs or enter a DOI — and get the same AI breakdown on any paper.

Get started

AI-generated analysis — This is SciGrove's AI interpretation of the paper, not peer-reviewed content. Always refer to the original paper.

Simply

Even tiny changes in how you ask a medical question can make smart AI models give different or wrong answers, showing they aren't always reliable for important healthcare decisions.

In depth
The paper systematically evaluates the robustness of Large Language Models (LLMs) in healthcare to minor input changes, specifically lexical and syntactic prompt perturbations. It reveals that both general-purpose and medical-specific LLMs are highly sensitive, with small rephrasing or reordering of prompts leading to altered clinical advice, diagnostic variability, and even harmful outputs, underscoring critical safety concerns for real-world deployment.

Key Takeaways

  • 1
    Medical LLMs are not intrinsically safe; minor prompt variations can significantly alter clinical advice and lead to diagnostic variability.
  • 2
    The study categorizes and evaluates lexical and syntactic perturbations, showing that syntactic changes often cause more severe performance degradation than simple lexical substitutions.
  • 3
    A systematic sensitivity analysis framework is proposed, utilizing tools like BioSyn and Universal Sentence Encoder (USE) to quantify robustness and identify critical vulnerabilities in LLMs for healthcare.

Conceptual Flow

HIGH LEVEL
1
Methodology: How LLM Robustness is Evaluated

The study checks if medical AI models give consistent answers even when questions are slightly rephrased, using special tools to measure how much the meaning changes.

Original Medical Question
Slightly Change Words or Order
Changed Medical Questions
AI Model Answers
2
Results: Impact of Prompt Changes on Medical LLMs

They found that even small changes to questions can make medical AI models give wrong diagnoses or ignore important instructions, showing they are not yet safe for hospitals.

Slightly Changed Questions
AI Model Gives Answers
Correct Answer
Wrong Answer
Ignored Instructions