How Reliably Do Large Language Models Reproduce Vital Pulp Therapy Guidelines? A Mixed-Effects Evaluation of Guideline-Concordance and Error Directionality
Healthcare (Switzerland), vol.14, no.12, 2026 (SCI-Expanded, SSCI, Scopus)
- Publication Type: Article / Article
- Volume: 14 Issue: 12
- Publication Date: 2026
- Doi Number: 10.3390/healthcare14121605
- Journal Name: Healthcare (Switzerland)
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Social Sciences Citation Index (SSCI), Scopus, CINAHL, Health Research Premium Collection (ProQuest)
- Keywords: error directionality, guideline-concordance, large language models, mixed-effects logistic regression, prompt engineering, vital pulp therapy
- Gazi University Affiliated: Yes
Abstract
Background: Large language models (LLMs) are increasingly consulted for clinical guidance, yet their reliability in protocol-sensitive domains remains insufficiently characterized. This study evaluated the ability of widely accessible LLMs to reproduce guideline-defined decision thresholds in vital pulp therapy (VPT), with emphasis on guideline-concordance accuracy, professional-role prompting, short-term response stability, and decision-level error directionality. Methods: Twenty-six binary yes/no questions were derived from an internationally recognized evidence-based guideline for VPT. Four LLMs—GPT-5, GPT-4o, DeepSeek-V3, and Gemini 2.5 Flash—were queried under non-prompted and professional-role-prompted conditions by two independent operators across three daily sessions over three consecutive days. Descriptive analyses were complemented by mixed-effects logistic regression in R to account for repeated responses clustered within guideline-derived questions. Results: Overall guideline-concordance accuracy was high across models. Gemini showed the highest observed accuracy under non-prompted conditions; DeepSeek showed the highest under prompted conditions. In the mixed-effects model, Gemini demonstrated significantly higher odds of guideline-concordant responses than GPT-5 under non-prompted conditions, whereas DeepSeek outperformed GPT-5 and GPT-4o under prompted conditions. The model × prompt interaction showed a trend toward significance but did not reach the conventional threshold. Day and within-day time point were not significantly associated with accuracy, supporting short-term response stability. Error-direction analysis revealed model-specific patterns: Gemini showed consistently low false-positive rates but increased false-negative responses under prompted conditions; DeepSeek showed reduced false-positive and no false-negative responses under prompted conditions. Conclusions: Average accuracy alone is insufficient to characterize the reliability of LLM-generated clinical guidance. Evaluation in protocol-sensitive domains should incorporate guideline-concordance, prompt responsiveness, short-term stability, and decision-level error directionality.