Large language models in professionalism and ethical decision-making education: a transferable process for generating concordance of professional judgment items


İş-Kara T., KIYAK Y. S.

Academic medicine : journal of the Association of American Medical Colleges, cilt.101, sa.10, ss.1248-1252, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 101 Sayı: 10
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1093/acamed/wvag147
  • Dergi Adı: Academic medicine : journal of the Association of American Medical Colleges
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Educational research abstracts (ERA), EMBASE, MEDLINE, MLA - Modern Language Association Database, MLA International Bibliography
  • Sayfa Sayıları: ss.1248-1252
  • Anahtar Kelimeler: ethics, large language models, learning by concordance, medical education, professionalism
  • Gazi Üniversitesi Adresli: Evet

Özet

PROBLEM: Concordance-based tools, such as the Concordance of Professional Judgment (CoPJ) tool, allow students to practice ethical decision-making and professionalism under uncertain conditions, but developing high-quality CoPJ items (including writing vignettes and convening expert panels) is resource intensive. Thus, programs often lack the practical infrastructure needed to integrate the CoPJ into curricula at the scale necessary for repeated learner practice. Large language models (LLMs) may help address this burden, but a transferable process to generate CoPJ items aligned with local ethics frameworks and expert oversight is needed. APPROACH: This study implemented and evaluated a transferable process for generating CoPJ items using general purpose LLMs. Using a generic zero-shot prompt, ChatGPT-5 Thinking and Gemini 2.5 Pro were each prompted to generate 1 CoPJ item for all 28 topics in the Turkish Medical Association's ethical declarations (56 items total) in August 2025. OUTCOMES: Between August and September 2025, 2 faculty reviewers independently rated each LLM-generated item using a 20-criterion checklist and a 4-level global suitability judgment (accepted, minor revision, major revision, rejected), with adjudication as needed. Gemini 2.5 Pro achieved 96.4% overall global suitability, with 27/28 items judged to be suitable (accepted or minor revision). ChatGPT-5 Thinking's overall global suitability was 89.3% (25/28 items judged to be suitable). Gemini 2.5 Pro had an overall mean checklist success score of 97.5% (19.5/20), and ChatGPT-5 Thinking had a score of 89.0% (17.8/20). NEXT STEPS: Future work will examine how LLM-generated CoPJ items function in real learning environments, including a planned learner-based comparison with human-written items. Additional studies across languages, professional programs, and cultural contexts will further help determine the broader adaptability of this process.