The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts
Ankara University Journal of Faculty of Educational Sciences, cilt.59, sa.2, ss.549-590, 2026 (Hakemli Dergi)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 59 Sayı: 2
- Basım Tarihi: 2026
- Doi Numarası: 10.30964/auebfd.1807153
- Dergi Adı: Ankara University Journal of Faculty of Educational Sciences
- Sayfa Sayıları: ss.549-590
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- Gazi Üniversitesi Adresli: Evet
Özet
This study investigated whether artificial intelligence (AI) systems could evaluate the content validity of B1-level English reading comprehension items in a manner comparable to human experts. A 25-item multiple-choice test was developed and rated by four human experts and four AI models (GPT-4o, Gemini 2.0, Replika, and Chatfuel). To ensure methodological transparency, a standardized few-shot prompting framework was employed. Each model was accessed in a new session to avoid context drift, and identical prompts were provided to maintain consistency across evaluations. The Content Validity Ratio (CVR) and Item Content Validity Index (I-CVI) were calculated and compared using the Wilcoxon Signed-Rank Test. The results revealed no statistically significant difference between AI- and human-generated scores, suggesting that AI evaluators produced ratings similar to those of human experts. Although fine-tuning could not be performed due to the closed-source nature of the models, standardized prompting procedures enhanced reliability and replicability. These findings indicate that AI tools can complement expert judgment in content validity assessment when methodological control is ensured. The study highlights the potential of hybrid human–AI frameworks in educational measurement, while emphasizing the importance of model transparency and context management.