Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model
BAYBURT EĞİTİM FAKÜLTESİ DERGİSİ, cilt.21, sa.51, ss.1187-1210, 2026 (TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 21 Sayı: 51
- Basım Tarihi: 2026
- Doi Numarası: 10.35675/befdergi.1788625
- Dergi Adı: BAYBURT EĞİTİM FAKÜLTESİ DERGİSİ
- Derginin Tarandığı İndeksler: TR DİZİN (ULAKBİM)
- Sayfa Sayıları: ss.1187-1210
- Gazi Üniversitesi Adresli: Evet
Özet
This study examines the inter-rater reliability (IRR) of three AI tools (ChatGPT 3.5, Gemini, YouChat) and four human raters in evaluating English writing skills using the multifaceted Rasch model. The evaluation focuses on the consistency and dependability of scores assigned by the AI tools compared to human judgment. The sample comprised 206 students whose responses were evaluated using a holistic scoring rubric based on the Common European Framework of Reference for Languages (CEFR) writing competencies. Findings indicate that ChatGPT 3.5 outperforms Gemini and YouChat regarding accuracy, precision, and recall, with an overall high IRR among all raters. The results suggest that AI tools can effectively supplement human raters, providing consistent and reliable assessments, which has significant implications for educational evaluations and the potential integration of AI in various scoring scenarios. The study contributes to understanding the role of AI in educational assessment and emphasizes the potential of AI tools to improve the scalability and efficiency of scoring processes. The study also discusses the importance of AI tools in providing prompt feedback, analyzing large datasets, and ensuring scoring consistency across different contexts. It also highlights the need for hybrid models that combine AI's consistency with human assessors' contextual understanding.