Evaluating Large Language Models in Turkish Short Answer Scoring: Validity, Reliability, and Fairness Perspectives


Creative Commons License

Kara A., YILDIRIM S.

Sakarya University Journal of Computer and Information Sciences, cilt.9, sa.3, ss.980-994, 2026 (Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 9 Sayı: 3
  • Basım Tarihi: 2026
  • Doi Numarası: 10.35377/saucis...1835608
  • Dergi Adı: Sakarya University Journal of Computer and Information Sciences
  • Derginin Tarandığı İndeksler: Scopus, Applied Science & Technology Source, Central & Eastern European Academic Source (CEEAS), Directory of Open Access Journals, TR DİZİN (ULAKBİM)
  • Sayfa Sayıları: ss.980-994
  • Anahtar Kelimeler: Artificial intelligence in education, Automated scoring, Evaluation rubric, Large Language Models (LLM), Validity and fairness
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Atatürk Üniversitesi Adresli: Evet

Özet

This study examines the performance of large language models (LLMs) in Turkish short-answer assessments within the measurement and evaluation theory framework. The GPT, Gemini, Gemma, and LLaMA models were evaluated under zero-shot and one-shot conditions with rubric support. The results show that LLMs have high self-consistency, but decision reliability can vary depending on prompt format and example sensitivity. Formulating rubrics with clear and concrete performance indicators increases model-human alignment and assessment fairness. Furthermore, error analyses revealed that while models often exhibit systematic low-scoring tendencies, they also display an ‘inverted-U’ error pattern, showing higher reliability at score extremes but struggling significantly with evaluating partial knowledge (intermediate scores). The results indicate that LLMs can support teachers in formative assessment when properly structured rubrics are used, but ethical oversight and pedagogical responsibility remain indispensable in final decisions.