Abstract
The integration of Artificial Intelligence (AI) in language assessment, particularly in evaluating speaking skills, has introduced opportunities for greater consistency, efficiency, and scalability in educational contexts. This paper studiesthe reliability of AI-assisted speaking assessment compared to human-mediated evaluation, with a focus on inter-rater and intra-rater reliability in English as a Foreign Language (EFL) learning. Thispaper explores the strengths and limitations of AI in automated scoring, such as its capacity for standardization, alongside challenges related to validity, bias, and interpretability of results. This study reviews discrepancies between human and AI scoring due to subjective judgment and training limitations. The study emphasizes the need for standardized rubrics, rater training, and AI model calibration to enhance reliability. This paperconcludes by proposing a hybrid assessment framework in which AI complements human raters, supported by methodological and technical improvements in speech recognition and natural language processing. This approach aims to optimize speaking proficiency evaluations while maintaining fairness and educational integrity.
Cite
CITATION STYLE
Tleshova, Z., Tusselbayeva, Z., Ichshanova, A., Urazbekova, A., Zhenisbayeva, M., & Orymbayev, A. (2025). RELIABILITY OF AI IN FOREIGN LANGUAGE SPEAKING ASSESSMENT: COMPARING AUTOMATED AND HUMAN SCORING AMONG UNDERGRADUATE IT STUDENTS IN KAZAKHSTAN. National Center for Higher Education Development, 2(50). https://doi.org/10.59787/2413-5488-2025-50-2-18-31
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.