Abstract
The process of associating elements of learning content with concepts or skills that this content helps students to master is one of the critical steps in developing personalized educational systems. When these associations are properly established, the system can infer the growth of student understanding of separate knowledge components from the logs of their interactions with associated learning content and use it to adapt the learning process accordingly by targeting gaps in individual students’ knowledge. Unfortunately, crafting these links between learning content and knowledge components is a very time- and expertise-demanding process that has traditionally been performed manually by domain experts with the help of knowledge engineers. Recently, the power of Large Language Models has motivated a new generation of research on concept extraction from textual learning content. The work presented in this paper contributes to this trend while introducing two important innovations. First, our concept extraction process is guided by a human-authored ontology of the target domain - Python programming. Second, alongside a traditional expert evaluation of the concept extraction quality, we apply two additional validation approaches: one based on using an educational data mining technique (learning curves) and another utilizing the pedagogical expertise of teaching the target domain (learning content placement).
Author supplied keywords
Cite
CITATION STYLE
Hendrawan, R. A., de Alencar, R. S., Micheli, A., Brusilovsky, P., Barria-Pineda, J., & Sosnovsky, S. (2026). Data-Driven Evaluation of LLM-Based Ontology Concept Extraction from Programming Learning Content. In 16th International Learning Analytics and Knowledge Conference, LAK 2026 (pp. 526–535). Association for Computing Machinery, Inc. https://doi.org/10.1145/3785022.3785106
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.