Abstract
Educational chatbots are increasingly deployed to scaffold student learning, yet educators lack scalable ways to assess the cognitive depth of these dialogues in situ. Bloom’s taxonomy provides a principled lens for characterizing reasoning, but manual tagging of conversational turns is costly and difficult to scale for learning analytics. We present a reproducible high-confidence pseudo-labeling pipeline for multi-label Bloom classification of Socratic student–chatbot exchanges. The dataset comprises 6716 utterances collected from conversations between a Socratic chatbot and 34 undergraduate statistics students at Nanyang Technological University. From three chronologically selected workbooks with expert Bloom annotations, we trained and compared two labeling tracks: (i) a calibrated classical approach using SentenceTransformer (all-MiniLM-L6-v2) embeddings with one-vs-rest Logistic Regression, Linear SVM, XGBoost, and MLP, followed by per-class precision–recall threshold tuning; and (ii) a lightweight LLM track using GPT-4o-mini after supervised fine-tuning. Class-specific thresholds tuned on 5-fold cross-validation were then applied in a single pass to assign high-confidence pseudo-labels to the remaining unlabeled exchanges, avoiding feedback-loop confirmation bias. Fine-tuned GPT-4o-mini achieved the highest prevalence-weighted performance (micro-F1 (Formula presented.)), whereas calibrated classical models yielded stronger balance across Bloom levels (best macro-F1 (Formula presented.) with Linear SVM; best classical micro-F1 (Formula presented.) with Logistic Regression). Both model families reflect the corpus skew toward lower-order cognition, with LLMs excelling on common patterns and linear models better preserving rarer higher-order labels, while results should be interpreted as a proof-of-concept given limited gold labeling, the approach substantially reduces annotation burden and provides a practical pathway for Bloom-aware learning analytics and future real-time adaptive chatbot support.
Author supplied keywords
Cite
CITATION STYLE
Lee, K. W., Ang, Y. S., & Lai, J. W. (2025). Learning Analytics with Scalable Bloom’s Taxonomy Labeling of Socratic Chatbot Dialogues. Computers, 14(12). https://doi.org/10.3390/computers14120555
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.