Contextualized Visual Storytelling for Conversational Chatbot in Education

2Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Interactive visual storytelling through conversational agents offers a means to enhance early childhood language learning. We developed a picture-guided conversational chatbot, driven by dense image captioning, for early childhood mother tongue language learning. However, state-of-the-art image captioning systems fall short in meeting the educational needs of young learners. They often lack cultural contextualization, use vocabulary that exceeds children’s developmental level, and fail to align with curriculum-relevant learning goals. We investigated a contextualized dense image captioning framework, which augments dense image captioning with cultural and curriculum-aligned keyword retrieval through a Retrieval-Augmented Generation (RAG) module. This enables the generation of culturally appropriate, age-level suitable, and educationally anchored captions that enhance learner engagement and pedagogical relevance. We demonstrate that our approach outperforms existing captioning models in terms of linguistic appropriateness, and curriculum and cultural alignment. The contextualized dense image captioning framework supports the development of culturally grounded, education-oriented conversational agents for young learners.

Cite

CITATION STYLE

APA

Tan, H. L., Gu, Y., Li, L., Leong, M. C., & Chen, N. F. (2025). Contextualized Visual Storytelling for Conversational Chatbot in Education. In ICMI 2025 - Companion Publication of the 27th International Conference on Multimodal Interaction (pp. 185–189). Association for Computing Machinery, Inc. https://doi.org/10.1145/3747327.3764895

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free