Abstract
Interactive visual storytelling through conversational agents offers a means to enhance early childhood language learning. We developed a picture-guided conversational chatbot, driven by dense image captioning, for early childhood mother tongue language learning. However, state-of-the-art image captioning systems fall short in meeting the educational needs of young learners. They often lack cultural contextualization, use vocabulary that exceeds children’s developmental level, and fail to align with curriculum-relevant learning goals. We investigated a contextualized dense image captioning framework, which augments dense image captioning with cultural and curriculum-aligned keyword retrieval through a Retrieval-Augmented Generation (RAG) module. This enables the generation of culturally appropriate, age-level suitable, and educationally anchored captions that enhance learner engagement and pedagogical relevance. We demonstrate that our approach outperforms existing captioning models in terms of linguistic appropriateness, and curriculum and cultural alignment. The contextualized dense image captioning framework supports the development of culturally grounded, education-oriented conversational agents for young learners.
Author supplied keywords
Cite
CITATION STYLE
Tan, H. L., Gu, Y., Li, L., Leong, M. C., & Chen, N. F. (2025). Contextualized Visual Storytelling for Conversational Chatbot in Education. In ICMI 2025 - Companion Publication of the 27th International Conference on Multimodal Interaction (pp. 185–189). Association for Computing Machinery, Inc. https://doi.org/10.1145/3747327.3764895
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.