Tibetan Few-Shot Learning Model With Deep Contextualised Two-Level Word Embeddings

1Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

Few-shot learning is the task of identifying new text categories from a limited set of training examples. The two key challenges in few-shot learning are insufficient understanding of new samples and imperfect modelling. The uniqueness of low-resource languages lies in their limited linguistic resources, which directly leads to the difficulty for models to learn sufficiently rich feature representations from limited samples. As a minority language, Tibetan few-shot learning requires further exploration. With limited data resources, if the model's understanding of text is noncontextual, it cannot provide sufficiently distinctive feature representations, limiting its performance in few-shot learning. Therefore, this paper proposed a few-shot learning architecture called two-level word embeddings matching networks (TWE-MN). TWE-MN is specifically designed to enhance the model's representational capacity and optimise its generalisation capabilities in data-scarce environments. As this paper focuses on Tibetan few-shot learning tasks, a pretrained Tibetan language model, BoBERT, was constructed. BoBERT, as the pre-embedding layer of TWE-MN, in combination with the BoBERT-augmented full-context embedding, can capture feature information from local to global levels. This paper evaluated the performance of TWE-MN in Tibetan few-shot learning tasks and Tibetan text classification tasks. The experimental results show that TWE-MN outperformed vanilla MN in all Tibetan few-shot learning tasks, with an average accuracy improvement of 4.5%–6.5% and up to 6.8% at most. In addition, this paper also explores the potential of TWE-MN in other NLP tasks, such as text classification and machine translation.

Cite

CITATION STYLE

APA

Zhang, Z., Yu, Y., Wang, X., Feng, X., Li, Y., Shen, J., … Cai, J. (2025). Tibetan Few-Shot Learning Model With Deep Contextualised Two-Level Word Embeddings. CAAI Transactions on Intelligence Technology, 10(5), 1394–1410. https://doi.org/10.1049/cit2.70047

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free