Urdu Word Sense Disambiguation: Leveraging Contextual Stacked Embedding, Siamese Transformer Encoder 1DCNN-BiLSTM, and Gloss Data Augmentation

N/ACitations
Citations of this article
20Readers
Mendeley users who have this article in their library.

Abstract

Word Sense Disambiguation (WSD) in Natural Language Processing (NLP) is crucial for discerning the correct meaning of words with multiple senses in various contexts. Recent advancements in this field, particularly Deep Learning (DL) and sophisticated language models such as BERT and GPT, have significantly improved WSD performance. However, challenges persist, especially with languages like Urdu, which are known for their linguistic complexity and limited digital resources compared to English. This study addresses the challenge of advancing WSD in Urdu by developing and applying tailored Data Augmentation (DA) techniques. We introduce an innovative approach, Prompt Engineering with Retrieval Augmented Generation (RAG), leveraging GPT-3.5-turbo to generate context-sensitive Gloss Definitions (GD). Additionally, we employ sentence-level and word-level DA techniques, including Back Translation (BT) and Masked Word Prediction (MWP). To enhance sentence understanding, we combine three BERT embedding models: mBERT, mDistilBERT, and Roberta_Urdu, facilitating a more nuanced comprehension of sentences and improving word disambiguation in complex linguistic contexts. Furthermore, we propose a novel network architecture merging Transformer Encoder (TE)-CNN and TE-BiLSTM models with Multi-Head Self-Attention (MHSA), One-Dimensional Convolutional Neural Network (1DCNN), and Bidirectional Long Short-Term Memory (BiLSTM). This architecture is tailored to address polysemy and capture short- and long-range dependencies critical for effective WSD in Urdu. Empirical evaluations on Lexical Sample (LS) and All Word (AW) tasks demonstrate the effectiveness of our approach, achieving an 88.9% F1 score on the LS and a 79.2% F1 score on AW tasks. These results underscore the importance of language-specific approaches and the potential of DA and advanced modeling techniques in overcoming challenges associated with WSD in languages with limited resources.

Cite

CITATION STYLE

APA

Ahmed, A., Huang, D., Arafat, S. Y., & Rashid, K. I. (2025). Urdu Word Sense Disambiguation: Leveraging Contextual Stacked Embedding, Siamese Transformer Encoder 1DCNN-BiLSTM, and Gloss Data Augmentation. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(5). https://doi.org/10.1145/3719293

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free