Boundary and Contextual Structure Extraction for Named Entity Recognition

0Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Named Entity Recognition (NER) models trained on small datasets often struggle to identify new entities, particularly in Thai, which presents unique linguistic challenges. These challenges include unclear word boundaries due to the absence of explicit delimiters, complex contextual structures that alter word meanings based on surrounding context, and limited annotated resources, such as insufficient datasets containing officially recognized named entities. As a result, NER models frequently exhibit poor generalization, classification errors, and reduced information extraction performance, especially when encountering newly coined names of persons, organizations, or locations. To address these challenges, this research proposes a novel framework for developing Thai NER models that leverages contextual learning and sentence structure analysis to enhance entity recognition, particularly for unseen entities. The proposed approach begins by identifying entity features from dictionaries or databases, including word components, root words, and structural patterns, to improve entity recognition and expand lexical coverage. Additionally, the Boundary Contextual Structures Extraction Framework is applied to analyze sentence structures and extract clear contextual features of entities, supporting the prediction of unseen entities by relying on word positioning patterns, relationships with neighboring words, and distinctive entity characteristics that may emerge in the future. Furthermore, accurate word segmentation and contextual learning techniques are employed to mitigate issues related to unclear word boundaries in Thai. At the core of this framework is a Transformer-based model that utilizes the Attention mechanism to capture long-distance word dependencies and contextual meanings. The Transformer efficiently processes word sequences, while the Attention mechanism emphasizes words or parts of sentences crucial for model decisions, thereby improving entity boundary detection and classification accuracy. To evaluate the model’s performance, a new Thai dataset of 5,000 sentences published in 2024 containing emerging entities was constructed. Experimental results demonstrate that the model accurately extracts entities that have never appeared in the training data, achieving an accuracy rate of 83% for Person, 62% for Location, and 81% for Organization. These findings highlight the model’s strong capacity for identifying new entities, particularly in the categories of person and organization names. In conclusion, the proposed model significantly enhances the recognition of new entities and overall Thai NER performance, demonstrating its effectiveness in overcoming linguistic and data limitations and adapting efficiently to evolving language usage.

Cite

CITATION STYLE

APA

Kuntichod, W., Lee, W., Srisomboon, K., Pipanmekaporn, L., & Prayote, A. (2025). Boundary and Contextual Structure Extraction for Named Entity Recognition. IEEE Access, 13, 187785–187798. https://doi.org/10.1109/ACCESS.2025.3626034

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free