Thai Finger-Spelling using Vision Transformer

2Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

In this paper, we present a finger-spelling recognition system that is based on Thai Sign Language (TFS) and employs a deep learning model called vision transformer. We extracted the 15 characters of the Thai alphabet from publicly available and our collected datasets to establish the recognition system. To train the learning model, we employed four EVA-02 vision transformer models, each of which showed impressive performance across different model sizes. We conducted four experiments to determine the most effective performance model. In Experiment 1, we directly trained the model to compare its performance. In Experiment 2, we used augmentation techniques to generate additional datasets. Experiment 3 utilized the Test-Time Augmentation (TTA) technique to generate test images with random variations. Lastly, in Experiment 4, we used Pseudo-Labelling (labeling labeled and unlabeled data) in each batch to train the model network. Furthermore, we developed a mobile application that collects user image data and provides helpful information related to finger-spelling, such as meanings, gestures, and usage examples.

Cite

CITATION STYLE

APA

Chaowanawatee, K., Silanon, K., & Kliangsuwan, T. (2023). Thai Finger-Spelling using Vision Transformer. International Journal of Advanced Computer Science and Applications, 14(11), 348–353. https://doi.org/10.14569/IJACSA.2023.0141135

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free