Hybrid model integrating LeViT transformer and distillation techniques for pattern detection and dance classification

0Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Artificial intelligence (AI) has become an integral part of modern life, extending its impact into the preservation of cultural heritage. This study applies state-of-the-art vision transformer models for the classification of traditional Chinese dance types. The main aim is to introduce a lightweight Vision Transformer (LeViT) architecture integrated with distillation mechanism to capture complex cues and visual patterns from images. It is evaluated alongside the standard Vision Transformer (ViT) and a baseline Convolutional Neural Network (CNN) model, which remains a benchmark in computer vision research. The dataset comprises images of three Chinese dance forms that follows standard preprocessing techniques including noise removal, geometric augmentation, and contrast enhancement. After feature extraction and supervised training, the proposed LeViT model demonstrated superior performance in terms of accuracy and stability compared with ViT and CNN. The findings confirm that transformer-based architectures, particularly LeViT, effectively capture subtle motion cues and fine-grained spatial patterns within dance imagery, offering a promising framework for future research in AI-driven cultural preservation and visual arts classification.

Cite

CITATION STYLE

APA

Wang, Y. (2026). Hybrid model integrating LeViT transformer and distillation techniques for pattern detection and dance classification. Scientific Reports, 16(1). https://doi.org/10.1038/s41598-025-26035-8

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free