Abstract
We focus on developing a lightweight model for resource-constrained devices, building on MobileViT, a hybrid model that combines the strengths of Transformers and CNNs to balance high accuracy and computational efficiency for image classification. Transformers, while effective at capturing global information, often have higher computational costs than CNNs due to the complexity of their self-attention mechanism. To address this, we introduce the Token Merging (ToMe) technique into MobileViT to reduce computational costs. However, because the number of tokens changes during merging, ToMe cannot be directly applied to MobileViT without adjustments. We propose simple methods, specifically reshaping features and removing skip connections, to resolve this issue. Additionally, we make adjustments to MobileViT’s structure to better support the application of ToMe. Our approach improves inference efficiency while retaining a competitive level of accuracy. The resulting models achieve a balance between performance and computational speed, offering a practical solution for hybrid architectures. This work shows the potential of ToMe-based techniques to broaden the range of lightweight model options, catering to diverse application requirements.
Author supplied keywords
Cite
CITATION STYLE
Yasukura, M., Yoshioka, M., & Inoue, K. (2024). Reducing Computational Cost in MobileViT for Edge-Oriented Models Through Token Merging †. Electronics (Switzerland), 13(24). https://doi.org/10.3390/electronics13245009
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.