Add-Vit: CNN-Transformer Hybrid Architecture for Small Data Paradigm Processing

N/ACitations
Citations of this article
22Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

The vision transformer(ViT), pre-trained on large datasets, outperforms convolutional neural networks (CNN) in computer vision(CV). However, if not pre-trained, the transformer architecture doesn’t work well on small datasets and is surpassed by CNN. Through analysis, we found that:(1) the division and processing of tokens in the ViT discard the marginalized information between token. (2) the isolated multi-head self-attention (MSA) lacks prior knowledge. (3) the local inductive bias capability of stacked transformer block is much inferior to that of CNN. We propose a novel architecture for small data paradigms without pre-training, named Add-Vit, which uses progressive tokenization with feature supplementation in patch embedding. The model’s representational ability is enhanced by using a convolutional prediction module shortcut to connect MSA and capture local features as additional representations of the token. Without the need for pre-training on large datasets, our best model achieved 81.25% accuracy when trained from scratch on the CIFAR-100.

Cite

CITATION STYLE

APA

Chen, J., Wu, P., Zhang, X., Xu, R., & Liang, J. (2024). Add-Vit: CNN-Transformer Hybrid Architecture for Small Data Paradigm Processing. Neural Processing Letters, 56(3). https://doi.org/10.1007/s11063-024-11643-8

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free