Abstract
In this paper, we present an approach for the multi‐label classification of remote sensing images based on data‐efficient transformers. During the training phase, we generated a second view for each image from the training set using data augmentation. Then, both the image and its augmented version were reshaped into a sequence of flattened patches and then fed to the transformer encoder. The latter extracts a compact feature representation from each image with the help of a self‐attention mechanism, which can handle the global dependencies between different regions of the high‐resolution aerial image. On the top of the encoder, we mounted two classifiers, a token and a distiller classifier. During training, we minimized a global loss consisting of two terms, each corresponding to one of the two classifiers. In the test phase, we considered the average of the two classifiers as the final class labels. Experiments on two datasets acquired over the cities of Trento and Civezzano with a ground resolution of two‐centimeter demonstrated the effectiveness of the proposed model.
Author supplied keywords
Cite
CITATION STYLE
Bashmal, L., Bazi, Y., Al Rahhal, M. M., Alhichri, H., & Ajlan, N. A. (2021). UAV image multi‐labeling with data‐efficient transformers. Applied Sciences (Switzerland), 11(9). https://doi.org/10.3390/app11093974
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.