Abstract
The existing generative Zero-Shot Learning (ZSL) methods only consider the unidirectional alignment from the class semantics to the visual features while ignoring the alignment from the visual features to the class semantics, which fails to construct the visual-semantic interactions well. In this paper, we propose to generate visual features based on an auto-encoder framework paired with multi-modality adversarial networks respectively for visual and semantic modalities to reinforce the visual-semantic interactions with a bidirectional alignment, which ensures the generated visual features to fit the real visual distribution and to be highly related to the semantics. The encoder aims at generating real-like visual features while the decoder forces both the real and the generated visual features to be more related to the class semantics. To further capture the discriminative information of the generated visual features, both the real and generated visual features are forced to be classified into the correct classes via a classification network. Experimental results on four benchmark datasets show that the proposed approach is particularly competitive on both the traditional ZSL and the generalized ZSL tasks.
Author supplied keywords
Cite
CITATION STYLE
Ji, Z., Dai, G., & Yu, Y. (2020). Multi-modality adversarial auto-encoder for zero-shot learning. IEEE Access, 8, 9287–9295. https://doi.org/10.1109/ACCESS.2019.2962298
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.