EAT: An Enhancer for Aesthetics-Oriented Transformers

41Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Transformers have shown great potential in various vision tasks, but none of them have surpassed the best CNN model on image aesthetics assessment (IAA) tasks. IAA is a challenging task in multimedia systems that requires attention to both foreground and background, as well as robustness to noisy and redundant labels. The global and dense attention mechanism of Transformers, designed for saliency-oriented tasks, may miss important aesthetic information in the background, increase the computational cost and slow down the convergence on IAA tasks. To address these issues, we propose an Enhancer for Aesthetics-Oriented Transformers (EAT). EAT uses a deformable, sparse and data-dependent attention mechanism that learns where to focus and how to refine attention by offsets. EAT also guides the offsets to balance the attention between foreground and background according to dedicated rules. Our EAT-enhanced Transformers outperform the previous methods on four representative datasets with fewer training epochs. Code is available in https://github.com/woshidandan/Image-Aesthetics-Assessment

Cite

CITATION STYLE

APA

He, S., Ming, A., Zheng, S., Zhong, H., & Ma, H. (2023). EAT: An Enhancer for Aesthetics-Oriented Transformers. In MM 2023 - Proceedings of the 31st ACM International Conference on Multimedia (pp. 1023–1032). Association for Computing Machinery, Inc. https://doi.org/10.1145/3581783.3611881

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free