Multi-Modality Deep Network for Extreme Learned Image Compression

N/ACitations
Citations of this article
11Readers
Mendeley users who have this article in their library.

Abstract

Image-based single-modality compression learning approaches have demonstrated exceptionally powerful encoding and decoding capabilities in the past few years, but suffer from blur and severe semantics loss at extremely low bitrates. To address this issue, we propose a multimodal machine learning method for text-guided image compression, in which the semantic information of text is used as prior information to guide image compression for better compression performance. We fully study the role of text description in different components of the codec, and demonstrate its effectiveness. In addition, we adopt the image-text attention module and image-request complement module to better fuse image and text features, and propose an improved multimodal semantic-consistent loss to produce semantically complete reconstructions. Extensive experiments, including a user study, prove that our method can obtain visually pleasing results at extremely low bitrates, and achieves a comparable or even better performance than state-of-the-art methods, even though these methods are at 2× to 4× bitrates of ours.

Cite

CITATION STYLE

APA

Jiang, X., Tan, W., Tan, T., Yan, B., & Shen, L. (2023). Multi-Modality Deep Network for Extreme Learned Image Compression. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023 (Vol. 37, pp. 1033–1041). AAAI Press. https://doi.org/10.1609/aaai.v37i1.25184

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free