Compression of Generative Pre-trained Language Models via Quantization

43Citations
Citations of this article
132Readers
Mendeley users who have this article in their library.

Abstract

The increasing size of generative Pre-trained Language Models (PLMs) have greatly increased the demand for model compression. Despite various methods to compress BERT or its variants, there are few attempts to compress generative PLMs, and the underlying difficulty remains unclear. In this paper, we compress generative PLMs by quantization. We find that previous quantization methods fail on generative tasks due to the homogeneous word embeddings caused by reduced capacity, and varied distribution of weights. Correspondingly, we propose a token-level contrastive distillation to learn distinguishable word embeddings, and a module-wise dynamic scaling to make quantizers adaptive to different modules. Empirical results on various tasks show that our proposed method outperforms the state-of-the-art compression methods on generative PLMs by a clear margin. With comparable performance with the full-precision models, we achieve 14.4× and 13.4× compression rates on GPT-2 and BART, respectively.

Cite

CITATION STYLE

APA

Tao, C., Hou, L., Zhang, W., Shang, L., Jiang, X., Liu, Q., … Wong, N. (2022). Compression of Generative Pre-trained Language Models via Quantization. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 4821–4836). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.acl-long.331

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free