Decoder-Only LLMs can be Masked Auto-Encoders

1Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Modern NLP workflows (e.g., RAG systems) require different models for generation and embedding tasks, where bidirectional pre-trained encoders and decoder-only Large Language Models (LLMs) dominate respective tasks. Structural differences between models result in extra development costs and limit knowledge sharing between tasks. In this work, we present UniMAE, a novel unsupervised training method that transforms a Decoder-Only LLM into a Uni-Directional Masked Auto-Encoder. UniMAE compresses high-quality semantic information into the [EOS] embedding while preserving the generation capabilities of LLMs. Comprehensive evaluations across 56 MTEB datasets demonstrate that UniMAE can achieve state-of-the-art results under unsupervised settings with merely 100 training steps, establishing the first effective approach to unifying generation and representation learning in decoder-only architectures.

Cite

CITATION STYLE

APA

Qiao, D., Gao, Y., Yang, Z., Yang, D., Wu, Z., Lu, P., … Zhang, M. (2025). Decoder-Only LLMs can be Masked Auto-Encoders. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 2, pp. 713–723). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-short.57

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free