Abstract
Self-supervised learning (SSL) methods such as Word2vec, BERT, and GPT have shown great effectiveness in language understanding. Contrastive learning, as a recent SSL approach, has attracted increasing attention in NLP. Con-trastive learning learns data representations by predicting whether two augmented data in-stances are generated from the same original data example. Previous contrastive learning methods perform data augmentation and con-trastive learning separately. As a result, the augmented data may not be optimal for con-trastive learning. To address this problem, we propose a four-level optimization framework that performs data augmentation and contras-tive learning end-to-end, to enable the augmented data to be tailored to the contrastive learning task. This framework consists of four learning stages, including training machine translation models for sentence augmentation, pretraining a text encoder using contrastive learning, finetuning a text classification model, and updating weights of translation data by minimizing the validation loss of the classification model, which are performed in a unified way. Experiments on datasets in the GLUE benchmark (Wang et al., 2018a) and on da-tasets used in Gururangan et al. (2020) dem-onstrate the effectiveness of our method.
Cite
CITATION STYLE
Fang, H., & Xie, P. (2022). An End-to-End Contrastive Self-Supervised Learning Framework for Language Understanding. Transactions of the Association for Computational Linguistics, 10, 1324–1340. https://doi.org/10.1162/tacl_a_00521
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.