Generalized entropy regularization or: There's nothing special about label smoothing

Clara Meister; Elizabeth Salesky; Ryan Cotterell

Conference ProceedingsOPEN ACCESS

Generalized entropy regularization or: There's nothing special about label smoothing

Proceedings of the Annual Meeting of the Association for Computational Linguistics (2020) 6870-6886

DOI: 10.18653/v1/2020.acl-main.615

43Citations

150Readers

Abstract

Prior work has explored directly regularizing the output distributions of probabilistic models to alleviate peaky (i.e. over-confident) predictions, a common sign of overfitting. This class of techniques, of which label smoothing is one, has a connection to entropy regularization. Despite the consistent success of label smoothing across architectures and datasets in language generation tasks, two problems remain open: (1) there is little understanding of the underlying effects entropy regularizers have on models, and (2) the full space of entropy regularization techniques is largely unexplored. We introduce a parametric family of entropy regularizers, which includes label smoothing as a special case, and use it to gain a better understanding of the relationship between the entropy of a trained model and its performance on language generation tasks. We also find that variance in model performance can be explained largely by the resulting entropy of the model. Lastly, we find that label smoothing provably does not allow for sparse distributions, an undesirable property for language generation models, and therefore advise the use of other entropy regularization methods in its place. Our code is available online at https://github.com/rycolab/entropyRegularization.

Cite

CITATION STYLE

APA

Meister, C., Salesky, E., & Cotterell, R. (2020). Generalized entropy regularization or: There’s nothing special about label smoothing. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 6870–6886). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.acl-main.615

Generalized entropy regularization or: There's nothing special about label smoothing

Abstract

Cite

Register to see more suggestions