Behind the mask: Random and selective masking in transformer models applied to specialized social science texts

4Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Transformer models such as BERT and RoBERTa are increasingly popular in the social sciences to generate data through supervised text classification. These models can be further trained through Masked Language Modeling (MLM) to increase performance in specialized applications. MLM uses a default masking rate of 15 percent, and few works have investigated how different masking rates may affect performance. Importantly, there are no systematic tests on whether selectively masking certain words improves classifier accuracy. In this article, we further train a set of models to classify fake news around the coronavirus pandemic using 15, 25, 40, 60 and 80 percent random and selective masking. We find that a masking rate of 40 percent, both random and selective, improves within-category performance but has little impact on overall performance. This finding has important implications for scholars looking to build BERT and RoBERTa classifiers, especially those where one specific category is more relevant to their research.

Cite

CITATION STYLE

APA

Timoneda, J. C., & Vera, S. V. (2025). Behind the mask: Random and selective masking in transformer models applied to specialized social science texts. PLoS ONE, 20(2 February). https://doi.org/10.1371/journal.pone.0318421

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free