Detecting independent pronoun bias with partially-synthetic data generation

9Citations
Citations of this article
81Readers
Mendeley users who have this article in their library.

Abstract

We report that state-of-the-art parsers consistently failed to identify “hers” and “theirs” as pronouns but identified the masculine equivalent “his”. We find that the same biases exist in recent language models like BERT. While some of the bias comes from known sources, like training data with gender imbalances, we find that the bias is amplified in the language models and that linguistic differences between English pronouns that are not inherently biased can become biases in some machine learning models. We introduce a new technique for measuring bias in models, using Bayesian approximations to generate partially-synthetic data from the model itself.

Cite

CITATION STYLE

APA

Monarch, R., & Morrison, A. (2020). Detecting independent pronoun bias with partially-synthetic data generation. In EMNLP 2020 - 2020 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference (pp. 2011–2017). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.emnlp-main.157

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free