Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach

1Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Bias and stereotypes in language models can cause harm, especially in sensitive areas like content moderation and decision-making. This paper addresses bias and stereotype detection by exploring how jointly learning these tasks enhances model performance. We introduce StereoBias, a unique dataset labeled for bias and stereotype detection across five categories: religion, gender, socio-economic status, race, profession, and others, enabling a deeper study of their relationship. Our experiments compare encoder-only models and fine-tuned decoder-only models using QLoRA. While encoder-only models perform well, decoder-only models also show competitive results. Crucially, joint training on bias and stereotype detection significantly improves bias detection compared to training them separately. Additional experiments with sentiment analysis confirm that the improvements stem from the connection between bias and stereotypes, not multi-task learning alone. These findings highlight the value of leveraging stereotype information to build fairer and more effective AI systems.

Cite

CITATION STYLE

APA

Tomar, A., Murthy, R., & Bhattacharyya, P. (2025). Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 17304–17317). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.889

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free