Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings

3Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Detecting toxic content using language models is important but challenging. While large language models (LLMs) have demonstrated strong performance in understanding Chinese, recent studies show that simple character substitutions in toxic Chinese text can easily confuse the state-of-the-art (SOTA) LLMs. In this paper, we highlight the multimodal nature of Chinese language as a key challenge in deploying LLMs in toxic Chinese detection. First, we propose a taxonomy of 3 perturbation strategies and 8 specific approaches in toxic Chinese content. Then, we curate a dataset based on this taxonomy, and benchmark 9 SOTA LLMs (from both the US and China) to assess if they can detect perturbed toxic Chinese text. Additionally, we explore cost-effective enhancement solutions like in-context learning (ICL) and supervised fine-tuning (SFT). Our results reveal two important findings. (1) LLMs are less capable of detecting perturbed multimodal Chinese toxic contents. (2) ICL or SFT with a small number of perturbed examples may cause the LLMs to “overcorrect”: misidentify many normal Chinese contents as toxic.

Cite

CITATION STYLE

APA

Yang, S., Cui, S., Hu, C., Wang, H., Zhang, T., Huang, M., … Qiu, H. (2025). Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 14382–14396). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.742

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free