Abstract
Large vision-language models (LVLMs) have achieved significant progress and been widely applied across various industries. However, a series of significant challenges have also been posed, such as the hallucination, where the generated attributes of a target are inconsistent with the information in the image. Recent works have primarily focused on improving generation performance by training the multimodal alignment module of the model with additional data. However, this approach incurs higher costs and introduces redundant information due to the reliance on large amounts of additional data. To address these challenges, we propose a novel multimodal mitigation framework, termed MH-PEFT. Specifically, we introduce an effective data augmentation strategy that leverages existing data to generate additional attribute-related information, thereby enhancing the model’s generative capabilities. Furthermore, we develop a PEFT-based training approach specifically tailored to mitigate multimodal hallucination in LVLMs. Notably, our framework features a plug-and-play design, which enhances its robustness and facilitates its deployment in real-world scenarios. Extensive experiments demonstrate that our method achieves state-of-the-art performance across multiple benchmark datasets.
Author supplied keywords
Cite
CITATION STYLE
Li, F. (2025). MH-PEFT: Mitigating Hallucinations in Large Vision-Language Models through the PEFT Method. In Proceedings of the 2025 2nd International Conference on Generative Artificial Intelligence and Information Security, GAIIS 2025 (pp. 137–142). Association for Computing Machinery, Inc. https://doi.org/10.1145/3728725.3728747
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.