Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

1Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

With the growing size of Large Vision-Language Models (LVLMs), network pruning techniques designed to compress these models for deployment in resource-constrained environments have attracted significant attention. However, we observe that pruning frequently results in a degradation in safety performance. To address this issue, we propose a novel and lightweight approach, named Hierarchical Safety Realignment (HSR). HSR operates by first quantifying the contribution of each attention head to safety, identifying the most critical ones, and then selectively restoring neurons directly within these attention heads that play a pivotal role in maintaining safety. This process hierarchically realigns the safety of pruned LVLMs, progressing from the attention head level to the neuron level. We validate HSR across various models and pruning strategies, consistently achieving notable improvements in safety performance. To the best of our knowledge, this is the first work explicitly focused on restoring safety in LVLMs post-pruning. The code will be available at https://github.com/TheShineyue/HSR.

Cite

CITATION STYLE

APA

Li, Y., Yi, X., Shi, D., de Melo, G., Wang, X., & Wang, L. (2025). Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 7600–7612). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.394

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free