Abstract
Finetuning serves as the critical adaptation mechanism for multimodal large language models, bridging their pretrained knowledge with specialized downstream task requirements. This paper reviews recent finetuning advances across three key dimensions: (1) efficiency-oriented methods that reduce resource costs; (2) capability-specific techniques enhancing specialized multimodal skills; and (3) task-unifying approaches that bridge understanding and generation. We demonstrate how these directions transform multimodal large language models from versatile foundations into adaptive, human-aligned systems, providing researchers with a structured roadmap for developing next-generation multimodal AI.
Cite
CITATION STYLE
Wang, Z., Li, L., & Chen, L. (2025). Recent advances in finetuning multimodal large language models. AI Magazine, 46(3). https://doi.org/10.1002/aaai.70025
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.