Multimodal Fine-Tuning of LLMs for Robust Document Visual Question Answering

3Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Document Visual Question Answering (DocVQA) necessitates comprehension of both the spatial layout and the textual content. Multimodal pretraining is a foundational component of existing vision-language models, including LayoutLM. However, they frequently lack integration with potent Large Language Models (LLMs). This work addresses this gap by fine-tuning Flan-T5 on the SP-DocVQA dataset using both text and bounding box information across multiple context categories. This spatial-textual alignment allows the model to attain an ANLS score of 76% solely through the text modality. In order to integrate visual comprehension, we implement a multimodal pipeline that coordinates cropped word images with the LLM embedding space through a novel pretraining task. Additionally, we present two DocVQA strategies that incorporate visual word embeddings to improve document comprehension. Empirical findings indicate that models utilizing bounding box information substantially surpass those employing text-only or layout-aware inputs, especially for spatially-grounded inquiries. In a pre-task evaluation, PT2 outperforms PT1 with significant enhancements in ANLS (+20%) and Accuracy (+24%), however it exhibits a minor decline in GTIP (–6.1%).

Cite

CITATION STYLE

APA

Tripathi, S., Tabrez Nafis, M., Hussain, I., & Saudagar, A. K. J. (2025). Multimodal Fine-Tuning of LLMs for Robust Document Visual Question Answering. IEEE Access, 13, 174611–174623. https://doi.org/10.1109/ACCESS.2025.3615201

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free