Enhancing Technical Question Answering Quality Through Multimodal Document Segmentation

0Citations
Citations of this article
12Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Technical documents present unique multimodal understanding challenges due to their heterogeneous information elements requiring both localized visual interpretation and global contextual reasoning. We introduce a novel two-tiered augmentation strategy that uniquely bridges this gap by generating semantically enriched document fragments through layout-aware segmentation and multimodal annotation. Unlike conventional RAG approaches that process documents as atomic text units, our methodology preserves visual integrity of technical elements while creating a hierarchical evidence structure that balances global context (via Page-Augmented Retrieval) with precise local evidence (via Localized Fragment Augmentation). This approach addresses the fundamental limitation in document AI: the trade-off between preserving spatial relationships and providing task-relevant context. Our method achieves state-of-the-art results, improving GPT-4’s DesignQA performance by 22.4% and demonstrating strong generalizability on ScienceQA (85.3% MC accuracy) and MMMU (66.7% overall). By enabling accurate interpretation of design constraints and functional requirements directly from technical documentation, our framework advances AI systems for engineering design verification and regulatory compliance.

Cite

CITATION STYLE

APA

Lvov, D., Smirnov, I., Volokha, V., Laushkina, A., & Boukhanovsky, A. (2026). Enhancing Technical Question Answering Quality Through Multimodal Document Segmentation. IEEE Access, 14, 12733–12743. https://doi.org/10.1109/ACCESS.2026.3655813

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free