Abstract
Technical documents present unique multimodal understanding challenges due to their heterogeneous information elements requiring both localized visual interpretation and global contextual reasoning. We introduce a novel two-tiered augmentation strategy that uniquely bridges this gap by generating semantically enriched document fragments through layout-aware segmentation and multimodal annotation. Unlike conventional RAG approaches that process documents as atomic text units, our methodology preserves visual integrity of technical elements while creating a hierarchical evidence structure that balances global context (via Page-Augmented Retrieval) with precise local evidence (via Localized Fragment Augmentation). This approach addresses the fundamental limitation in document AI: the trade-off between preserving spatial relationships and providing task-relevant context. Our method achieves state-of-the-art results, improving GPT-4’s DesignQA performance by 22.4% and demonstrating strong generalizability on ScienceQA (85.3% MC accuracy) and MMMU (66.7% overall). By enabling accurate interpretation of design constraints and functional requirements directly from technical documentation, our framework advances AI systems for engineering design verification and regulatory compliance.
Author supplied keywords
Cite
CITATION STYLE
Lvov, D., Smirnov, I., Volokha, V., Laushkina, A., & Boukhanovsky, A. (2026). Enhancing Technical Question Answering Quality Through Multimodal Document Segmentation. IEEE Access, 14, 12733–12743. https://doi.org/10.1109/ACCESS.2026.3655813
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.