Abstract
Bridge health diagnosis plays a vital role in ensuring structural safety and extending service life while reducing maintenance costs. Traditional structural health monitoring approaches rely on sensor-based measurements, which are costly, labor-intensive, and limited in coverage. To address these challenges, we propose a three-phase solution that integrates the Dynamic Lightweight Vision-Language Model (DL-VLM), domain adaptation, and knowledge-enhanced reasoning. First, as the core of the framework, the DL-VLM consists of three components: a visual information encoder with multi-scale feature selection, a text encoder for processing inspection-related language, and a multimodal alignment module. Second, to enhance practical applicability, we further introduce domain-specific fine-tuning on the Bridge-SHM dataset, enabling the model to acquire specialized knowledge of bridge construction, defects, and structural components. Third, a knowledge retrieval augmentation module is incorporated, leveraging external knowledge graphs and vector-based retrieval to provide contextually relevant information and improve diagnostic reasoning. Experiments on high-resolution bridge inspection datasets demonstrate that DL-VLM achieves competitive diagnostic accuracy while substantially reducing computational cost. The combination of domain-specific fine-tuning and knowledge augmentation significantly improves performance on specialized tasks, supporting efficient and practical deployment in real-world structural health monitoring scenarios.
Author supplied keywords
Cite
CITATION STYLE
Liang, S., He, Z., Gui, H., & Liu, F. (2026). DL-VLM: A Dynamic Lightweight Vision-Language Model for Bridge Health Diagnosis. Big Data and Cognitive Computing, 10(1). https://doi.org/10.3390/bdcc10010003
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.