Abstract
Effective machine learning for natural hazard prediction and monitoring depends on timely access to high-quality, event-specific datasets and models capable of adapting to evolving environmental dynamics (e.g., those induced by climate change). Equally important is model explainability, which enhances trust by clarifying decision-making processes and enabling insight into observed hazard patterns. This article introduces a novel approach for the automated construction of multimodal hazard datasets tailored for supervised learning. Central to our method is an ontology-driven self-labeling pipeline that semantically annotates each data element using concepts from a modular, integrated ontology encompassing geographic, hazard, sensor, spatial, and temporal dimensions. This enriched semantic representation facilitates the rapid generation of event-specific datasets and supports reuse across hazard types. Furthermore, embedding ontological descriptors into machine learning outputs enables explainable AI through semantic reasoning, enhancing the interpretability and transparency of predictions. Our pipeline allows for dynamic dataset creation, model adaptation to newly emerging patterns, and live semantic querying over a knowledge graph. Each dataset instance encapsulates a rich semantic narrative including hazard type, evolution stage, and contextual variables such as land cover, socio-environmental indicators, and historical weather data.
Author supplied keywords
Cite
CITATION STYLE
Grujdin, I., & Datcu, M. (2025). Ontology-Driven Pipeline for the Automated Generation of Multimodal Datasets for Supervised Learning in Natural Hazard Models. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 18, 27113–27127. https://doi.org/10.1109/JSTARS.2025.3622513
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.