Multi-view infrastructure damage assessment method based on multimodal image analysis
Main Article Content
Abstract
Massive destruction of civil infrastructure caused by the war in Ukraine has created an urgent need for automated, reliable, and scalable methods for assessing structural damage to buildings. This study presents a comparative evaluation of state-of-the-art Vision-Language Models (VLM) and a hybrid computer vision approach for classifying infrastructure damage using multi-view data. A specialized dataset was curated based on ground-level imagery of war-damaged buildings collected across various regions of Ukraine, where each infrastructure object is represented by two to eight photographs, including pre- and post-damage images when available. The degree of damage is classified into five categories ranging from undamaged to completely destroyed. The proposed methodology combines a multi-view analysis approach with semantic segmentation based on a U-Net architecture with a ResNet encoder, alongside a rule-based heuristic model that estimates the degree of structural damage using pixel-level metrics. Information from different viewpoints is aggregated to form a unified object-level assessment, which reduces ambiguity caused by occlusions and limited viewing angles. The proposed approach is experimentally compared with several modern multimodal VLM, including ChatGPT version 5.5 (High Intelligence), Claude Sonnet version 5 (Medium), Gemini version 3.1 Pro (High), and Grok version 4 (Fast), using metrics such as Accuracy, Precision, Recall, F1-score, as well as overestimation and underestimation coefficients. The experimental evaluation, conducted on a specialized dataset containing images of damaged infrastructure objects, that modern VLMs outperform the baseline segmentation model in overall classification accuracy while providing robust reasoning in complex structural damage scenarios. Meanwhile, the proposed segmentation-based hybrid approach provides interpretable pixel-level damage assessment and remains suitable for deployment in resource-constrained environments. The developed multi-view assessment methodology serves as a key analytical component of an intelligent multi-functional GIS designed to support the inventory of construction and demolition waste, optimize its recycling logistics, and promote sustainable post-war recovery based on circular economy principles.

