| ||||
| ||||
![]() Title:Semantic Video Understanding for Automated Facility Maintenance Inspection Conference:2026 SCEM Tags:1. Facility Maintenance Inspection, 2. Video-based Inspection, 3. Vision-Language Models, 4. Building Information Modeling (BIM) and 5. Graph-based Comparison Abstract: Facility maintenance inspection is critical for ensuring the safety, functionality, and operational reliability of building assets. However, traditional inspection processes are time-consuming, labor-intensive, and heavily dependent on manual documentation. Although automated approaches have been proposed, they typically involve complex processing pipelines, high computational costs, and reliance on large, manually annotated domain-specific datasets. To address these limitations, this paper presents an automated, video-based facility maintenance inspection framework that integrates Vision-Language Models (VLMs) with Building Information Modeling (BIM) to systematically compare as-designed and as-observed building conditions. The proposed framework operates in three stages. First, IFC-based BIM data are transformed into a hierarchical BIM Graph that encodes the spatial and hierarchical relationships among building stories, spaces, and facility assets. Second, a pretrained VLM (Qwen3-VL-8B-Instruct) analyzes egocentric walkthrough videos to identify visited spaces and detect facility assets, which are then structured into a Video Graph representing the as-observed condition. Third, Graph Edit Distance (GED) is applied to quantify structural discrepancies between the two graphs. A case study conducted in the Civil Engineering Research Building at National Taiwan University yielded a GED of 210 against a maximum possible edit cost of 550, corresponding to a normalized similarity score of 61.9%. Semantic Video Understanding for Automated Facility Maintenance Inspection ![]() Semantic Video Understanding for Automated Facility Maintenance Inspection | ||||
| Copyright © 2002 – 2026 EasyChair |
