PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 17, 2026˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences1 citationsOpen Access

3D Building Model Segmentation using GNN and ViT

View Full Paper
HRH. RashidanIMI. A. MuslimanARAlias Abdul Rahman

Key Points

  • The study aims to enhance the accuracy of component segmentation in 3D building models using advanced neural networks.
  • Applied a Graph Neural Network (GNN) to analyze building mesh structure.
  • Rendered multi-view 2D projections (orthographic and perspective) for visual pattern extraction.
  • Utilized a Vision Transformer (ViT) to identify elements like windows and doors.
  • Implemented a consensus fusion method to merge GNN and ViT predictions.
  • The pipeline shows improved accuracy compared to a GNN baseline.
  • Clearer gains are observed in classifying small or visually ambiguous components.
  • Achieved better classwise consistency across all tested categories.

Abstract

Abstract. Reliable semantics in 3D building models support practical urban tasks such as planning, asset inventory, and maintenance. This paper presents an approach that pairs graph-based geometry (GNN) with image-based appearance (ViT) to improve component segmentation. A Graph Neural Network (GNN) is first applied to the building mesh to capture structural cues and produce initial labels. Multi-view 2D projections (orthographic and perspective) are then rendered and processed with a Vision Transformer (ViT) to recover visual patterns related to windows, doors, roofs, and walls. The two streams are reconciled through a simple consensus fusion that projects ViT predictions back onto the 3D geometry and refines the labels. In experiments, the proposed pipeline improves accuracy and classwise consistency over a GNN baseline, with clearer gains on small or visually ambiguous elements.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Rashidan et al. (2026) studied this question.

synapsesocial.com/papers/696b2631d2a12237a93498b2https://doi.org/10.5194/isprs-archives-xlviii-4-w17-2025-279-2026
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Semantic Segmentation of Building Models with Deep Learning in CityGML2024 · 2 citations
  2. 2OV3DSeg-VGGT: Open-Vocabulary 3D Segmentation with Visual Geometry-Grounded Transformers2026
  3. 3A lightweight hybrid ViT-GNN framework for data-centric land cover mapping in the amazon biome using graph structural priors2026
  4. 4GreedyViG: Dynamic Axial Graph Construction for Efficient Vision GNNs2024
  5. 5PointViG: A Lightweight GNN-based Model for Efficient Point Cloud Analysis2024