3D reconstruction and classification of aircraft are two active research areas in optical remote sensing image processing which are of great significance for applications such as airport monitoring and intelligence analysis. The traditional approaches usually focus only on one of these two tasks, and all these methods suffer from inherent limitations. In the field of 3D reconstruction, most current methods require multiple-view images as input, which is rarely feasible in remote sensing. However, single-view 3D reconstruction is an inherently ill-posed problem. Existing methods, including voxel generation and mesh template deformation, still suffer from limited accuracy and poor shape fidelity. In the field of image classification, the existing methods are mainly based on deep learning. These methods require a large amount of labeled data, and they may also be misled by the color and texture features of the target in the dataset. In this paper, we propose a unified framework for simultaneous 3D reconstruction and classification, specifically tailored for aircraft targets in optical remote sensing imagery. The key innovations are threefold: First, we introduce the Signed Distance Field (SDF) implicit representation to build a prior-guided 3D reconstruction framework pre-trained on 3D model datasets. Second, to achieve the reconstruction process with a single image as input, we design a new joint optimization pipeline. We propose a novel dual-kernel differentiable rendering method, which is fused behind the SDF generation network for iterative optimization of the implicit code and pose parameters. Third, a gated feature fusion module is developed to combine the optimal latent vector from reconstruction with the classification backbone. This integration enables the joint output of 3D meshes and category labels within a unified loop. The resulting optimal latent code plays a dual role as a generative seed for high-fidelity 3D reconstruction and as a low-dimensional feature representation for target classification. Quantitative evaluations validate the superiority of our joint framework. Compared with the strong mesh-based competitor AtlasNet, the proposed method yields a 12.2% boost in mean F-score. In object classification, leveraging the 3D implicit geometric features boosts the performance to a peak accuracy of 97.88%, outperforming advanced remote sensing backbones such as RSMamba and EAM by 2.03% and 2.54%. Additionally, ablation studies confirm the indispensability of our key designs, revealing that our dual-task feature fusion strategy brings an absolute gain of 1.18% in classification accuracy, while omitting the clustering prior stages and the dual-kernel rendering method leads to a 30.4% and 10.1% degradation in Chamfer distance.
Wang et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: