Deep semantic interpretation of high-resolution remote sensing (RS) imagery is critical for refined Earth system monitoring; however, current data-driven approaches are impeded by the “semantic gap” inherent in existing datasets, which typically feature flat taxonomies, coarse categories, and insufficient fine-grained semantic descriptions. To mitigate these limitations, this paper presents LuoJia-FG, a large-scale, attribute-driven, and hierarchical fine-grained land-cover dataset tailored for next-generation vision–language models. The dataset comprises 119,619 multimodal triplets based on Gaofen-2, Ziyuan-3, and aerial imagery, with spatial resolutions ranging from 0.5 to 2.0 m. Furthermore, LuoJia-FG is distinguished by three core innovations. First, it establishes a deep three-level hierarchical taxonomy expanding from 8 primary classes to 52 secondary and 95 tertiary fine-grained classes, offering unprecedented semantic granularity. Second, it bridges the pixel–knowledge gap through a novel attribute-coding system that programmatically generates rich, attribute-driven text descriptions based on physical properties such as phenology and canopy density. Third, to demonstrate the utility of these multimodal annotations, we propose the CLIP-guided Hierarchical Classification Network (CLIP-HCNet) as a robust benchmark. This framework effectively leverages the dataset’s attribute-driven text descriptions as semantic priors to resolve visual ambiguity among spectrally similar fine-grained categories. Experiments verify that LuoJia-FG constitutes a challenging testbed and that incorporating attribute-driven semantic priors significantly enhances hierarchical classification accuracy, opening new avenues for text-guided geospatial understanding. The LuoJia-FG benchmark dataset is publicly available at https://doi.org/10.57760/sciencedb.33946, and the source code is available at https://github.com/WHUXyb/LuoJia-FG.
Xiong et al. (Sun,) studied this question.