Abstract Stylized, non-photorealistic historical images pose a significant challenge to standard computer vision techniques, limiting the potential for large-scale quantitative analysis in art history and the digital humanities. This article addresses this “style gap’ by developing and validating a zero-shot deep learning workflow for object segmentation. Using the Qing Dynasty album Complete Illustrations of Liaozhai (Liaozhai Quan Tu) as a case study, we apply a vision-language model (CLIPSeg) to systematically identify, quantify, and analyze the furniture depicted in 204 illustrations. The quantitative results reveal a highly codified “arrangement grammar” underpinning the album’s visual narrative. We interpret this grammar through a “Body, Skeleton, Soul” framework: the functional Body is established by the numerical dominance of tables and chairs; the spatial Skeleton is constructed by the strong co-occurrence of screens and cabinets; and the symbolic Soul is manifested in the immense visual weight of the screen, which signifies authority and serves as a psychological threshold. This study makes a dual contribution: methodologically, it presents a transferable workflow for the computational analysis of stylized visual corpora; empirically, it decodes a systemic visual syntax in Qing Dynasty popular art, advancing a “distant reading” paradigm that moves from analyzing individual objects to understanding their relational grammar.
Li et al. (Wed,) studied this question.