Video text retrieval is a hot research topic in artificial intelligence, with the core challenge being the semantic gap between visual dynamic features and discrete linguistic symbols. In recent years, with the development of large-scale models, cross-modal modeling capabilities have significantly improved, driving continuous evolution in retrieval methods regarding granularity modeling strategies. This article provides a systematic review of research methods in video text retrieval, categorizing them into single-granularity retrieval and multi-granularity retrieval. Single-granularity retrieval focuses on modeling a single semantic layer. Coarse-grained methods achieve efficient retrieval through global feature matching using pre-trained models, but they suffer from incomplete semantic coverage. Fine-grained methods enhance semantic analysis accuracy through local alignment mechanisms, but they are constrained by inherent limitations. In contrast, multi-granularity retrieval combines global scene understanding and local detail perception through hierarchical feature fusion strategies, with typical technical approaches including dynamic fusion frameworks. Analysis results indicate that multi-granularity retrieval can more comprehensively capture cross-modal semantic associations, providing a more effective solution for video text retrieval.
Wang et al. (Wed,) studied this question.