In untrimmed video data of indoor scenes, actions exhibit complex temporal relationships, such as co-occurrence and compositional dependencies, and multi-granularity semantic hierarchies. Detecting actions under these intricate relationships is a challenging task. Many existing methods primarily rely on complex temporal annotations to address the problem of multi-granularity action detection. However, this approach often lacks explicit modeling of hierarchical dependencies between actions. This oversight results in limited capability to handle multi-granularity action detection across complex temporal scales. To address the limitations, this paper proposes a novel multi-granular action detection method based on highlighted a set of categories hierarchical prior knowledge graph, which is first introduced in this paper. For representing the hierarchical prior knowledge graph, a Quantified Action Hierarchical prior Construction(QAHC) approach is proposed, through which the semantics of action categories are integrated with the actions’ temporal patterns to express the hierarchical dependencies and logical relationships among actions both qualitatively and quantitatively in the form of a directed acyclic graph. Based on the hierarchical prior knowledge graph, we design a Prior Highlight guided Action Detection(PHAD) network. It employs a prior highlighted mechanism to select sufficient and necessary prior knowledge in real-time to guide the efficient detection of multi-granularity actions. Experimental results show that our method achieves excellent accuracy on the Charades (+1.49) and STS (+1.12) datasets.
Wang et al. (Sat,) studied this question.