Key points are not available for this paper at this time.
With the continuous advancement of the livestreaming industry, streamers as content producers pose significant challenges to the timeliness of regulatory response mechanisms, emerging as a critical weak link in cyberspace governance. Human-object interaction (HOI) detection plays a pivotal role in understanding multimodal livestreaming videos. In mainstream healthy online ecosystems, normal HOI categories dominate, while rare ones are extremely scarce, i.e., long-tail distribution that hinders HOI models from effectively detecting streamer violations. Driven by the transformative potential of foundation models (FMs) in multimodal video understanding, we propose a knowledge-driven memory network (KdM-Net) for long-tailed HOI in livestreaming, leveraging the extensive capabilities of the contrastive language-image pretraining (CLIP) model. First, human-object (HO) pairs are generated and modeled using general object detector/tracker. After converting each HOI label into a short sentence description, text embeddings are extracted via the CLIP text encoder to initialize classifier weights. Notably, we introduce a visual-textual knowledge transfer strategy to align visual and text features, complementing for the multimodal knowledge deficit of rare categories that plague long-tailed HOI distributions. Finally, a knowledge-driven memory module is designed to dynamically assign adaptive weights and attention to interaction categories based on their long-tailed distribution characteristics, mitigating model forgetting tail data and enhancing HOI detection performance. Experimental results demonstrate that KdM-Net achieves HOI detection accuracies of 37.33
Zhang et al. (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: