PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 22, 2018IEEE Transactions on Circuits and Systems for Video Technology128 citations

Two-Stream Collaborative Learning With Spatial-Temporal Attention for Video Classification

View Full Paper
YPYuxin PengYZYunzhen ZhaoJZJunchao Zhang

Key Points

Key points are not available for this paper at this time.

Abstract

Video classification is highly important and has widespread applications, such as video search and intelligent surveillance. Video naturally contains both static and motion information, which can be represented by frames and optical flow, respectively. Recently, researchers have generally adopted deep networks to capture the static and motion information separately, which has two main limitations. First, the coexistence relationship between spatial and temporal attention is ignored, although they should be jointly modeled as the spatial and temporal evolutions of video to learn discriminative video features. Second, the strong complementarity between static and motion information is ignored, although they should be collaboratively learned to enhance each other. To address the above two limitations, this paper proposes the two-stream collaborative learning with spatial-temporal attention (TCLSTA) approach, which consists of two models. First, for the spatial-temporal attention model, the spatial-level attention emphasizes the salient regions in a frame, and the temporal-level attention exploits the discriminative frames in a video. They are mutually enhanced to jointly learn the discriminative static and motion features for better classification performance. Second, for the static-motion collaborative model, it not only achieves mutual guidance between static and motion information to enhance the feature learning but also adaptively learns the fusion weights of static and motion streams, thus exploiting the strong complementarity between static and motion information to improve video classification. Experiments on four widely used data sets show that our TCLSTA approach achieves the best performance compared with more than 10 state-of-the-art methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Peng et al. (2018) studied this question.

synapsesocial.com/papers/6a1f37ccd09bc027e48330a0https://doi.org/10.1109/tcsvt.2018.2808685
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Temporal Segment Networks: Towards Good Practices for Deep Action Recognition2016 · 3,991 citations
  2. 2A Duality Based Approach for Realtime TV-L 1 Optical Flow2007 · 1,702 citations
  3. 3Discovering Primary Objects in Videos by Saliency Fusion and Iterative Appearance Estimation2015 · 37 citations
  4. 4Histograms of Oriented Gradients for Human Detection2005 · 32,130 citations
  5. 5lp-Norm Multiple Kernel Learning2011 · 429 citations