This article introduces a novel approach for fast summarization of user-generated videos (UGVs). Different from other types of videos where the semantic content might vary greatly over time, most UGVs contain only a single shot with relatively consistent high-level semantics and emotional content. Therefore, a few representative segments, which can be selected based on segment-level semantic and emotional recognition results, are generally sufficient for a summary. In addition, due to the poor shooting quality of many UGVs, factors such as camera shaking and lighting conditions are also considered to achieve more pleasant summaries. This article is part of a special issue on quality modeling.
No takes yet. Share an insight, caveat, or question.
Xu et al. (2016) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: