Key points are not available for this paper at this time.
What emotion do we feel when we see a situation? Multimodal sentiment analysis has been used to answer this question, but most of the research considers only low-level perceptual information such as textual, acoustic, and visual features. However, these features are not appropriate for the classification of situations as it is difficult to depict real-life complexities with low-level features. In this paper, we propose an emotion prediction framework which identifies polarity of emotion in situations using high-level contextual information, namely, location, people and time. Before predicting emotions, the framework structures data into `situation' segments and labels each segment based on our carefully designed annotation guideline. Our approach is tested with various situations in TV sitcoms as a substitute for real-life situations. Experimental results indicate that contextual information is more effective than textual or acoustic features in determining emotions induced by situations.
Jin et al. (Sun,) studied this question.