PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 20, 2016344 citationsOpen Access

MOSI: Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis in Online Opinion Videos

AZAmir ZadehRZRowan ZellersEPEli Pincus

Key Points

  • This research aims to develop a corpus for analyzing sentiment and subjectivity in online opinion videos.
  • Introduced the Multimodal Opinion-level Sentiment Intensity dataset (MOSI) with annotations for sentiment and subjectivity.
  • Annotated video features including per-frame visual and per-millisecond audio characteristics.
  • Presented baselines for future research and a new approach to multimodal fusion.
  • Successfully created the first opinion-level annotated corpus for sentiment analysis in videos.
  • Provided rigorous annotations for sentiment intensity and features, paving the way for future studies.
  • Established new baselines for sentiment and subjectivity research in multimedia.”

Abstract

People are sharing their opinions, stories and reviews through online video sharing websites every day. Studying sentiment and subjectivity in these opinion videos is experiencing a growing attention from academia and industry. While sentiment analysis has been successful for text, it is an understudied research question for videos and multimedia content. The biggest setbacks for studies in this direction are lack of a proper dataset, methodology, baselines and statistical analysis of how information from different modality sources relate to each other. This paper introduces to the scientific community the first opinion-level annotated corpus of sentiment and subjectivity analysis in online videos called Multimodal Opinion-level Sentiment Intensity dataset (MOSI). The dataset is rigorously annotated with labels for subjectivity, sentiment intensity, per-frame and per-opinion annotated visual features, and per-milliseconds annotated audio features. Furthermore, we present baselines for future studies in this direction as well as a new multimodal fusion approach that jointly models spoken words and visual gestures.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zadeh et al. (2016) studied this question.

synapsesocial.com/papers/6a08001a98f34196d2735d7chttps://doi.org/10.48550/arxiv.1606.06259
Ask AI
Helpful
Bookmark
Share
View Full Paper