PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 21, 2026Discover Mental Health1 citationsOpen Access

A unified multimodal learning framework for sentiment analysis and mental health indicators from YouTube videos

PSPriyanshu SatapathyOCOnushka ChauhanDKDeepika Kumar

Key Points

  • This study aims to analyze sentiments in YouTube videos and explore their relationship with mental health indicators.
  • Developed a multimodal deep learning framework integrating text, audio, and visual data.
  • Trained and evaluated the model on a diverse dataset of YouTube content.
  • Used transformer-based fusion for stable performance across different communication styles.
  • The multimodal model showed higher accuracy and fewer misclassifications than unimodal approaches.
  • Strong correlations were found between sentiment trajectories and emotional instability indicators.
  • Contradictions between facial expressions and verbal sentiments indicated potential mild distress signals.

Abstract

This study presents a multimodal deep learning framework designed to analyze sentiment patterns in YouTube videos and explore their association with early indicators of mental well-being. The approach integrates textual transcripts, vocal characteristics, and facial expressions into a unified representation to capture the emotional depth that individual modalities often miss. The model was trained and evaluated on a curated dataset of diverse YouTube content, and the results consistently showed that the fused architecture performed better than unimodal baselines. Compared with text-only or audio-only systems, the multimodal model achieved higher accuracy and fewer misclassifications, particularly in cases where speakers displayed subtle or mixed emotions. The integration of vocal cues such as pitch variation, speaking rate, and stress patterns helped clarify emotional ambiguity, while visual features such as micro-expressions, gaze direction, and facial tension added further clarity to the sentiment shifts within the videos. Transformer-based fusion delivered the most stable performance, demonstrating strong generalization across varied communication styles and recording conditions. In addition to reporting classification outcomes, the study examined how specific multimodal patterns correlate with non-clinical markers of mental health. Consistent associations were observed between fluctuating sentiment trajectories and indicators such as emotional instability, sustained negative tone, and reduced expressive variability. Instances where facial expressions contradicted verbal sentiment also showed relevance for identifying mild distress signals. These findings suggest that multimodal emotional cues can offer valuable insights into the affective state of content creators and may support research on digital well-being. The analysis also revealed challenges related to background noise, varying video quality, and inconsistent facial visibility, which influenced the reliability of certain features. Despite these limitations, the study demonstrates that combining audio, visual, and textual information provides a more complete and reliable picture of sentiment expression on social media platforms. The proposed framework offers a foundation for future systems aimed at understanding online emotional behavior and contributes to ongoing discussions on the responsible use of machine learning in mental-health-oriented applications.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Satapathy et al. (2026) studied this question.

synapsesocial.com/papers/69994b88873532290d01fab9https://doi.org/10.1007/s44192-026-00388-6
Ask AI
Helpful
Bookmark
Share
View Full Paper