PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 17, 2025Electronics3 citationsOpen Access

Facial and Speech-Based Emotion Recognition Using Sequential Pattern Mining

View Full Paper
YSYounghun SongKCKyungyong Chung

Key Points

  • The proposed method captures emotional transitions in dialogue effectively, overcoming limitations of static emotion classification.
  • Utilizing the MELD dataset, emotion sequences are generated based on utterance order, demonstrating improved accuracy.
  • Facial detection is achieved through DeepFace, while speech is transcribed by Whisper and classified with a BERT model.
  • The method employs a weighted voting scheme and LSTM-based classification for predicting overall emotional flow.

Abstract

We propose a multimodal emotion recognition framework that integrates facial expressions and speech transcription (where text is derived from the transcribed speech), with a particular focus on effectively modeling the continuous changes and transitions of emotional states during conversation. Existing studies have primarily relied on single modalities (text or facial expressions). They often perform static emotion classification at specific time points. This approach limits their ability to capture abrupt emotional shifts or the structural patterns of emotional flow within dialogues. To address these limitations, this paper utilizes the MELD dataset to construct emotion sequences based on the order of utterances and introduces an analytical approach using Sequential Pattern Mining (SPM). Facial expressions are detected using DeepFace, while speech is transcribed with Whisper and passed through a BERT-based emotion classifier to infer emotions. The proposed method fuses multimodal results through a weighted voting scheme to generate emotion label sequences for each utterance. These sequences are then used to construct an emotion transition matrix, apply change-point detection, perform SPM, and train an LSTM-based classification model to predict the overall emotional flow of the dialogue. This approach goes beyond single-point judgments by capturing the contextual flow and dynamics of emotions and demonstrates superior performance compared to existing methods through experimental validation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Song et al. (2025) studied this question.

synapsesocial.com/papers/68f199b7de32064e504dc70chttps://doi.org/10.3390/electronics14204015
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Multimodal Machine Learning: A Survey and Taxonomy2018 · 4,688 citations
  2. 2Transformers: State-of-the-Art Natural Language Processing2020 · 8,308 citations
  3. 3MultiMAE-DER: Multimodal Masked Autoencoder for Dynamic Emotion Recognition2024 · 13 citations
  4. 4Masked Graph Learning With Recurrent Alignment for Multimodal Emotion Recognition in Conversation2024 · 55 citations
  5. 5Automated emotion recognition: Current trends and future perspectives2022 · 174 citations