Key points are not available for this paper at this time.
In this paper we study how external affective information should be integrated into a compact transformer-based text classifier. Rather than treating affective features as signals to be appended directly to the representation, we examine whether their contribution should be controlled through lightweight fusion mechanisms. The comparison focuses on scalar-gated fusion versus plain concatenation, using DistilBERT as the textual backbone and four affective resources: the NRC VAD Lexicon, VAD-BERT, Ekman-style emotion scores, and SenticNet. The evaluation is conducted on two English corpora with different label structures: a seven-class MentalHealth dataset and the fine-grained GoEmotions benchmark. Across both corpora, scalar gating consistently matches or outperforms concatenation in terms of Macro-F1. On MentalHealth, scalar gating improves all directly comparable configurations. On GoEmotions, it achieves the best overall Macro-F1 and improves most matched comparisons. Beyond static feature integration, we introduce affective flow (EmoFlow) representations derived from VAD-BERT, which model the evolution of valence, arousal, and dominance across segments of a text. These dynamic representations do not surpass the strongest static lexical resources in absolute performance, but they provide consistent improvements within the VAD-BERT family, particularly when combined with scalar gating or cross-attention. Our contribution is twofold. First, we show that a lightweight scalar gate provides an effective and interpretable mechanism for adaptively integrating low-dimensional affective side information into transformer-based classifiers. Second, we introduce affective flow representations that explicitly model how affect evolves within a document, enabling the analysis of both adaptive resource selection and intra-document affective dynamics. Together, these results suggest that the key issue is not only which affective resources to use, but also when and how they should influence the model.
Calvo et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: