Comparative study evaluates genre prediction from reviews using feature-based and transformer models, highlighting performance differences.
This study investigates multi-label movie genre prediction from user-written reviews in which textual content is inherently subjective and the movies reviewed naturally belong to multiple genres. To address extreme class imbalance and label sparsity in the IMDb Large Movie Review Dataset, 234 fine-grained genre labels are consolidated into 35 parent categories using a deterministic genre-mapping strategy. A unified experimental pipeline evaluates traditional feature-based models (TF-IDF vectorization with Logistic Regression and Linear SVM), a sequence-based BiLSTM with self-attention using GloVe embeddings, and transformer-based architectures (DistilBERT and RoBERTa) under consistent evaluation metrics. Experimental analyses indicate that transformer-based architectures outperform alternative approaches, with RoBERTa achieving the best performance (Macro-F1 = 0.518, Micro-F1 = 0.576). The results indicate that genre consolidation enhances robustness under long-tailed label distributions. Moreover, contextualized transformer representations better capture implicit and subjective cues. The results further clarify practical trade-offs between predictive performance and computational efficiency across model families.
No takes yet. Share an insight, caveat, or question.
Davityan et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: