With the emergence of Large Language Models (LLMs), there has been growing interest in AI-assisted literary creation and its sociocultural implications. Recent studies report that AI-generated poetry can receive higher quality ratings than human-written works (Porter & Machery, 2024), suggesting a significant shift in our understanding of 'creativity' as a uniquely human domain. Additionally, research utilizing fine-tuned models like GPoeT-2, which generates and evaluates poetry in specific forms (Lo et al., 2022), demonstrates how AI is becoming an innovative platform for literary research and creation. This study aims to construct a specialized dataset for emotion classification in Korean modern poetry, contributing to computer-based literary research and AI-assisted poetic creation. The recent Nobel Prize in Literature awarded to Han Kang has heightened global interest in Korean literature, making it particularly timely to systematically analyze and preserve the unique emotional expressions and aesthetics of Korean literary works through digital technology. Our research focuses on the KOTE (Korean Online That-gul Emotions) dataset, which utilizes the KcELECTRA model—a BERT-based ELECTRA model fine-tuned on Korean comments (Jeon et al., 2022). While KOTE provides a valuable resource with 50,000 online comments classified into 44 emotions using Korean emotional lexicon-based clustering, its foundation in informal online language limits its applicability to analyzing the sophisticated emotional expressions found in modern poetry, including archaic language and poetic devices. To address these limitations, we collected 316 poems from five prominent Korean modern poets: Kim So-wol (85 poems), Yun Dong-ju (88 poems), Im Hwa (20 poems), Yi Sang (31 poems), and Han Yong-un (92 poems). We scraped their major works from public domain sources, including "The Azure Daisy" (진달래 꽃), "Sky, Wind, Stars, and Poetry" (하늘과 바람과 별과 시), "The Korea Strait" (현해탄), and "The Silence of Love" (님의 침묵). These texts were divided into line-level (4,000 entries) and poem-level (380 entries) datasets, totaling 4,380 annotated texts. For text preprocessing, we implemented automated Korean-Chinese character conversion with parallel notation, and developed a systematic approach to handle archaic expressions while preserving their original poetic nuances. The labeling process included multiple rounds of cross-validation among annotators to ensure consistency and reliability in emotional classification, particularly for cases involving complex literary devices such as metaphors and classical rhetorical expressions. Our study employed several digital technologies to optimize the dataset for machine learning: BERT-based tokenization and embedding processing Automatic text segmentation based on maximum length (512 tokens) One-hot encoding for multi-label classification Cross-validation for labeling reliability assessment Through this process, we constructed a total of 4,380 annotated texts, comprising 4,000 line-level and 380 poem-level entries. The emotional labeling was performed by three researchers specializing in Korean literature and digital humanities. The labeled dataset was then combined with KOTE to fine-tune a KcELECTRA model. For performance validation, we randomly split the entire dataset into training (80%), validation (10%), and test (10%) sets, with the test set completely isolated to ensure independent evaluation. Performance assessment was conducted under three experimental conditions (KOTE-only, KOTE+line-level, KOTE+line-level+poem-level), using quantitative metrics including accuracy, F1 scores, and ROC-AUC for statistical analysis. This approach allowed us to objectively verify the generalization capability of both the dataset and the model. The dataset and model developed in this study serve as foundational resources for sophisticated emotional analysis of Korean modern poetry, bridging traditional literary studies with digital methodologies. This approach enables the analysis of emotional patterns in large-scale texts through "distant reading" methodologies, offering systematic insights into the temporal and cultural emotions embedded in literary works. Furthermore, by integrating with generative AI technology and deepening the learning of poetic language emotions, it can evolve into an innovative literary creation tool that inspires human creative processes. The model's practical applications extend to various domains, including computer-assisted poetry translation, educational tools for teaching Korean literature, and computational analysis of emotional patterns in cross-cultural literary studies. This multifaceted approach not only serves academic research but also provides practical tools for literature education and translation work. We plan to make our dataset and model publicly available to facilitate further research in computational literary analysis. Future work could extend this methodology to other genres of Korean literature and explore cross-cultural emotional patterns in Asian literature more broadly, potentially revealing shared emotional expressions and cultural connections across different literary traditions. Moreover, this study provides an empirical foundation for globalizing Korean literature by systematically analyzing its emotional characteristics in the digital age. This foundation will enable more accurate conveyance of cultural contexts and emotional nuances in translation and international presentation of Korean literature. Particularly in the current era of active literary production and consumption through online platforms, our data-driven emotion analysis methodology offers crucial insights into understanding and developing literary communication in digital environments. Furthermore, the methodology developed in this study presents a model applicable to emotional literary research across different languages and cultures. By contributing to the analysis and understanding of unique emotional expressions in various cultures through digital technology, it is expected to open new horizons for cross-cultural comparative research in digital humanities. Ultimately, this study represents a significant attempt to explore new possibilities in literary research and creation through the convergence of humanistic insight and digital technology.
Lim et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: