Decoding the Poetic Language of Emotion in Korean Modern Poetry: Insights from a Human-Labeled Dataset and AI Modeling

Date:

This study introduces a human-labeled dataset of Korean modern poetry annotated with fine-grained emotion categories, and examines how well AI models can learn the emotional language of poetry.

Author

Iro Lim, Haein Ji, and Byungjun Kim

Abstract

This presentation reports on the construction of a specialized emotion classification dataset for Korean modern poetry, built to address the limitations of existing Korean emotion resources such as KOTE, which were developed on colloquial and non-literary text. The team compiled 316 poems by five major poets and produced 4,380 annotated texts, labeled at both the line and the poem level. Preparing early modern Korean verse for computational analysis required substantial preprocessing: automated Korean-Chinese character conversion with parallel notation, and systematic handling of archaic expressions and orthographic variation. Emotion classification models were then built using BERT-based tokenization and a fine-tuned KcELECTRA model trained on the combined datasets. The work bridges traditional literary studies and digital methodologies, enabling distant reading of poetic affect while supporting AI-assisted literary applications such as translation assistance and educational tools. The authors plan a public release of the dataset in order to advance computational literary analysis across languages and cultures.

DOI: 10.5281/zenodo.18752715