KPoEM: A Human-Annotated Dataset for Emotion Classification and RAG-Based Poetry Generation in Korean Modern Poetry
Published in The Review of Korean Studies, 2026
Online link
Dataset
Download paper here
Abstract
This study introduces KPoEM (Korean Poetry Emotion Mapping), a novel dataset that serves as a foundation for both emotion-centered analysis and generative applications in modern Korean poetry. Despite advancements in NLP, poetry remains underexplored due to its complex figurative language and cultural specificity. We constructed a multi-label dataset of 7,622 entries (7,007 line-level and 615 work-level), annotated with 44 fine-grained emotion categories, drawn from the works of five influential Korean poets. The KPoEM emotion classification model, fine-tuned through a sequential strategy—moving from general-domain emotion corpus (KOTE) to the specialized KPoEM dataset—achieved a micro F1-score of 0.60, significantly outperforming the baseline model (0.43). The model demonstrates an enhanced ability to identify temporally and culturally specific emotional expressions while preserving core poetic sentiments. Furthermore, applying the structured emotion dataset to a Retrieval-Augmented Generation (RAG)-based poetry generation model demonstrates the feasibility of generating poetic texts that reflect the emotional and cultural sensibilities of Korean literature. This integrated approach strengthens the connection between computational techniques and literary analysis, opening new pathways for quantitative emotion research and generative poetics.

Recommended citation: Lim, I., Ji, H., & Kim, B. (2026). KPoEM: A Human-Annotated Dataset for Emotion Classification and RAG-Based Poetry Generation in Korean Modern Poetry. The Review of Korean Studies, 29(1), 161-206.
