GOLEMcoref: A Multilingual Coreference Dataset of Fiction
Published in Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026, Volume 2: Short Papers), 2026
Online link
Github
Download paper here
Abstract
We present a multilingual coreference dataset of 827k tokens of fiction in 7 languages: Bahasa Indonesia, Chinese, Dutch, English, Italian, Korean, and Spanish. The dataset includes full stories of diverse lengths, ranging from 500 to 17k words. We discuss our annotation scheme focusing on characters and language-specific challenges we encountered. Finally we present evaluation results of a neural coreference system trained on our dataset. We show that jointly training a system across all languages provides a strong improvement over monolingually trained models. The dataset is available under a creative commons license in CoNLL-2012 and CorefUD format at https://github.com/GOLEM-lab/GOLEMcoref/.
Recommended citation: van Cranenburgh, A., Yang, X., Alvanita, Di Domenico, C. N., Ferragud, M., Graciotti, A., Ion, A. G., Kim, B., Park, S., Visser Solissa, N., Zhou, X., & Pianzola, F. (2026). GOLEMcoref: A Multilingual Coreference Dataset of Fiction. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 472-480.
