Building a Korean Literature TEI/XML Database: From Data Acquisition to AI-Enhanced Reading

Date:

A hands-on workshop presenting a complete pipeline for building TEI/XML databases of Korean literature, demonstrated through a corpus of 30 works from the 1910s-1940s.

Author

Byungjun Kim, Haein Ji, Gayeon Kim, Chaeyeon Jeong, Hagyeong Lee, and Seonyeong Park

Abstract

This workshop walks participants through a complete pipeline for building TEI/XML databases of Korean literature, demonstrated on a corpus of 30 works from the 1910s-1940s. The session covers four stages. Participants first learn text harvesting from Korean Wikisource, including the practical problems of source selection, orthographic variation, and provenance in early modern Korean texts. They then work with LLM-based auto-tagging, using large language models to propose TEI markup and evaluating where machine suggestions require human correction. The third stage addresses collaborative encoding practices — schema design, encoding guidelines, and workflows that keep a distributed team consistent in its markup. Finally, participants build RAG-enhanced literary analysis on top of the encoded corpus, using retrieval-augmented generation to support question-driven reading of the collection. The workshop is aimed at researchers and students who want a reproducible route from raw text to a queryable, AI-enhanced literary database, and it treats encoding not as a preliminary chore but as an interpretive act that shapes what can later be asked of the corpus.

DOI: 10.5281/zenodo.21830474