Overview This poster presents MagicTagger, a prototype workflow for transforming archival Russian-language folktales into a queryable, provenance-aware, and reusable data (submission number - 213). The project addresses a recurring problem in computational folklore: archival texts and typological classifications are often distributed across catalogues, manuscript images, local spreadsheets, and even intangible expert knowledge. As Tangherlini notes, “many … research collections do not exist in machine-readable form,” while folklore records systematisation became what Ilyefalvi calls the “immensity of the systematization” (Tangherlini, n.d.; Ilyefalvi 2018). Case Study The case study focuses on Russian-language tales of magic preserved in the Estonian Folklore Archives. These collections preserve folktale texts recorded across different villages on the Estonian-Russian border, from more than 20 narrators, from 20’s till 60’s. These texts exemplify practical archival constraints: manuscript-based sources, heterogeneous metadata, incomplete machine-readable access, and rights-sensitive contextual information. Even when manuscript pages have been scaned, “hand-written pages are not machine-readable,” which means that text extraction, normalisation, and data-readiness must be treated as part of the research workflow (Järv and Sarv 2013). MagicTagger treats this case as a knowledge-management problem. Without stable identifiers, controlled vocabularies, and explicit links between records, it is difficult to retrieve variants, compare them systematically, or reuse classifications. This is also why computational access to folklore is a cultural heritage issue: as Bascom stressed that “classification is a vexing first-order problem in folklore” (Bascom, n.d.). In this project, a folktale type is an abstract narrative pattern used to group different variants of the same story. It is a comparative unit that allows tales recorded in different places, decades, and collection contexts to be studied together. MagicTagger uses the Aarne–Thompson–Uther (ATU) classification system, the most widely used international tale-type system in folklore studies (Uther 2011). Each ATU type has a stable number and title, such as “ATU 510A, Cinderella”, making it possible to retrieve and compare variants across archival records. This is methodologically important because ATU “allows structure where diversity would otherwise only be apparent” (D’Huy 2019). In MagicTagger, ATU concepts are represented as stable SKOS-style resources, allowing the user to move from a type code to concrete archival records, variant sets, and exportable structured data. Workflow and Interface MagicTagger implements a complete pipeline: archival records are curated into a corpus index; tale texts and metadata are normalised; records are represented as a knowledge graph; the graph is validated and queried; and the resulting data are made accessible through a web interface. The interface supports two main workflows: Explore, for corpus-level discovery and type-mediated comparison, and Classify, for classifier-assisted enrichment of new Russian tale texts. The project’s outputs include RDF/JSON-LD exports, SHACL validation reports, classifier results, and reusable data bundles, so the system is not only an interface but also a reproducible research workflow. Knowledge Graph The knowledge model follows a reuse-first strategy. The profile combines widely used vocabularies: DCTERMS for core description, SKOS for tale-type concepts, PROV-O for provenance, and a lightweight CIDOC-CRM for archival entities. This design follows the FAIR principle that reusable data should be machine-readable and organised so that machines can “automatically find and use the data”. The graph also responds to the FAIR requirement that data should have a “globally unique and lasting identifier” and be “described with rich metadata” (Wilkinson et al. 2016). Classifier-Assisted Enrichment The classifier component is presented as a controlled enrichment layer, not as an automatic authority. Given an external Russian tale text, the system returns the three most likely ATU candidates, together with scores, confidence signalling (high or low), and classification run metadata. These suggestions are designed for expert review: high-confidence cases can support faster metadata enrichment, while uncertain cases remain visible as requiring interpretation. This follows GLAM-oriented guidance that AI should support cataloguing workflows, not replace human cataloguers, and that enrichments require review before becoming authoritative metadata (Europeana Copyright Community Steering Group 2025). Scope and Limitations The current project version is a pilot. It focuses on Russian-language magic tales and does not claim to solve cross-linguistic folktale comparison or fully automatic classification. The classifier is constrained by the small size of the training data (117 texts) and by the quality of text recognition from manuscript sources (out of 510 checked pages 137 were excluded because of the bad HTR quality). For this reason, uncertainty, provenance, and expert validation are treated as core parts of the workflow. Contribution The project will highlight three aspects of the applications prototype. First, it will provide a quantitative overview of the corpus, including distribution by collection, ATU type, narrator, collector, place, and recording period. Second, it will demonstrate type-mediated retrieval by selecting an ATU type and moving from a typological category to concrete archival records. Third, it will demonstrate classifier-assisted enrichment and export of both corpus data and prediction results as reusable RDF/JSON-LD.
Evgeniia Vdovichenko (Sun,) studied this question.