Purpose This study aims to demonstrate the application of text extraction methods to digitized travel records for synthesizing constituents for the construction of a knowledge base such that the complicated interrelations between various entities mentioned within multiple tourism scenes discussed in travel texts can be aptly represented. It is envisaged that the generated knowledge corpus can be used as supporting resource rich in historical content for different knowledge-based applications for deepening their understanding. Design/methodology/approach Text extraction was carried out on a selected digitized travel record to extract Resource Description Framework triples, which were mapped with an OWL ontology-based vocabulary constructed for this purpose using the Protege ontology editor. Relevant Python libraries were used. Findings The refined triples along with the ontology were loaded into Ontotext’s GraphDB for the purpose of generating a knowledge corpus. The knowledge corpus retained both (used during the time period of the travel text and present usage) types of spellings used for the entities and thus will be useful for the knowledge organization process under consideration as well as remain open for deployment in modern-day information retrieval for disambiguation purposes. Originality/value Studies on generation of knowledge graphs from travel-related documents exist but are few in number, especially those that cover South-Asian collections. This study intends to serve as a road map for South-Asian knowledge institutions by adding up to their efforts to improve access to their collections using innovative methods.
Ghosh et al. (Mon,) studied this question.