Abstract This article explores the process of determining whole–part relationships in cataloging digital libraries. In contemporary library practices, accurately determining how an entity or work relates to others in the collection is considered an essential task for effective information organization and retrieval. The “Biblioteca Nacional de España” (Spanish National Library) has made significant efforts to achieve this objective, but the results are far from optimal due to the extensive human efforts required. We present a methodology capable of automatically identifying whole–part relationships and propose an online navigation tool for analyzing the structure of authors’ works by users and researchers. To facilitate easy adoption and widespread use, our methodology directly utilizes text extracted by Optical Character Recognition (OCR) technologies, eliminating the need for further preprocessing. Additionally, we introduce and provide access to Spanish19BNE, a dataset containing approximately 12,000 books in Spanish written by 559 different authors, with their original manifestations held in the Biblioteca Nacional de España collection. This dataset can be used for digital library research and computational analysis. Experimental results on this dataset demonstrate the viability of the proposed approach.
Campos et al. (2026) studied this question.