The increasing flow of documents in digital libraries, archives and electronic document management systems makes the standardization, adaptation and automation of the process of creating metadata an urgent scientific problem. Metadata directly affects the efficiency of document search, identification, semantic interpretation, long-term storage and intersystem exchange. However, while standardized description based on MARC21, a flexible approach to creating a dynamic field, and intelligent methods based on deep learning, cover these requirements separately, the issue of their full integration into a single methodological system has not been sufficiently resolved. In this study, an integrated hybrid model for describing electronic documents based on standardized, flexible, and intelligent metadata was proposed. A mixed electronic document corpus of 1500 documents was formed for evaluation. The corpus consisted of books, dissertations, scientific articles, archival documents, and heterogeneous electronic documents, with 300 samples selected from each group. Key metadata elements for each document were manually identified and used as ground truth. According to experimental results, the MARC21-based constructor achieved 96.8% structural compatibility and 95.6% metadata completeness, but the average description time was 6.8 min. The dynamic field approach achieved 93.4% structural compatibility and 94.1% metadata completeness, and reduced the description time to 4.1 min. The deep learning-based intelligent module achieved a structural matching score of 91.7%, a metadata extraction score of 93.8% F1, and reduced the processing time to 1.9 min. The proposed hybrid model achieved a structural matching score of 95.9%, a metadata F1 score of 95.1%, and an average description time of 2.3 min. The results showed that the hybrid model is a balanced solution between metadata quality, flexibility, and automation.
Dauletov et al. (Fri,) studied this question.