Technical report demonstrates an auditable OAI-PMH harvesting workflow in digital repository systems, highlighting robust metadata quality assurance and provenance tracking.
This technical note presents the EPrints OAI-PMH Import Workflow, a production-informed and auditable methodology for harvesting, selecting, mapping, importing, enriching, and validating scholarly metadata and full-text documents in EPrints repositories. The workflow separates scientific scope, technical harvestability, and rights/licensing eligibility; distinguishes structural import fidelity from semantic and bibliographic quality assurance; and adopts a controlled write cycle based on snapshot → preflight → apply → verify. Particular attention is given to source identity, DOI semantics, collision detection, destination-field semantics, multilingual metadata, structured references, creator handling, controlled classification, field-level provenance, and anomaly taxonomy. Publication PDFs are treated as targeted authoritative evidence when structured source metadata are incomplete, rather than as blind fallback sources. The methodology is derived from production work carried out in 2026 on RM Open Archive and E-LIS, with case studies involving University of Milan journals, Serena/SHARE, En la España Medieval, AIB Studi, and a JLIS.it pilot. The JLIS.it case demonstrates a layered-source strategy combining OAI/OJS core metadata, structured citation_reference bibliography data, and authoritative publication-PDF evidence for multilingual enrichment. Release v0.1.0 is archived in Zenodo under DOI 10.5281/zenodo.22122826; the Concept DOI for all versions is 10.5281/zenodo.22122825. The workflow is publicly available on GitHub and distributed under the MIT License.
No takes yet. Share an insight, caveat, or question.
Roberto Delle Donne (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: