PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 6, 2026KNOWLEDGE ORGANIZATION0 citationsOpen Access

Automatic Classification of Records and Archives as Data: A Survey of Experiments Using Machine Learning

View Full Paper
EWEduardo WatanabeRSRenato Tarciso Barbosa de Sousa

Key Points

  • The study explores the application of AI and ML for classifying records and archives while identifying key principles for their integration.
  • Synthesis of archival science, classification theory, and information science to propose guidelines for AI application.
  • Comparative analysis of 24 machine learning experiments evaluating adherence to archival guidelines.
  • 75% of experiments acknowledge the principle of provenance.
  • 25% show acknowledgment of diverse perspectives in their classification.
  • Only 16.7% adhere to explainability guidelines.
  • 0% demonstrate algorithmic accountability, showing a significant gap in ethical practices.

Abstract

This paper investigates the application of machine learning (ML) to the automatic classification of records and archives, framing it as a critical challenge in Knowledge Organization (KO). As digitization creates massive volumes of uncategorized data, the following research question arises: how can fundamental archival principles—such as provenance, original order, and hierarchical description—be translated into this new computational paradigm? This study first synthesizes, based on a multidisciplinary review of archival science, classification theory, KO, computer science, and information science, a proposal of six fundamental guidelines for the responsible application of artificial intelligence (AI) in records and archives. These guidelines connect traditional archival theory with the modern imperatives of trustworthy and explainable AI. Second, we conduct a comparative analysis of 24 published ML experiments, assessing their adherence to these guidelines. Our analysis reveals a significant and troubling disconnect. While most experiments acknowledge the principle of provenance (75.0%), they demonstrate profound neglect of guidelines related to diverse perspectives (25.0%), explainability (16.7%), and, most critically, algorithmic accountability (0.0%). The results indicate that current practices often succeed in basic content categorization but fail in the more sophisticated archival task of preserving archives’ evidentiary and relational integrity by treating records as decontextualized data. The study calls for urgently developing an Archival AI Lifecycle—a framework that weaves archival principles, classification theory, and knowledge organization into AI development, safeguarding archival practice’s intellectual and ethical integrity in the digital age.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Watanabe et al. (2026) studied this question.

synapsesocial.com/papers/69faa25e04f884e66b533049https://doi.org/10.31083/ko53184
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Utilising machine-learning tools to increase access to archival collections2025
  2. 2Automated classification and storage of archival resources based on machine learning algorithms2026
  3. 3The Prospects and Challenges of Artificial Intelligence Technology in Archival Management2024 · 3 citations
  4. 4From hype to practice: Leveraging artificial intelligence to enhance archival discovery and accessibility2026
  5. 5Exploring the Use of Artificial Intelligence in Archival Activities2026