Key points are not available for this paper at this time.
• We propose DOcument-level Relation Extraction optiMizing the long taIl (DOREMI), an iterative system tailored for long-tail relations to enhance the distantly supervised dataset through disagreement-driven annotations which, to our knowledge, is the first DocRE model featuring a human-in-the-loop strategy for denoising. • We demonstrate that measuring the disagreement between multiple models is a good proxy to identify Hard-To-Classify examples and yields substantial performance improvements with negligible human effort. • We release two Denoised Distantly Supervised Datasets (DDSs), one based on DocRED and one on Re-DocRED, which can be used to train any DocRE model. These DDSs greatly improve the prediction of long-tail relations and complement existing denoising approaches, as confirmed by our experimental evaluation. Document-Level Relation Extraction (DocRE) presents significant challenges due to its reliance on cross-sentence context and the long-tail distribution of relation types, where many relations have scarce training examples. In this work, we introduce DO cument-level R elation E xtraction opti M izing the long ta I l (DOREMI), an iterative framework that enhances underrepresented relations through minimal yet targeted manual annotations. Unlike previous approaches that rely on large-scale noisy data or heuristic denoising, DOREMI actively selects the most informative examples to improve training efficiency and robustness. DOREMI can be applied to any existing DocRE model and is effective at mitigating long-tail biases, offering a scalable solution to improve generalization on rare relations.
Menotti et al. (Fri,) studied this question.