Key points are not available for this paper at this time.
The prevalence of government-funded dataset usage has yet to be comprehensively tracked and understood. The lack of a standardized citation methodology has thus far prevented the government from understanding dataset usage in a transparent, accessible way. In this work, we seek to build on recent successes in natural language processing techniques and a recent Kaggle competition to develop an extensible framework for extracting government dataset usage from scientific publications. Further, we apply the developed techniques to over 50,000 scientific articles from Elsevier's ScienceDirect collection. Finally, we show that improvements to the submitted algorithms along with ensembling improved overall performance on an evaluation dataset.
Hausen et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: