Key points are not available for this paper at this time.
We present PubMed 200k RCT, a new dataset based on PubMed for sequential classification. The dataset consists of approximately 200, 000 of randomized controlled trials, totaling 2. 3 million sentences. Each of each abstract is labeled with their role in the abstract using one the following classes: background, objective, method, result, or conclusion. purpose of releasing this dataset is twofold. First, the majority of for sequential short-text classification (i. e. , classification of texts that appear in sequences) are small: we hope that releasing a new dataset will help develop more accurate algorithms for this task. Second, an application perspective, researchers need better tools to efficiently through the literature. Automatically classifying each sentence in an would help researchers read abstracts more efficiently, especially in where abstracts may be long, such as the medical field.
Dernoncourt et al. (Mon,) studied this question.