PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 9, 2020508 citationsOpen Access

Data and its (dis)contents: A survey of dataset development and use in machine learning research

APAmandalynne PaulladaUniversity of WashingtonIRInioluwa Deborah RajiUniversity of California, BerkeleyEBEmily M. BenderUniversity of Washington

Key Points

Key points are not available for this paper at this time.

Abstract

Datasets have played a foundational role in the advancement of machine learning research. They form the basis for the models we design and deploy, as well as our primary medium for benchmarking and evaluation. Furthermore, the ways in which we collect, construct and share these datasets inform the kinds of problems the field pursues and the methods explored in algorithm development. However, recent work from a breadth of perspectives has revealed the limitations of predominant practices in dataset collection and use. In this paper, we survey the many concerns raised about the way we collect and use data in machine learning and advocate that a more cautious and thorough understanding of data is necessary to address several of the practical and ethical issues of the field.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Paullada et al. (2020) studied this question.

synapsesocial.com/papers/6a097a4316dfdfe7ed342076https://doi.org/10.1016/j.patter.2021.100336
Ask AI
Helpful
Bookmark
Share
View Full Paper