PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 1, 1983ACM Transactions on Database Systems180 citationsOpen Access

Duplicate record elimination in large data files

View Full Paper
DBDina BittonDDDavid J. DeWitt

Key Points

Key points are not available for this paper at this time.

Abstract

The issue of duplicate elimination for large data files in which many occurrences of the same record may appear is addressed. A comprehensive cost analysis of the duplicate elimination operation is presented. This analysis is based on a combinatorial model developed for estimating the size of intermediate runs produced by a modified merge-sort procedure. The performance of this modified merge-sort procedure is demonstrated to be significantly superior to the standard duplicate elimination technique of sorting followed by a sequential pass to locate duplicate records. The results can also be used to provide critical input to a query optimizer in a relational database system.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bitton et al. (1983) studied this question.

synapsesocial.com/papers/6a20d3e410699ec7be2a9b87https://doi.org/10.1145/319983.319987
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Sorting and Searching in Multisets1976 · 79 citations
  2. 2Implementing a relational database by means of specialzed hardware1979 · 252 citations
  3. 3System R: Relational Approach to Database Management1989 · 214 citations
  4. 4System R1976 · 1,041 citations
  5. 5The Art of Computer Programming1968 · 16,195 citations