We formulate and propose the template detection problem, and suggest a practical solution for it based on counting frequent item sets. We show that the use of templates is pervasive on the web. We describe three principles, which characterize the assumptions made by hypertext information retrieval (IR) and data mining (DM) systems, and show that templates are a major source of violation of these principles. As a consequence, basic "pure" implementations of simple search algorithms coupled with template detection and elimination show surprising increases in precision at all levels of recall.
No takes yet. Share an insight, caveat, or question.
Bar-Yossef et al. (2002) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: