PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 21, 200958 citations

Web spam identification through language model analysis

View Full Paper
JMJuan Martínez-RomoLALourdes Araujo

Key Points

Key points are not available for this paper at this time.

Abstract

This paper applies a language model approach to different sources of information extracted from a Web page, in order to provide high quality indicators in the detection of Web Spam. Two pages linked by a hyperlink should be topically related, even though this were a weak contextual relation. For this reason we have analysed different sources of information of a Web page that belongs to the context of a link and we have applied Kullback-Leibler divergence on them for characterising the relationship between two linked pages. Moreover, we combine some of these sources of information in order to obtain richer language models. Given the different nature of internal and external links, in our study we also distinguished these types of links getting a significant improvement in classification tasks. The result is a system that improves the detection of Web Spam on two large and public datasets such as WEBSPAM-UK2006 and WEBSPAM-UK2007.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Martínez-Romo et al. (2009) studied this question.

synapsesocial.com/papers/6a1bc376d54006be995eea7bhttps://doi.org/10.1145/1531914.1531920
Ask AI
Helpful
Bookmark
Share
View Full Paper