PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 1, 2000454 citationsOpen Access

An experimental comparison of naive Bayesian and keyword-based anti-spam filtering with personal e-mail messages

IAIon AndroutsopoulosJKJohn KoutsiasKCKonstantinos V. Chandrinos

Key Points

Key points are not available for this paper at this time.

Abstract

The growing problem of unsolicited bulk e-mail, also known as “spam”, has generated a need for reliable anti-spam e-mail filters. Filters of this type have so far been based mostly on manually constructed keyword patterns. An alternative approach has recently been proposed, whereby a Naive Bayesian classifier is trained automatically to detect spam messages. We test this approach on a large collection of personal e-mail messages, which we make publicly available in “encrypted” form contributing towards standard benchmarks. We introduce appropriate cost-sensitive measures, investigating at the same time the effect of attribute-set size, training-corpus size, lemmatization, and stop lists, issues that have not been explored in previous experiments. Finally, the Naive Bayesian filter is compared, in terms of performance, to a filter that uses keyword patterns, and which is part of a widely used e-mail reader.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Androutsopoulos et al. (2000) studied this question.

synapsesocial.com/papers/6a0da83fcae7912d2fa52910https://doi.org/10.1145/345508.345569
Ask AI
Helpful
Bookmark
Share
View Full Paper