PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 26, 2012EURASIP Journal on Information Security72 citationsOpen Access

phishGILLNET—phishing detection methodology using probabilistic latent semantic analysis, AdaBoost, and co-training

VRVenkatesh RamanathanHWHarry Wechsler

Key Points

Key points are not available for this paper at this time.

Abstract

Identity theft is one of the most profitable crimes committed by felons. In the cyber space, this is commonly achieved using phishing. We propose here robust server side methodology to detect phishing attacks, called phishGILLNET, which incorporates the power of natural language processing and machine learning techniques. phishGILLNET is a multi-layered approach to detect phishing attacks. The first layer (phishGILLNET1) employs Probabilistic Latent Semantic Analysis (PLSA) to build a topic model. The topic model handles synonym (multiple words with similar meaning), polysemy (words with multiple meanings), and other linguistic variations found in phishing. Intentional misspelled words found in phishing are handled using Levenshtein editing and Google APIs for correction. Based on term document frequency matrix as input PLSA finds phishing and non-phishing topics using tempered expectation maximization. The performance of phishGILLNET1 is evaluated using PLSA fold in technique and the classification is achieved using Fisher similarity. The second layer of phishGILLNET (phishGILLNET2) employs AdaBoost to build a robust classifier. Using probability distributions of the best PLSA topics as features the classifier is built using AdaBoost. The third layer (phishGILLNET3) further expands phishGILLNET2 by building a classifier from labeled and unlabeled examples by employing Co-Training. Experiments were conducted using one of the largest public corpus of email data containing 400,000 emails. Results show that phishGILLNET3 outperforms state of the art phishing detection methods and achieves F -measure of 100%. Moreover, phishGILLNET3 requires only a small percentage (10%) of data be annotated thus saving significant time, labor, and avoiding errors incurred in human annotation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ramanathan et al. (2012) studied this question.

synapsesocial.com/papers/6a6fe78da7fbea1e4407e1cfhttps://doi.org/10.1186/1687-417x-2012-1
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Proceedings of the 11th annual SIGCSE conference on Innovation and technology in computer science education2006 · 26 citations
  2. 2Lexical URL analysis for discriminating phishing and legitimate websites2011 · 38 citations
  3. 3Implementing a Web Browser with Phishing Detection Techniques2011 · 11 citations
  4. 4Antiphishing through Phishing Target Discovery2011 · 42 citations
  5. 5Probabilistic Latent Semantic Indexing2017 · 4,063 citations