PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 20260 citationsOpen Access

Intelligent Phishing Website Detection Using Machine Learning And URL Feature Analysis

View Full Paper
MSMrs. V. SuvarnaMSMallidi Mohana SudhaGVGunturi Satyasai Phani Amrutha Sri Varshini

Key Points

  • The research aims to develop a machine learning framework for detecting phishing websites based on URL features.
  • Utilized a dataset of legitimate and phishing URLs from publicly available sources.
  • Applied data preprocessing techniques to ensure dataset consistency.
  • Implemented multiple machine learning algorithms including Random Forest and AdaBoost.
  • Evaluated models using stratified cross-validation for reliable prediction.
  • Conducted feature importance analysis to identify key attributes influencing detection.
  • The Random Forest classifier achieved high detection accuracy.
  • Ensemble learning models showed superior performance in distinguishing between legitimate and phishing sites.
  • Evaluation metrics indicated high accuracy, precision, recall, and F1-score for the proposed framework.
  • The system effectively identifies fraudulent URLs, enhancing user protection.

Abstract

Phishing attacks have become one of the most common cybersecurity threats, targeting users by creating fraudulent websites that mimic legitimate platforms to steal sensitive information such as login credentials, financial data, and personal identity details. Traditional phishing detection approaches, such as blacklist-based systems and manual verification methods, are often inefficient and unable to detect newly emerging phishing websites in real time. Therefore, intelligent and automated detection mechanisms are required to improve cybersecurity and protect users from online fraud. This study proposes an efficient machine learning–based framework for detecting phishing websites using URL and domain-based features. The proposed system utilizes a dataset containing both legitimate and phishing website URLs collected from publicly available repositories. Data preprocessing techniques are applied to clean and normalize the dataset, ensuring consistency and improving model performance. Multiple machine learning algorithms including Logistic Regression, Decision Tree, Random Forest, AdaBoost, and Gradient Boosting are implemented and evaluated using stratified cross-validation techniques to ensure reliable prediction results. Among the evaluated models, ensemble learning algorithms demonstrate superior performance due to their ability to combine multiple weak learners and reduce prediction errors. In particular, the Random Forest classifier achieves high detection accuracy by analyzing key URL characteristics such as domain name structure, prefix and suffix usage, DNS records, URL length, and IP address patterns. The experimental results show that the ensemble model effectively distinguishes between legitimate and phishing websites with high accuracy, precision, recall, and F1-score.Furthermore, feature importance analysis is performed to identify the most influential attributes contributing to phishing detection, enabling better understanding of model behaviour and improving system transparency. The proposed framework provides a scalable and automated solution for detecting malicious websites, helping users identify fraudulent URLs before interacting with them. Overall, the proposed machine learning framework enhances phishing detection capability, improves cybersecurity awareness, and provides an efficient tool for protecting users against online phishing attacks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Suvarna et al. (2026) studied this question.

synapsesocial.com/papers/69e07d732f7e8953b7cbe5e9https://doi.org/10.5281/zenodo.19563893
Ask AI
Helpful
Bookmark
Share
View Full Paper