PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 9, 2025Engineering Technology & Applied Science Research8 citationsOpen Access

A Hybrid Heuristic-Machine Learning Framework for Phishing Detection Using Multi-Domain Feature Analysis

View Full Paper
AJAshvini JadhavPCPankaj Chandre

Key Points

  • The framework achieves an impressive accuracy exceeding 95% on real-world datasets, enhancing phishing detection reliability and efficiency.
  • Utilizing feature selection techniques such as Recursive Feature Elimination contributes to improved model performance while reducing noise.
  • Machine learning classifiers like XGBoost and Random Forest excel, with testing accuracies of 96.7% and 99.85% on specific datasets.
  • The composite suspicion score methodology marries heuristic and ML predictions for a comprehensive phishing detection system.

Abstract

This study introduces a hybrid phishing detection framework that combines machine learning with heuristic rule-based techniques to provide accurate, scalable, and policy-compliant detection across a variety of phishing types. The proposed method uses diverse datasets, including URL patterns, email headers, and HTML content, organized in a layered manner, allowing flexible analysis even when some features are missing. Feature selection techniques, such as variance thresholding and Recursive Feature Elimination (RFE), are applied to improve learning efficiency and reduce noise. Several classifiers, including Random Forest (RF), XGBoost, Gradient Boosting (GB), and CatBoost, are trained on optimized features, and their outputs are combined using voting to boost overall reliability. The system also includes a rule-based engine aligned with India's national Email Policy, incorporating heuristic checks such as non-government domains, missing authentication (SPF/DKIM/DMARC), use of insecure protocols, foreign IPs, phishing URLs, and other threat indicators. Each rule is weighted and contributes to a composite suspicion score, which is explainable and policy-mapped. These heuristic signals are used both directly and as features for the machine learning models, allowing for layered, interpretable AI. The final phishing score balances the contribution of both heuristic and ML predictions and is compared against an optimized threshold to determine whether an input is phishing or safe. Experimental results on benchmark datasets demonstrate that heuristic-guided feature selection, combined with hybrid data integration, significantly improves performance, achieving an average accuracy exceeding 95% in real-world datasets. Individual models, including CatBoost and XGBoost, demonstrated outstanding performance, achieving training accuracies of up to 100% and testing accuracies of 96.7% and 96.4%, respectively, for URL datasets. For email header analysis, RF achieved the highest accuracy at 99.85%. The findings underscore the significance of feature engineering in developing scalable and reliable phishing detection systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jadhav et al. (2025) studied this question.

synapsesocial.com/papers/68e70db790569dd607ee651ehttps://doi.org/10.48084/etasr.11548
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Advancing Email Spam Classification using Machine Learning and Deep Learning Techniques2024 · 14 citations
  2. 2Can serious gaming tactics bolster spear-phishing and phishing resilience? : Securing the human hacking in Information Security2024 · 25 citations
  3. 3A Survey of Machine Learning-Based Solutions for Phishing Website Detection2021 · 187 citations
  4. 4Look before you leap: Detecting phishing web pages by exploiting raw URL and HTML characteristics2023 · 117 citations
  5. 5Enhancing Phishing Email Detection through Ensemble Learning and Undersampling2023 · 33 citations