Robust phishing URL detection using FastText embeddings and deep sequential models
Shah Noor, Sibghat Ullah Bazai, Muhammad Imran Ghafoor, Alamgir Naushad, Abid Mehmood, Qazi Mudassar Ilyas, Muhammad Nasir Mumtaz BhuttaPhishing attacks are increasingly prevalent, leading to the loss of assets, financial information, personal data, and other sensitive information to malicious actors. Various detection methods have been developed to protect users from such attacks. Traditional phishing Uniform Resource Locator (URL) detection methods, such as blacklisting and heuristics, rely on maintaining lists of known phishing URLs. However, these techniques struggle to identify new unlisted URLs and are expensive to maintain. In recent years, there has been a shift toward deep learning techniques, which are particularly efficient for training large data systems and handling undefined features. This research aims to develop and compare seven deep-learning models to identify phishing URLs more accurately and efficiently. The proposed model utilizes FastText-based word embeddings to preprocess input URLs, generating a robust vector representation. This model is then applied to various deep learning-based classification models, including Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM) network, Bidirectional LSTM (BiLSTM), CNN-LSTM hybrid model, CNN-BiLSTM hybrid model, Gated Recurrent Unit (GRU), and Bidirectional GRU (BiGRU). The experimental results demonstrate that the BiLSTM model with FastText embeddings outperforms alternative models, attaining an accuracy of 98.62%. The proposed technique was tested on benchmark data from an online repository and then validated on an independent external dataset to assess its resilience and generalization capabilities, yielding good phishing-detection performance across multiple data sources. Even though some earlier research has achieved an accuracy above 99%, many of these studies ignore token-level URL structures, avoid hybrid architectures, or depend on small or unbalanced datasets. To overcome these constraints, we combine FastText embeddings with BiLSTM networks to improve generalization and robustness, especially for identifying obfuscated and unique phishing URLs. This results in a more dependable and scalable detection system.