Robust Detection of Phishing URLs Under Obfuscation Attacks

Keywords: Data Augmentation, Canonicalization, Machine Learning, Obfuscation, Phishing, Web Address

Abstract

Purpose. Based on a dual-view approach, a lexical detection protocol for phishing web addresses was developed and experimentally validated under syntactic transformations that preserve the destination host within the adopted threat model.

Method. The study is empirical. An open dataset of phishing and legitimate web addresses compiled by Prasad and Chandra was used. After data cleaning, 235,366 unique records remained. The training, validation, and test subsets were partitioned by registrable domain. A character-level baseline model, a dual-view model, and the proposed model trained with data augmentation were compared. The threat model included host case variation, explicit default ports, percent encoding, fragments, trailing dots, user-info prefixes, removable path segments, and combinations of these transformations. Model performance was assessed using precision, recall, F1-score, area under the receiver operating characteristic curve (ROC-AUC), and false-positive rate.

Findings. On the clean, domain-disjoint test set, the baseline and augmentation-trained models achieved F1-scores of 0.9962 and 0.9960, respectively; under the baseline-adaptive transfer-testing scenario, the corresponding values were 0.9943 and 0.9961.

Theoretical implications. The study extends methodological approaches to evaluating lexical phishing detectors by substantiating the joint use of raw and canonical web address representations and invariance-oriented training within a formalized obfuscation-adversary model.

Practical implications. The proposed protocol is suitable for local pre-filtering of web addresses in email gateways, browser extensions, and monitoring systems without requiring network queries. It can also support security-policy decisions concerning web-address filtering.

Originality/value. The study demonstrates that performance on clean data alone is insufficient to substantiate claims of robustness within the adopted threat model. It further shows that training-time augmentation improves invariance to the studied class of semantically equivalent obfuscations without causing a practically significant degradation in performance on clean data.

Limitations/future research. Further temporal, cross-dataset, and browser-dependent evaluations are required before deployment in operational systems.

Paper type. Empirical.

Downloads

Download data is not yet available.

References

Aljofey, A., Jiang, Q., Qu, Q., Huang, M., & Niyigena, J.-P. (2020). An effective phishing detection model based on character level convolutional neural network from URL. Electronics, 9(9), 1514. https://doi.org/10.3390/electronics9091514.

Anti-Phishing Working Group. (2026). Phishing activity trends report: 4th quarter 2025. https://docs.apwg.org/reports/apwg_trends_report_q4_2025.pdf.

Arp, D., Quiring, E., Pendlebury, F., Warnecke, A., Pierazzi, F., Wressnegger, C., Cavallaro, L., & Rieck, K. (2022). Dos and don’ts of machine learning in computer security. In 31st USENIX Security Symposium (pp. 3971–3988). https://www.usenix.org/conference/usenixsecurity22/presentation/arp.

Berners-Lee, T., Fielding, R., & Masinter, L. (2005). Uniform resource identifier (URI): Generic syntax (RFC 3986). RFC Editor. https://doi.org/10.17487/RFC3986.

Biggio, B., & Roli, F. (2018). Wild patterns: Ten years after the rise of adversarial machine learning. Pattern Recognition, 84, 317–331. https://doi.org/10.1016/j.patcog.2018.07.023.

Bozkir, A. S., Dalgic, F. C., & Aydos, M. (2023). GramBeddings: A new neural network for URL based identification of phishing web pages through n-gram embeddings. Computers & Security, 124, 102964. https://doi.org/10.1016/j.cose.2022.102964.

Efron, B., & Tibshirani, R. J. (1993). An introduction to the bootstrap. Chapman & Hall. https://doi.org/10.1201/9780429246593.

European Union Agency for Cybersecurity. (2025). ENISA threat landscape 2025. Publications Office of the European Union. https://doi.org/10.2824/1946374.

Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. https://doi.org/10.1016/j.patrec.2005.10.010.

Fielding, R., Nottingham, M., & Reschke, J. (2022). HTTP semantics (RFC 9110). RFC Editor. https://doi.org/10.17487/RFC9110.

Ghalechyan, H., Israyelyan, E., Arakelyan, A., Hovhannisyan, G., & Davtyan, A. (2024). Phishing URL detection with neural networks: An empirical study. Scientific Reports, 14, 25134. https://doi.org/10.1038/s41598-024-74725-6.

Gupta, B. B., Yadav, K., Razzak, I., Psannis, K., Castiglione, A., & Chang, X. (2021). A novel approach for phishing URLs detection using lexical based machine learning in a real-time environment. Computer Communications, 175, 47–57. https://doi.org/10.1016/j.comcom.2021.04.023.

Hannousse, A., & Yahiouche, S. (2021). Towards benchmark datasets for machine learning based website phishing detection: An experimental study. Engineering Applications of Artificial Intelligence, 104, 104347. https://doi.org/10.1016/j.engappai.2021.104347.

Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804. https://doi.org/10.1016/j.patter.2023.100804.

Kaufman, S., Rosset, S., Perlich, C., & Stitelman, O. (2012). Leakage in data mining. ACM Transactions on Knowledge Discovery from Data, 6(4), 1–21. https://doi.org/10.1145/2382577.2382579.

Kwon, H., & Lee, J. (2026). Crafting evasive phishing URLs: Exploiting tokenizer vulnerabilities in transformer-based detection systems. International Journal of Machine Learning and Cybernetics, 17(8), 390. https://doi.org/10.1007/s13042-026-03239-6.

Maneriker, P., Stokes, J. W., Lazo, E. G., Carutasu, D., Tajaddodianfar, F., & Gururajan, A. (2021). URLTran: Improving phishing URL detection using transformers. In MILCOM 2021—IEEE Military Communications Conference (pp. 197–204). https://doi.org/10.1109/MILCOM52596.2021.9653028.

McNemar, Q. (1947). Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2), 153–157. https://doi.org/10.1007/BF02295996.

Pendlebury, F., Pierazzi, F., Jordaney, R., Kinder, J., & Cavallaro, L. (2019). TESSERACT: Eliminating experimental bias in malware classification across space and time. In 28th USENIX Security Symposium (pp. 729–746). https://www.usenix.org/conference/usenixsecurity19/presentation/pendlebury

Pillai, M. J., Remya, S., Devika, V., Ramasubbareddy, S., & Cho, Y. (2024). Evasion attacks and defense mechanisms for machine learning-based web phishing classifiers. IEEE Access, 12, 19375–19387. https://doi.org/10.1109/ACCESS.2023.3342840.

Prasad, A., & Chandra, S. (2023). PhiUSIIL phishing URL dataset (Version 2) [Data set]. Mendeley Data. https://doi.org/10.17632/shwpxscxy2.2.

Prasad, A., & Chandra, S. (2024). PhiUSIIL: A diverse security profile empowered phishing URL detection framework based on similarity index and incremental learning. Computers & Security, 136, 103545. https://doi.org/10.1016/j.cose.2023.103545.

Sabir, B., Babar, M. A., Gaire, R., & Abuadbba, A. (2022). Reliability and robustness analysis of machine learning based phishing URL detectors. IEEE Transactions on Dependable and Secure Computing. https://doi.org/10.1109/TDSC.2022.3218043.

Safi, A., & Singh, S. (2023). A systematic literature review on phishing website detection techniques. Journal of King Saud University—Computer and Information Sciences, 35(2), 590–611. https://doi.org/10.1016/j.jksuci.2023.01.004.

Sahingoz, O. K., Buber, E., Demir, O., & Diri, B. (2019). Machine learning based phishing detection from URLs. Expert Systems with Applications, 117, 345–357. https://doi.org/10.1016/j.eswa.2018.09.029.

Sahoo, D., Liu, C., & Hoi, S. C. H. (2017). Malicious URL detection using machine learning: A survey. arXiv. https://doi.org/10.48550/arXiv.1701.07179.

Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432.

Tamal, M. A., Islam, M. K., Bhuiyan, T., & Sattar, A. (2024). Dataset of suspicious phishing URL detection. Frontiers in Computer Science, 6, 1308634. https://doi.org/10.3389/fcomp.2024.1308634.

Tang, L., & Mahmoud, Q. H. (2022). A deep learning-based framework for phishing website detection. IEEE Access, 10, 1509–1521. https://doi.org/10.1109/ACCESS.2021.3137636.

Weinberger, K., Dasgupta, A., Langford, J., Smola, A., & Attenberg, J. (2009). Feature hashing for large scale multitask learning. In Proceedings of the 26th International Conference on Machine Learning (pp. 1113–1120). https://doi.org/10.1145/1553374.1553516.


Abstract views: 10
PDF Downloads: 3
Published
2026-08-31
How to Cite
Feshchenko, R., Vitvitska, K., & Luchyk, V. (2026). Robust Detection of Phishing URLs Under Obfuscation Attacks, 16(4), 192-207. https://doi.org/10.33445/sds.2026.16.4.10
Section
Information and Cyber Security