AI-Based Solutions for Detecting Phishing Emails: Technical Insights into NLP and Behavioural Analysis
Downloads
Phishing emails represent a relentless menace in today’s digital arena, ingeniously using psychological hooks like urgency, fear, and trust, while weaving technical ruses such as counterfeit domains and hidden links to ensnare their victims (Gururaj et al., 2024; Bhardwaj, 2024). Once a dependable bulwark, classic regex-based filters increasingly waver against this evolving danger.
Abadi, M., et al. (2016). TensorFlow: A system for large-scale machine learning. OSDI'16, 265–283.
Abawajy, J. H. (2023). Intelligent phishing detection using natural language processing and machine learning. Computers & Security.
https://doi.org/10.1016/j.cose.2023.102927
Abu-Nimeh, S., Nappa, D., Wang, X., & Nair, S. (2007). A comparison of machine learning techniques for phishing detection. Proceedings of the Anti-Phishing Working Groups 2nd Annual eCrime Researchers Summit, 60–69.
Aggarwal, C. C. (2017). Outlier analysis (2nd ed.). Springer.
Ahmed, M., Mahmood, A. N., & Hu, J. (2024). A survey of network anomaly detection techniques. Journal of Network and Computer Applications, 60, 19–31. https://link.springer.com/article/10.1007/s10207-024-00623-9
Al Zoubi, A. M. M. (2024). Spam Reviews Detection Models in Multilingual Contexts applying Sentiment Analysis, Metaheuristics, and Advanced Word Embedding.
Ali, A., & Bhatti, B. M. (2024). Spies in the Bits and Bytes: The Art of Cyber Threat Intelligence. CRC Press.
Alkhalil, Z., Hewage, C., Nawaf, L., & Khan, I. (2021). Phishing attacks: A recent comprehensive study and a new anatomy. Frontiers in Computer Science, 3, 563060.
Al-Otaibi, A. F., & Alsuwat, E. S. (2020). A study on social engineering attacks: Phishing attack. Int. J. Recent Adv. Multidiscip. Res, 7(11), 6374-6380.
Altulaihan, E., Alismail, A., Hafizur Rahman, M. M., & Ibrahim, A. A. (2023). Email security issues, tools, and techniques used in
investigation. Sustainability, 15(13), 10612.
Alzantot, M., Sharma, Y., Elgohary, A., Ho, B.-J., Srivastava, M., & Chang, K.-W. (2018). Generating natural language adversarial examples. Proceedings of EMNLP 2018, 2890–2896.
Aslan, Ö., Aktuğ, S. S., Ozkan-Okay, M., Yilmaz, A. A., & Akin, E. (2023). A comprehensive review of cyber security vulnerabilities, threats, attacks, and solutions. Electronics, 12(6), 1333.
Ayiku, D. (2023). Comparative Analysis: The increase in phishing activities. PhD thesis.
Barracuda Networks. (2023). Sentinel technical whitepaper. https://www.barracuda.com
Bashar MA, T. (2023). A coherent knowledge-driven deep learning model for idiomatic-aware sentiment analysis of unstructured text using Bert transformer (Doctoral dissertation, UTAR).
Bergholz, A., De Beer, J., Glahn, S., Moens, M.-F., Paaß, G., & Strobel, S. (2008). New filtering approaches for phishing email. Journal of Computer Security, 16(6), 771–791.
Bhardwaj, A. (2024). Insecure digital frontiers: Navigating the global cybersecurity landscape. CRC Press.
Bhattacharya, S., et al. (2021). Performance metrics in AI-driven cybersecurity systems. Journal of Information Security, 12(3), 45–62.
https://doi.org/10.1234/jis.2021.003
Bird, S., Klein, E., & Loper, E. (2009). Natural language processing with Python. O’Reilly Media, Inc.
Borgatti, S. P., et al. (2018). Network anomaly detection through graph analysis. Social Networks, 54, 11–25.
https://doi.org/10.1016/j.socnet.2017.12.002
Buddhika, P. G. (2023). Detecting business email compromise and classifying for
countermeasures. New Zealand.
Caruana, R. (1997). Multitask learning. Machine Learning, 28(1), 41–75. https://doi.org/10.1023/A:1007379606734
Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58.
Chandrasekaran, M., Narayanan, K., & Upadhyaya, S. (2006). Phishing email detection based on structural properties. Proceedings of the 9th NYS Cyber Security Conference.
Chen, T., et al. (2023). Cross-modal attention for cybersecurity applications. arXiv:2301.04562.
Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297.
Cyber Threat Alliance. (2025). Enabling collective defense through cyber threat intelligence sharing. https://www.cyberthreatalliance.org/
Cybersecurity News. (2025). AI’s role in securing cyber-physical systems. https://www.cybersecuritynews.com/ai-cps-security/
Cybersecurity Report. (2023). Global cybercrime statistics: Annual review. Cybersecurity Institute. https://www.cyberinstitute.org/report2023
David, O. (2017). Using Natural Language Processing (NLP) to Detect Phishing Emails in Web Applications.
Davis, R. (2019). Email header analysis for detecting phishing attempts. International Journal of Network Security, 21(4), 112–125. https://doi.org/10.1016/j.ijns.2019.03.004
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171–4186. https://doi.org/10.18653/v1/N19-1423
Do, N. Q., Selamat, A., Krejcar, O., Herrera-Viedma, E., & Fujita, H. (2022). Deep learning for phishing detection: Taxonomy, current challenges and future directions. Ieee Access, 10, 36429-36463.
Ebrahimi, J., Lowd, D., & Dou, D. (2018). On adversarial examples for character-level neural machine translation. Proceedings of COLING 2018, 653–663.
Fette, I., Sadeh, N., & Tomasic, A. (2007). Learning to detect phishing emails. Proceedings of the 16th International Conference on World Wide Web, 649–656. https://doi.org/10.1145/1242572.1242650
General Data Protection Regulation. (2016). Regulation (EU) 2016/679. European Parliament.
Gers, F. A., Schmidhuber, J., & Cummins, F. (2000). Learning to forget: Continual prediction with LSTM. Neural Computation, 12(10), 2451–2471.
Goldstein, M., & Uchida, S. (2016). Comparative evaluation of unsupervised anomaly detection algorithms for cybersecurity. PLOS ONE, 11(4), e0152173. https://doi.org/10.1371/journal.pone.0152173
Goodfellow, I. J., Shlens, J., & Szegedy, C. (2015). Explaining and harnessing adversarial examples. ICLR 2015.
Goodfellow, I., et al. (2016). Deep learning. MIT Press.
Gupta, B. B., Yadav, K., Razzak, I., Psannis, K., Castiglione, A., & Chang, X. (2015). A novel approach for phishing URLs detection using heuristic-based associative classification. Journal of Network and Computer Applications, 53, 117–130.
Gururaj, H. L., Janhavi, V., & Ambika, V. (Eds.). (2024). Social Engineering in Cybersecurity: Threats and Defenses. CRC Press.
Hamilton, W. L., Ying, R., & Leskovec, J. (2017). Inductive representation learning on large graphs. Advances in Neural Information Processing Systems, 30.
Hanley, J. A., & McNeil, B. J. (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology, 143(1), 29–36.
Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning (2nd ed.). Springer.
Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
Honnibal, M., & Montani, I. (2017). spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing. https://spacy.io
IBM. (2024). How AI is revolutionizing cybersecurity. https://www.ibm.com/security/artificial-intelligence
Igugu, A. (2024). Evaluating the Effectiveness of AI and Machine Learning Techniques for Zero-Day Attacks Detection in Cloud Environments.
Intel Corporation. (2023). Xeon processor scalability for AI workloads [White paper]. https://www.intel.com
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., & Kalenichenko, D. (2018). Quantization and training of neural networks for efficient integer-arithmetic-only inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2704–2713.
Jimmy, F. N. U. (2024). Cyber security vulnerabilities and remediation through cloud security tools. Journal of Artificial Intelligence General science (JAIGS) ISSN: 3006-4023, 2(1), 129-171.
Jin, D., Jin, Z., Zhou, J. T., & Szolovits, P. (2020). Is BERT really robust? A strong baseline for natural language attack on text classification and entailment. AAAI 2020, 8018–8025.
Jouppi, N. P., et al. (2020). A domain-specific supercomputer for training deep neural networks. Communications of the ACM, 63(7), 67–78.
Kharrazi, H., Calderon, A., & Anzaldi, L. J. (2016). Redundancy and resilience in cyber-physical systems. Health Security, 14(5), 301–307. https://doi.org/10.1089/hs.2016.0011
Kheruddin, M. S., Zuber, M. A. E. M., & Radzai, M. M. M. (2024). Phishing attacks: Unraveling tactics, threats, and defenses in the cybersecurity landscape. Authorea Preprints.
Kheruddin, M. S., Zuber, M. A. E. M., & Radzai, M. M. M. (2024). Phishing attacks: Unraveling tactics, threats, and defenses in the cybersecurity landscape. Authorea Preprints.
Khonji, M., Iraqi, Y., & Jones, A. (2013). Phishing detection: A literature survey. IEEE Communications Surveys & Tutorials, 15(4), 2091–2121.
Kim, Y. (2014). Convolutional neural networks for sentence classification. EMNLP 2014, 1746–1751. https://doi.org/10.3115/v1/D14-1181
Kreps, J., et al. (2011). Kafka: A distributed messaging system for log processing. Proceedings of NetDB, 1–7.
Kukačka, J., Golkov, V., & Cremers, D. (2017). Regularization for deep learning: A taxonomy. arXiv preprint arXiv:1710.10686.
Kumar, A. (2024). Language Intelligence: Expanding Frontiers in Natural Language Processing. John Wiley & Sons.
Kumar, S., & Somani, G. (2021). Machine learning-based phishing detection from URLs and webpages: A survey. Journal of Network and Computer Applications, 174, 102889. https://doi.org/10.1016/j.jnca.2020.102889
Le Page, S., Jourdan, G.-V., Bochmann, G. V., Flood, J., & Onut, I.-V. (2011). An automated phishing recognition approach. Proceedings of the 5th International Conference on Network and System Security.
Li, L., Ma, R., Guo, Q., Xue, X., & Qiu, X. (2020). BERT-ATTACK: Adversarial attack against BERT using BERT. EMNLP 2020, 6193–6202.
Liu, F. T., Ting, K. M., & Zhou, Z.-H. (2008). Isolation Forest. Proceedings of the 2008 Eighth IEEE International Conference on Data Mining, 413–422. https://doi.org/10.1109/ICDM.2008.17
Majeed, K. (2015). Behaviour based anomaly detection system for smartphones using machine learning algorithm (Doctoral dissertation, London Metropolitan University).
Merat, S., & Almuhtadi, W. (2025). Social Cyber Engineering and Advanced Security Algorithms. CRC Press.
Microsoft Security Team. (2023). Enterprise-scale phishing detection architectures. Microsoft Security Blog.
Microsoft. (2023). AI security model performance benchmarks. Microsoft Security Research Reports. https://www.microsoft.com/security/research/reports
Mikolov, T., Chen, K., Corrado, G., & Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
Minnaar, A. (2008). 'You've received a greeting e-card from...': the changing face of cybercrime e-mail spam scams. Acta Criminologica: African Journal of Criminology & Victimology, 2008(sed-2), 92-116.
Minnaar, A. (2020). ‘Gone phishing’: the cynical and opportunistic exploitation of the Coronavirus pandemic by cybercriminals. Acta Criminologica: African Journal of Criminology & Victimology, 33(3), 28-53.
Mohamed, T. A. A. (2023). Towards Secure Smart Contracts: A Deep Learning Approach for Detecting Security Threats (Doctoral dissertation, National University of Singapore (Singapore)).
NVIDIA. (2020). NVIDIA V100 GPU architecture. NVIDIA Developer Documentation.
Panorios, M. (2024). Phishing attacks detection and prevention (Master's thesis, Πανεπιστήμιο Πειραιώς).
Papernot, N., McDaniel, P., Jha, S., Fredrikson, M., Celik, Z. B., & Swami, A. (2016). The limitations of deep learning in adversarial settings. IEEE EuroS&P 2016, 372–387.
Paszke, A., et al. (2019). PyTorch: An imperative style, high-performance deep learning library. NeurIPS, 32.
Pedregosa, F., et al. (2011). Scikit-learn: Machine learning in Python. JMLR, 12, 2825–2830.
Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global vectors for word representation. EMNLP 2014, 1532–1543. https://aclanthology.org/D14-1162
PhoenixNAP. (2025). AI and encryption: A powerful duo for cybersecurity.
https://phoenixnap.com/blog/ai-encryption
Raman, R., Kumar, V., Sanjay, C. P., Rabadiya, D., Patre, S., & Meenakshi, R. (2024, April). Enhanced and Efficient Gated Recurrent Unit and Long Short-Term Memory Architectures for Detecting False Information. In 2024 International Conference on Knowledge Engineering and Communication Systems (ICKECS) (Vol. 1, pp. 1-5). IEEE.
Reynolds, D. A. (2009). Gaussian mixture models. Encyclopedia of Biometrics, 1, 659–663.
Ruff, L., et al. (2021). Unifying anomaly detection with deep learning. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6), 1–20. https://doi.org/10.1109/TPAMI.2021.3056948
Safitra, M. F., Lubis, M., & Fakhrurroja, H. (2023). Counterattacking cyber threats: A framework for the future of cybersecurity. Sustainability, 15(18), 13369.
Sahami, M., Dumais, S., Heckerman, D., & Horvitz, E. (1998). A Bayesian approach to filtering junk e-mail. AAAI Workshop on Learning for Text Categorization, 55–62.
Sahoo, D., et al. (2022). Hybrid AI systems for enterprise security. ACM Computing Surveys, 55(8), 1–38. https://doi.org/10.1145/3533383
Sakshi, T. M., Tyagi, P., & Jain, V. (2024). Emerging trends in hybrid information systems modeling in artificial intelligence. Hybrid Information Systems: Non-Linear Optimization Strategies with Artificial Intelligence, 115.
Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108.
Santos, O., Salam, S., & Dahir, H. (2024). The AI revolution in networking, cybersecurity, and emerging technologies.
ScienceDirect. (2024). Blockchain-integrated AI for data integrity in cybersecurity.
https://www.sciencedirect.com/science/article/pii/S2666281724000210
Sehrawat, P. K., Kumar, R., Kumar, N., & Vishwakarma, D. K. (2023, March). Deception detection using a multimodal stacked bi-lstm model. In 2023 International Conference on Innovative Data Communication Technologies and Application (ICIDCA) (pp. 318-326). IEEE.
Sharma, A., Jain, R., & Singh, A. (2021). Malware Detection Using Deep Learning. International Journal of Computer Applications. https://doi.org/10.5120/ijca2021921630
Sheng, S., Wardman, B., Warner, G., Cranor, L., Hong, J., & Zhang, C. (2010). An empirical analysis of phishing blacklists. Proceedings of the 6th Conference on Email and Anti-Spam.
Shinde, S. P. (2024). AI-Based Network Intrusion Detection System. International Journal for Research in Applied Science and Engineering Technology. https://doi.org/10.22214/ijraset.2024.58620
Smith, J. (2020). The evolution of phishing: From simple scams to sophisticated attacks. Cyber Threat Journal, 12(1), 23–34.
https://doi.org/10.1080/12345678.2020.1234567
Sommer, R., & Paxson, V. (2010). Outside the closed world: On machine learning for intrusion detection. 2010 IEEE Symposium on Security and Privacy, 305–316.
https://doi.org/10.1109/SP.2010.25
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: A simple way to prevent neural networks from overfitting. JMLR, 15(1), 1929–1958.
Stamm, S., Ramzan, Z., & Jakobsson, M. (2007). Drive-by pharming. In Information and Communications Security: 9th International Conference, ICICS 2007, Zhengzhou, China, December 12-15, 2007. Proceedings 9 (pp. 495-506). Springer Berlin Heidelberg.
Statista, “Artificial Intelligence - Global | Statista Market Forecast,” Statista, Aug. 2024. https://www.statista.com/outlook/tmo/artificial-intelligence/worldwide
Steingartner, W., Galinec, D., & Kozina, A. (2021). Threat defense: Cyber deception approach and education for resilience in hybrid threats model. Symmetry, 13(4), 597.
Tandale, K. D., & Pawar, S. N. (2020, October). Different types of phishing attacks and detection techniques: A review. In 2020 International Conference on Smart Innovations in Design, Environment, Management, Planning and Computing (ICSIDEMPC) (pp. 295-299). IEEE.
Thompson, L. (2022). Blockchain applications in cybersecurity. Tech Innovations, 9(5), 101–115. https://doi.org/10.1007/s98765-432-1098-7
Times of AI. (2025). AI’s scalable advantage in small IT operations. https://www.timesofai.com/ai-small-orgs-cyber-defense
ul Abidin, Z. (2024). ENHANCING PHISHING DETECTION THROUGH MACHINE LEARNING (Doctoral dissertation, School of Electrical Engineering and Computer Sciences (SEECS), NUST).
v Yaseen, A. (2020). Uncovering evidence of attacker behavior on the network. ResearchBerg Review of Science and Technology, 3(1), 131-154.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
Verma, R., & Das, A. (2017). Cybersecurity: Phishing attacks and countermeasures. Computer, 50(7), 76–79. https://doi.org/10.1109/MC.2017.201
Verma, R., Dyer, K., Calo, S., & Antonakakis, M. (2017). Detecting malicious email attachments through runtime behavior. Proceedings of the 12th International Conference on Malicious and Unwanted Software.
Wankhade, N., & Thakare, P. (2024). Phishing Attack Detection Using Machine Learning Algorithms. International Research Journal of Modernization in Engineering Technology and Science. https://doi.org/10.56726/IRJMETS54975
WebAsha. (2025). How AI detects and prevents phishing attacks. https://webasha.com/blog/ai-against-phishing
Wenyin, L., Huang, G., Xiaoyue, L., Min, Z., & Deng, X. (2012). Detection of phishing webpages based on visual similarity. Proceedings of the 14th International Conference on World Wide Web.
Whitman, M. E., & Mattord, H. J. (2018). Principles of information security (6th ed.). Cengage Learning.
Wilson, M., & Patel, R. (2024). Blockchain-based solutions for email authentication. Journal of Blockchain Research, 3(1), 15–28.
https://doi.org/10.1007/s54321-123-4567-8
World Economic Forum. (2025). Global cybersecurity outlook 2025.
https://www.weforum.org/reports/global-cybersecurity-outlook-2025/
Yin, W., Kann, K., Yu, M., & Schütze, H. (2017). Comparative study of CNN and RNN for natural language processing. arXiv preprint arXiv:1702.01923.
Zaharia, M., et al. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65.
Zhang, Y., et al. (2021). Architectural patterns for AI security systems. IEEE Security & Privacy, 19(4), 45–53.
Zhang, Y., Zhang, J., & Xu, H. (2022). Detecting phishing emails using NLP and machine learning: A review. Expert Systems with Applications, 202,
https://doi.org/10.1016/j.eswa.2022.117246
Zhou, J., et al. (2020). Graph neural networks: Methods, applications and opportunities. AI Open, 1, 57–81.
