Skip to main navigation Skip to search Skip to main content

Phishing detection with computational techniques and human effort

  • Gang LIU

Student thesis: Doctoral Thesis

Abstract

With the increasing development and popularity of e-Commerce, more and more correspondent e-Transactions are conducted online which require users to input their personal information to pass through the identification. Unfortunately, accompanied by e-Transactions, an illegal industry named as phishing emerges and catches a lot of attention in recent years. Phishing is a criminal trick of stealing personal information such as bank account, password, and credit card number through requesting people to access a fake web page which is similar to the legitimate one. Unwary online users can be easily deceived by these phishing web pages created by malicious people, also known as phishers, leading to the exposure of their sensitive information and huge financial loss. In this dissertation, we review the literature of current anti-phishing solutions and propose two novel approaches for tackling phishing problems. The first contribution is an approach to identifying the phishing target of a given (suspicious) web page by clustering a web page set consisting of its all associated web pages and the given web page itself. We first find its associated web pages, and then explore their relationships to the given web page as their features for clustering. Such relationships include link relationship, ranking relationship, text similarity, and web page layout similarity relationship. A DBSCAN clustering method is employed to find if there is a cluster around the given web page. If such cluster exists, we claim the given web page is a phishing web page and then find its phishing target (i.e., the legitimate web page it is attacking) from this cluster. Otherwise, we identify it as a legitimate web page. The second contribution is to leverage Wisdom of Crowds for improving human verification to fight phishing scams. In this research, we explore novel techniques for combating the phishing problem using computational techniques to improve human effort. Using tasks posted to the Amazon Mechanical Turk human effort market, we measure the accuracy of minimally trained humans in identifying potential phish, and consider methods for best taking advantage of individual contributions. Furthermore, we present our experiments using clustering techniques and vote weighting to improve the results of human effort in fighting phishing. We find that these techniques can increase coverage over and are significantly faster than existing blacklists used today. In summary, our theoretic and practical findings help us find the characteristics of phishing web pages. We can exploit these characteristics as the clues to fight back the phishing attacks. Two approaches are also implemented into applications. The experiments show that these approaches are effective to protect users from phishing attacks.
Date of Award3 Oct 2011
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorWenyin LIU (Supervisor)

Keywords

  • Phishing
  • Computer networks
  • Internet
  • Security measures

Cite this

'