Skip to main navigation Skip to search Skip to main content

Clustering and classification techniques for nominal data application

  • Piyang WANG

Student thesis: Doctoral Thesis

Abstract

In this thesis, clustering and classification techniques supporting nominal data types are studied. Clustering is the classification of objects into different groups, or the partitioning of a dataset into clusters, so that the data in each cluster can share some common trait - often proximity according to some defined distance measurements. Data clustering, used in many fields, including machine learning, data mining, and pattern recognition, is a usual technique for statistical data analysis. Feature selection is the technique, commonly used in machine learning, of selecting a subset of relevant features for building robust learning models. It also helps to acquire better understanding about their data by indicating which are the important features and how they are related with each other. Nominal data is categorical data where the order of the categories is arbitrary. Nominal data and numerical data are well structured data. Numerical data can be presented in terms of geometric coordinates but nominal data cannot. Transactional data is also called market-basket data and exhibits no dimension. Binary data is a special data type, which can be categorized as nominal data and discrete numerical data as well. Converting nominal data into binary data is a common practice to make use of numerical algorithm. There are diverse types of algorithms proposed to handle numerical data. However, there are few algorithms for handling nominal data (non-numerical data) because there is no Euclidean distance among nominal data. In this thesis, some newly developed clustering and feature selection algorithms which are especially for nominal data are suggested Clustering is a data-driven process and each cluster description describes the data structure from various perspectives. Therefore, there are many possible but no absolutely correct results. In this regard, cluster validity is required to evaluate the quality of cluster descriptions. In this thesis, the concept of Singleton Item is introduced to ease the above mentioned problem. Meanwhile, a transactional cluster validity (TCV) index is developed to quantify cluster descriptions. The search of singleton item has a similar effect of evaluating entropy of a dataset so that the best clusters among all the cluster descriptions are able to be determined. Clustering is a technique that can be used as both a stand-alone tool and a pre-processing tool while feature selection is an important pre-processing step to reduce the number of comparatively irrelevant features for further research. Most feature selection methods are supervised methods, which require class labels. However, most datasets do not have class labels. Hence, it is necessary to develop an unsupervised feature selection method. An unsupervised feature selection method for nominal data without data transformation is therefore proposed in this thesis. This proposed method combines the compactness and separation together with the concept of singleton item. Data mining has practical significance in many areas because it is the principle of dealing with large amounts of data and picking out related information. It is increasingly used in extracting information from the enormous datasets generated by modern experimental and observational methods. As a new application of nominal data analysis, “Chart of Accounts” (COA) in computational accounting system is introduced as an example. COA is used as a nominal dataset enabling nominal feature selection technique be used to determine the close-optimal number of accounting segments. Moreover, text classification assigns an electronic document to one or more categories based on its contents. With the amount of textual information become astronomically large, performing text classification efficiently is important in a way that it helps to accelerate the speed of searching documents from the gigantic publication pool. A new concept of long term is proposed for improving the classification accuracy by the method of Multi-layer Self-Organizing map (MLSOM). In summary, clustering and classification techniques for nominal data applications are proposed in this thesis. The concept of Singleton Item and the proposed transactional clustering validity (TCV) index are used to evaluate clustering results. The unsupervised nominal feature selection scheme (UFSN) ranking the importance of features is put forth. At last, feature ranking for the nominal data and classification for textual data are two applications in further.
Date of Award16 Feb 2009
Original languageEnglish
Awarding Institution
  • City University of Hong Kong
SupervisorWai Shing Tommy CHOW (Supervisor)

Keywords

  • Mathematical statistics
  • Cluster analysis
  • Classification

Cite this

'