In this thesis, clustering and classification techniques supporting nominal data types
are studied. Clustering is the classification of objects into different groups, or the
partitioning of a dataset into clusters, so that the data in each cluster can share some
common trait - often proximity according to some defined distance measurements.
Data clustering, used in many fields, including machine learning, data mining, and
pattern recognition, is a usual technique for statistical data analysis. Feature
selection is the technique, commonly used in machine learning, of selecting a subset
of relevant features for building robust learning models. It also helps to acquire
better understanding about their data by indicating which are the important features
and how they are related with each other.
Nominal data is categorical data where the order of the categories is arbitrary.
Nominal data and numerical data are well structured data. Numerical data can be
presented in terms of geometric coordinates but nominal data cannot. Transactional
data is also called market-basket data and exhibits no dimension. Binary data is a
special data type, which can be categorized as nominal data and discrete numerical
data as well. Converting nominal data into binary data is a common practice to
make use of numerical algorithm. There are diverse types of algorithms proposed to
handle numerical data. However, there are few algorithms for handling nominal data
(non-numerical data) because there is no Euclidean distance among nominal data. In
this thesis, some newly developed clustering and feature selection algorithms which
are especially for nominal data are suggested
Clustering is a data-driven process and each cluster description describes the data
structure from various perspectives. Therefore, there are many possible but no
absolutely correct results. In this regard, cluster validity is required to evaluate the
quality of cluster descriptions. In this thesis, the concept of Singleton Item is introduced to ease the above mentioned problem. Meanwhile, a transactional cluster
validity (TCV) index is developed to quantify cluster descriptions. The search of
singleton item has a similar effect of evaluating entropy of a dataset so that the best
clusters among all the cluster descriptions are able to be determined.
Clustering is a technique that can be used as both a stand-alone tool and a
pre-processing tool while feature selection is an important pre-processing step to
reduce the number of comparatively irrelevant features for further research. Most
feature selection methods are supervised methods, which require class labels.
However, most datasets do not have class labels. Hence, it is necessary to develop
an unsupervised feature selection method. An unsupervised feature selection
method for nominal data without data transformation is therefore proposed in this
thesis. This proposed method combines the compactness and separation together
with the concept of singleton item.
Data mining has practical significance in many areas because it is the principle of
dealing with large amounts of data and picking out related information. It is
increasingly used in extracting information from the enormous datasets generated by
modern experimental and observational methods. As a new application of nominal
data analysis, “Chart of Accounts” (COA) in computational accounting system is
introduced as an example. COA is used as a nominal dataset enabling nominal
feature selection technique be used to determine the close-optimal number of
accounting segments. Moreover, text classification assigns an electronic document
to one or more categories based on its contents. With the amount of textual
information become astronomically large, performing text classification efficiently is
important in a way that it helps to accelerate the speed of searching documents from
the gigantic publication pool. A new concept of long term is proposed for improving
the classification accuracy by the method of Multi-layer Self-Organizing map
(MLSOM).
In summary, clustering and classification techniques for nominal data applications are
proposed in this thesis. The concept of Singleton Item and the proposed
transactional clustering validity (TCV) index are used to evaluate clustering results.
The unsupervised nominal feature selection scheme (UFSN) ranking the importance
of features is put forth. At last, feature ranking for the nominal data and
classification for textual data are two applications in further.
| Date of Award | 16 Feb 2009 |
|---|
| Original language | English |
|---|
| Awarding Institution | - City University of Hong Kong
|
|---|
| Supervisor | Wai Shing Tommy CHOW (Supervisor) |
|---|
- Mathematical statistics
- Cluster analysis
- Classification
Clustering and classification techniques for nominal data application
WANG, P. (Author). 16 Feb 2009
Student thesis: Doctoral Thesis