TY - GEN
T1 - Mutual Information Based K-Labelsets Ensemble for Multi-Label Classification
AU - Wang, Ran
AU - Kwong, Sam
AU - Jia, Yuheng
AU - Huang, Zhiqi
AU - Wu, Lang
PY - 2018/7
Y1 - 2018/7
N2 - For solving multi-label classification problems, traditional random K-labelsets method has two main drawbacks: 1) the randomly selected label set may result in highly imbalanced data for single-label multi-class learning, and 2) the dependency relations among different labels in the same label set may cause serious information redundancy and overlap. Both of these two drawbacks can affect the generalization capability of the multi-label learner. In order to overcome these two problems, in this paper, we propose a K-labelsets ensemble method based on mutual information and joint entropy. First, the mutual information and joint entropy are adopted to evaluate the redundancy level and imbalance level of each K-labelset. Then, disjoint sampling is performed iteratively, where during each iteration, a number of K-labelsets with low mutual information are retained to be the candidates, and the one with the highest joint entropy is selected. Afterwards, for each selected K-labelset, the label powerset method is employed and a multi-class classification model is constructed. Finally, the multi-class models on different K-labelsets are integrated, and a voting based ensemble model is generated to perform the predictions for unseen samples. We conduct extensive experiments on real-world multi-label data sets. Experimental results demonstrate the effectiveness of the proposed method.
AB - For solving multi-label classification problems, traditional random K-labelsets method has two main drawbacks: 1) the randomly selected label set may result in highly imbalanced data for single-label multi-class learning, and 2) the dependency relations among different labels in the same label set may cause serious information redundancy and overlap. Both of these two drawbacks can affect the generalization capability of the multi-label learner. In order to overcome these two problems, in this paper, we propose a K-labelsets ensemble method based on mutual information and joint entropy. First, the mutual information and joint entropy are adopted to evaluate the redundancy level and imbalance level of each K-labelset. Then, disjoint sampling is performed iteratively, where during each iteration, a number of K-labelsets with low mutual information are retained to be the candidates, and the one with the highest joint entropy is selected. Afterwards, for each selected K-labelset, the label powerset method is employed and a multi-class classification model is constructed. Finally, the multi-class models on different K-labelsets are integrated, and a voting based ensemble model is generated to perform the predictions for unseen samples. We conduct extensive experiments on real-world multi-label data sets. Experimental results demonstrate the effectiveness of the proposed method.
KW - Classification
KW - K-labelsets
KW - Multi-label learning
KW - Mutual information
UR - https://www.scopus.com/pages/publications/85060496272
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-85060496272&origin=recordpage
U2 - 10.1109/fuzz-ieee.2018.8491677
DO - 10.1109/fuzz-ieee.2018.8491677
M3 - RGC 32 - Refereed conference paper (with host publication)
T3 - IEEE International Conference on Fuzzy Systems
BT - 2018 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE) - 2018 Proceedings
PB - IEEE
T2 - 2018 IEEE International Conference on Fuzzy Systems, FUZZ 2018
Y2 - 8 July 2018 through 13 July 2018
ER -