Abstract
[Purposes] In order to solve the high dimension problem caused by the similarity of Euclidean distance calculation, a class-based cosine distance clustering missing value imputation approach is proposed. [Methods] Firstly, the incomplete data set is divided into two different groups (G1 and GIM); secondly, the missing data in the GIM group is pre-filled by the clustering center; the cosine distance is used again to calculate the correlation ; finally, the data with the smallest distance from the G1 group is selected to fill the missing values. [Findings] The experimental results show that the proposed method outperforms other imputation methods for both categorical and mixed datasets. [Conclusions] The CBCIM-COS method significantly improves accuracy, recall and F1-score and imputation performance.
| Translated title of the contribution | A Study of Missing Value Imputation Methods for Class-based Cosine Distance Clustering |
|---|---|
| Original language | Chinese (Simplified) |
| Pages (from-to) | 28-35 |
| Journal | 河南科技 |
| Volume | 51 |
| Issue number | 8 (总第879) |
| DOIs | |
| Publication status | Published - Apr 2024 |
| Externally published | Yes |
Research Keywords
- 不完整数据
- 缺失值插补
- 聚类
- 余弦距离
- incomplete data
- missing value imputation
- clustering
- cosine distance
Fingerprint
Dive into the research topics of 'A Study of Missing Value Imputation Methods for Class-based Cosine Distance Clustering'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver