TY - GEN
T1 - Customer information system data pre-processing with feature selection techniques for non-technical losses prediction in an electricity market
AU - Nizar, Anisah Hanim
AU - Zhao, Jun Hua
AU - Dong, Zhao Yang
N1 - Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].
PY - 2006
Y1 - 2006
N2 - Non-technical losses (NTL) identification and prediction are important tasks for many utilities. Data from customer information system (CIS) can be used for NTL analysis. However, in order to accurately and efficiently perform NTL analysis, the original data from CIS need to be pre-processed before any detailed NTL analysis can be carried out. In this paper, we propose a feature selection based method for CIS data pre-processing in order to extract the most relevant information for further analysis such as clustering and classifications. By removing irrelevant and redundant features, feature selection is an essential step in data mining process in finding optimal subset of features to improve the quality of result by giving faster time processing, higher accuracy and simpler results with fewer features. Detailed feature selection analysis is presented in the paper. Both time-domain and load shape data are compared based on the accuracy, consistency and statistical dependencies between features. ©2006 IEEE.
AB - Non-technical losses (NTL) identification and prediction are important tasks for many utilities. Data from customer information system (CIS) can be used for NTL analysis. However, in order to accurately and efficiently perform NTL analysis, the original data from CIS need to be pre-processed before any detailed NTL analysis can be carried out. In this paper, we propose a feature selection based method for CIS data pre-processing in order to extract the most relevant information for further analysis such as clustering and classifications. By removing irrelevant and redundant features, feature selection is an essential step in data mining process in finding optimal subset of features to improve the quality of result by giving faster time processing, higher accuracy and simpler results with fewer features. Detailed feature selection analysis is presented in the paper. Both time-domain and load shape data are compared based on the accuracy, consistency and statistical dependencies between features. ©2006 IEEE.
KW - Data mining
KW - Feature selection
KW - Load profiling
KW - Non-technical loss (NTL) analysis
UR - https://www.scopus.com/pages/publications/43549091594
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-43549091594&origin=recordpage
U2 - 10.1109/ICPST.2006.321964
DO - 10.1109/ICPST.2006.321964
M3 - RGC 32 - Refereed conference paper (with host publication)
SN - 1424401119
SN - 9781424401116
T3 - 2006 International Conference on Power System Technology, POWERCON2006
BT - 2006 International Conference on Power System Technology, POWERCON2006
PB - IEEE
T2 - 2006 International Conference on Power System Technology, POWERCON2006
Y2 - 22 October 2006 through 26 October 2006
ER -