TY - GEN
T1 - A systematic study of parameter correlations in large scale duplicate document detection
AU - Ye, Shaozhi
AU - Wen, Ji-Rong
AU - Ma, Wei-Ying
N1 - Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].
PY - 2006
Y1 - 2006
N2 - Although much work has been done on duplicate document detection (DDD) and its applications, we observe the absence of a systematic study of the performance and scalability of large-scale DDD. It is still unclear how various parameters of DDD, such as similarity threshold, precision/recall requirement, sampling ratio, document size, correlate mutually. In this paper, correlations among several most important parameters of DDD are studied and the impact of sampling ratio is of most interest since it heavily affects the accuracy and scalability of DDD algorithms. An empirical analysis is conducted on a million documents from the TREC .GOV collection. Experimental results show that even using the same sampling ratio, the precision of DDD varies greatly on documents with different size. Based on this observation, an adaptive sampling strategy for DDD is proposed, which minimizes the sampling ratio within the constraint of a given precision threshold. We believe the insights from our analysis are helpful for guiding the future large scale DDD work. © Springer-Verlag Berlin Heidelberg 2006.
AB - Although much work has been done on duplicate document detection (DDD) and its applications, we observe the absence of a systematic study of the performance and scalability of large-scale DDD. It is still unclear how various parameters of DDD, such as similarity threshold, precision/recall requirement, sampling ratio, document size, correlate mutually. In this paper, correlations among several most important parameters of DDD are studied and the impact of sampling ratio is of most interest since it heavily affects the accuracy and scalability of DDD algorithms. An empirical analysis is conducted on a million documents from the TREC .GOV collection. Experimental results show that even using the same sampling ratio, the precision of DDD varies greatly on documents with different size. Based on this observation, an adaptive sampling strategy for DDD is proposed, which minimizes the sampling ratio within the constraint of a given precision threshold. We believe the insights from our analysis are helpful for guiding the future large scale DDD work. © Springer-Verlag Berlin Heidelberg 2006.
UR - https://www.scopus.com/pages/publications/33745789442
UR - https://www.scopus.com/record/pubmetrics.uri?eid=2-s2.0-33745789442&origin=recordpage
U2 - 10.1007/11731139_33
DO - 10.1007/11731139_33
M3 - RGC 32 - Refereed conference paper (with host publication)
SN - 3540332065
SN - 9783540332060
VL - 3918 LNAI
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 275
EP - 284
BT - Advances in Knowledge Discovery and Data Mining - 10th Pacific-Asia Conference, PAKDD 2006, Proceedings
PB - Springer Verlag
T2 - 10th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, PAKDD 2006
Y2 - 9 April 2006 through 12 April 2006
ER -