Skip to main navigation Skip to search Skip to main content

MSRA Columbus at GeoCLEF 2006

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

This paper describes the participation of Columbus Project of Microsoft Research Asia (MSRA) in the Geo-CLEF 2006 (a cross-language geographical retrieval track which is part of Cross Language Evaluation Forum). For location extraction from the corpus, we employ a gazetteer and rule based approach. We use the MSRA's IREngine as our text search engine. Both text indexing and geo-indexing (implicit location indexing and grid indexing) are considered in our system. We only participated in the Monolingual GeoCLEF evaluation (EN-EN) and submitted five runs based on different methods, including MSRAWhitelist, MSRAManual, MSRAExpansion, MSRALocal and MSRAText. In MSRAWhitelist, we expanded the unrecognized locations (such as former Yugoslavia) to several countries manually. In MSRAManual, based on the MSRAWhitelist, we manually modified several queries since these queries are too "natural language" and the keywords of the queries seldom appear in the corpus. In MSRAExpansion, first we use the original queries to search the corpus. Then we extract the locations from the returned documents and calculate the times each location appears in the documents. Finally we will get the top 10 most frequent location names and combine them with the original geo-terms in the queries. However, this may introduce some unrelated locations. In MSRALocal, we do not use white list or query expansion method to expand the query locations. We just utilize our location extraction module to extract the locations automatically from the queries. In MSRAText, we just utilize our pure text search engine "IREngine" to process the queries. The experimental results show that MSRAManual is the best run among the five ones and then the MSRAWhitelist approach. MSRALocal and MSRAText perform similarly. The MSRAExpansion performs worst due to the introduced unrelated locations. One conclusion is that if we only extract the locations from the topics automatically, the retrieval performance does not improve significantly. Another conclusion is that automatic query expansion will weaken the performance. This is because the topics are too difficult to be handled and the corpus may be not large enough. Perhaps, the automatic query expansion may perform better in the web-scale corpus. And we find that if the queries are formed manually, the performance will be improved significantly.
© Springer-Verlag Berlin Heidelberg 2007
Original languageEnglish
JournalCEUR Workshop Proceedings
Volume1172
DOIs
Publication statusPublished - 2006
Externally publishedYes
Event2006 Cross Language Evaluation Forum Workshop, CLEF 2006, co-located with the 10th European Conference on Digital Libraries, ECDL 2006 - Alicante, Spain
Duration: 20 Sept 200622 Sept 2006

Bibliographical note

Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].

Research Keywords

  • Geographic focus detection
  • Geographic information retrieval
  • Implicit location
  • Location name extraction and disambiguation

Fingerprint

Dive into the research topics of 'MSRA Columbus at GeoCLEF 2006'. Together they form a unique fingerprint.

Cite this