Abstract
This paper describes the participation of Columbus Project of Microsoft Research Asia (MSRA) in the Geo-CLEF 2006 (a cross-language geographical retrieval track which is part of Cross Language Evaluation Forum). For location extraction from the corpus, we employ a gazetteer and rule based approach. We use the MSRA's IREngine as our text search engine. Both text indexing and geo-indexing (implicit location indexing and grid indexing) are considered in our system. We only participated in the Monolingual GeoCLEF evaluation (EN-EN) and submitted five runs based on different methods, including MSRAWhitelist, MSRAManual, MSRAExpansion, MSRALocal and MSRAText. In MSRAWhitelist, we expanded the unrecognized locations (such as former Yugoslavia) to several countries manually. In MSRAManual, based on the MSRAWhitelist, we manually modified several queries since these queries are too "natural language" and the keywords of the queries seldom appear in the corpus. In MSRAExpansion, first we use the original queries to search the corpus. Then we extract the locations from the returned documents and calculate the times each location appears in the documents. Finally we will get the top 10 most frequent location names and combine them with the original geo-terms in the queries. However, this may introduce some unrelated locations. In MSRALocal, we do not use white list or query expansion method to expand the query locations. We just utilize our location extraction module to extract the locations automatically from the queries. In MSRAText, we just utilize our pure text search engine "IREngine" to process the queries. The experimental results show that MSRAManual is the best run among the five ones and then the MSRAWhitelist approach. MSRALocal and MSRAText perform similarly. The MSRAExpansion performs worst due to the introduced unrelated locations. One conclusion is that if we only extract the locations from the topics automatically, the retrieval performance does not improve significantly. Another conclusion is that automatic query expansion will weaken the performance. This is because the topics are too difficult to be handled and the corpus may be not large enough. Perhaps, the automatic query expansion may perform better in the web-scale corpus. And we find that if the queries are formed manually, the performance will be improved significantly.
© Springer-Verlag Berlin Heidelberg 2007
© Springer-Verlag Berlin Heidelberg 2007
| Original language | English |
|---|---|
| Journal | CEUR Workshop Proceedings |
| Volume | 1172 |
| DOIs | |
| Publication status | Published - 2006 |
| Externally published | Yes |
| Event | 2006 Cross Language Evaluation Forum Workshop, CLEF 2006, co-located with the 10th European Conference on Digital Libraries, ECDL 2006 - Alicante, Spain Duration: 20 Sept 2006 → 22 Sept 2006 |
Bibliographical note
Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].Research Keywords
- Geographic focus detection
- Geographic information retrieval
- Implicit location
- Location name extraction and disambiguation
Fingerprint
Dive into the research topics of 'MSRA Columbus at GeoCLEF 2006'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver