Abstract
An approach to postal address detection from webpages is proposed. The webpages are first segmented into text blocks based on their visual similarity. The text content in each block undergoes the recognition process, which employs a syntactic approach. The grammars of almost all possible patterns of postal addresses are built for this purpose. The results of our preliminary experiments on 44 webpages with 56 true addresses show that our approach can detect the postal addresses with a high precision (89.3%) and a low false alarms rate (3.8%). © 2005 IEEE.
| Original language | English |
|---|---|
| Title of host publication | Proceedings - International Workshop on Challenges in Web Information Retrieval and Integration, WIRI'05 |
| Pages | 40-45 |
| Volume | 2005 |
| DOIs | |
| Publication status | Published - 2005 |
| Event | International Workshop on Challenges in Web Information Retrieval and Integration, WIRI'05 - Tokyo, Japan Duration: 8 Apr 2005 → 9 Apr 2005 https://ieeexplore.ieee.org/xpl/conhome/10404/proceeding |
Publication series
| Name | Proceedings - International Workshop on Challenges in Web Information Retrieval and Integration, WIRI'05 |
|---|---|
| Volume | 2005 |
Conference
| Conference | International Workshop on Challenges in Web Information Retrieval and Integration, WIRI'05 |
|---|---|
| Place | Japan |
| City | Tokyo |
| Period | 8/04/05 → 9/04/05 |
| Internet address |
Bibliographical note
Publication details (e.g. title, author(s), publication statuses and dates) are captured on an “AS IS” and “AS AVAILABLE” basis at the time of record harvesting from the data source. Suggestions for further amendments or supplementary information can be sent to [email protected].Fingerprint
Dive into the research topics of 'Postal address detection from Web documents'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver