Skip to main navigation Skip to search Skip to main content

Incorporating word embeddings in the hierarchical dirichlet process for query-oriented text summarization

Research output: Chapters, Conference Papers, Creative and Literary WorksRGC 32 - Refereed conference paper (with host publication)peer-review

Abstract

The ever-growing amount of textual data available online creates the need for automatic text summarization tools. Probabilistic topic models are able to infer semantic relationships between sentences which is a key step of extractive summarization methods. However, they strongly rely on word co-occurrence patterns and fail to capture the actual semantic relationships between words such as synonymy, antonymy, etc. We propose a novel algorithm which incorporates pre-trained word embeddings in the probabilistic topic model in order to capture semantic similarities between sentences. These similarities provide the basis for a sentence ranking algorithm for query-oriented summarization. The summary is then produced by extracting highly ranked sentences from the original corpus. Our method is shown to outperform state-of-the-art algorithms on a benchmark dataset.
Original languageEnglish
Title of host publicationProceedings - 2017 IEEE 15th International Conference on Industrial Informatics (INDIN)
PublisherIEEE
Pages1037-1042
ISBN (Electronic)9781538608371
ISBN (Print)9781538608388
DOIs
Publication statusPublished - Jul 2017
EventIEEE 15th International Conference on Industrial Informatics INDIN 2017 : The Undergoing Industrial Informatics R-Evolution - Emden, Germany
Duration: 24 Jul 201726 Jul 2017
http://www.indin2017.i2ar.de/

Publication series

NameIEEE International Conference on Industrial Informatics (INDIN)
Volume2017
ISSN (Electronic)2378-363X

Conference

ConferenceIEEE 15th International Conference on Industrial Informatics INDIN 2017 : The Undergoing Industrial Informatics R-Evolution
PlaceGermany
CityEmden
Period24/07/1726/07/17
Internet address

Fingerprint

Dive into the research topics of 'Incorporating word embeddings in the hierarchical dirichlet process for query-oriented text summarization'. Together they form a unique fingerprint.

Cite this