Document expansion using relevant web documents for spoken document retrieval

Ryo Masumura, Akinori Ito, Yu Uno, Masashi Ito, Shozo Makino

Research output: Chapter in Book/Report/Conference proceedingConference contribution

3 Citations (Scopus)

Abstract

Recently, automatic indexing of a spoken document using a speech recognizer attracts attention. However, index generation from an automatic transcription has many problems because the automatic transcription has many recognition errors and Out-Of-Vocabulary words. To solve this problem, we propose a document expansion method using Web documents. To obtain important keywords which included in the spoken document but lost by recognition errors, we acquire Web documents relevant to the spoken document. Then, an index of the spoken document is generated by combining an index that generated from the automatic transcription and the Web documents. We propose a method for retrieval of relevant documents, and the experimental result shows that the retrieved Web document contained many OOV words. Next, we propose a method for combining the recognized index and the Web index. The experimental result shows that the index of the spoken document generated by the document expansion was closer to an index from the manual transcription than the index generated by the conventional method. Finally, we conducted a spoken document retrieval experiment, and the document-expansion-based index gave better retrieval precision than the conventional indexing method.

Original languageEnglish
Title of host publicationProceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE, 2010
DOIs
Publication statusPublished - 2010 Nov 29
Event6th International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2010 - Beijing, China
Duration: 2010 Aug 212010 Aug 23

Publication series

NameProceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2010

Other

Other6th International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2010
CountryChina
CityBeijing
Period10/8/2110/8/23

ASJC Scopus subject areas

  • Computer Science (miscellaneous)
  • Computer Science Applications

Fingerprint Dive into the research topics of 'Document expansion using relevant web documents for spoken document retrieval'. Together they form a unique fingerprint.

  • Cite this

    Masumura, R., Ito, A., Uno, Y., Ito, M., & Makino, S. (2010). Document expansion using relevant web documents for spoken document retrieval. In Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE, 2010 [5587854] (Proceedings of the 6th International Conference on Natural Language Processing and Knowledge Engineering, NLP-KE 2010). https://doi.org/10.1109/NLPKE.2010.5587854