Abstract
Bilingual web pages are widely used to mine translations of unknown terms. This study focused on an effective solution for obtaining relevant web pages, extracting translations with correct lexical boundaries, and ranking the translation candidates. This research adopted co-occurrence information to obtain the subject terms and then expanded the source query with the translation of the subject terms to collect effective bilingual search engine snippets. Afterwards, valid candidates were extracted fromsmall-sized, noisy bilingual corpora using an improved frequency change measurement that combines adjacent information. This research developed a method that considers surface patterns, frequency-distance, and phonetic features to elect an appropriate translation. The experimental results revealed that the proposed method performed remarkably well for mining translations of unknown terms.
Author supplied keywords
Cite
CITATION STYLE
Li, B., & Yao, J. (2019). Study on unknown term translation mining from google snippets. Information (Switzerland), 10(9). https://doi.org/10.3390/info10090267
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.