Categorizing and extracting information from multilingual HTML documents

5Citations
Citations of this article
4Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The amount of online information written in different natural languages and the number of non-English speaking Internet users have been increasing tremendously during the past decade. In order to provide high-performance access of multilingual information on the Internet, we have developed a data analysis and querying system (DatAQs) that (i) analyzes, identifies, and categorizes languages used in HTML documents, (ii) extracts information from HTML documents of interest written in different languages, (iii) allows the user to submit queries for retrieving extracted information in the same natural language provided by the query engine of DatAQs using a menu-driven user interface, and (iv) processes the user's queries (as Boolean expressions) to generate the results. DatAQs extracts information from HTML documents that belong to various data-rich, narrow-in-breadth application domains, such as car ads, house rentals, job ads, stocks, university catalogs, etc. The average F-measure on identifying HTML documents written in a particular natural language correctly is 89%, whereas the F-measure on categorizing HTML documents belonged to the car-ads application domain is 94%.

Cite

CITATION STYLE

APA

Lim, S. J., & Ng, Y. K. (2005). Categorizing and extracting information from multilingual HTML documents. In Proceedings of the International Database Engineering and Applications Symposium, IDEAS (Vol. 2005-January, pp. 415–422). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/IDEAS.2005.15

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free