A two-phase sampling technique to improve the accuracy of text similarities in the categorisation of hidden web databases

Yih Ling Hedley; Muhammad Younas; Anne James; Mark Sanderson

Journal Article

A two-phase sampling technique to improve the accuracy of text similarities in the categorisation of hidden web databases

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2004) 3306 516-527

DOI: 10.1007/978-3-540-30480-7_54

N/ACitations

4Readers

Get full text

Abstract

The larger amount of high quality and specialised information on the Web is stored in document databases, which is not indexed by general-purpose search engines such as Google and Yahoo. Such information is dynamically generated as a result of submitting queries to databases - which are referred to as Hidden Web databases. This paper presents a Two-Phase Sampling (2PS) technique that detects Web page templates from the randomly sampled documents of a database. It generates terms and frequencies that summarise the database content with improved accuracy. We then utilise such statistics to improve the accuracy of text similarity computation in categorisation. Experimental results show that 2PS effectively eliminates terms contained in Web page templates, and generates terms and frequencies with improved accuracy. We also demonstrate that 2PS improves the accuracy of text similarity computation required in the process of database categorisation. © Springer-Verlag 2004.

Cite

CITATION STYLE

APA

Hedley, Y. L., Younas, M., James, A., & Sanderson, M. (2004). A two-phase sampling technique to improve the accuracy of text similarities in the categorisation of hidden web databases. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 3306, 516–527. https://doi.org/10.1007/978-3-540-30480-7_54

A two-phase sampling technique to improve the accuracy of text similarities in the categorisation of hidden web databases

Abstract

Cite

Register to see more suggestions