An application of the nearest correlation matrix on web document classification

8Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

The Web document is organized by a set of textual data according to a predefined logical structure. It has been shown that collecting Web documents with similar structures can improve query efficiency. The XML document has no vectorial representation, which is required in most existing classification algorithms. The kernel method has been applied to represent structural data with pairwise similarity. In this case, a set of Web data can be fed into classification algorithms in the format of a kernel matrix. However, since the distance between a pair of Web documents is usually obtained approximately, the derived distance matrix is not a kernel matrix. In this paper, we propose to use the nearest correlation matrix (of the estimated distance matrix) as the kernel matrix, which can be fast computed by a Newton-type method. Experimental studies show that the classification accuracy can be significantly improved.

Cite

CITATION STYLE

APA

Qi, H., Xia, Z., & Xing, G. (2007). An application of the nearest correlation matrix on web document classification. Journal of Industrial and Management Optimization, 3(4), 701–713. https://doi.org/10.3934/jimo.2007.3.701

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free