Abstract
This paper proposes an integrated approach to automatic information extraction for Forums, Blogs and News web sites using wrapper. This paper presents a tree alignment and transfer learning method to generate the wrapper. The tree alignment algorithm is adopted to find the best matching structure of the input web pages. A kind of linear regression method is employed to get the weight of different tag-matching. For wrapper maintenance, this paper presents a method using a log likelihood ratio test for detecting the change points on the similarity series which gotten from the wrapper and input web pages. Experimental results show that the method achieves high accuracy and has steady performance. © 2011 AICIT.
Cite
CITATION STYLE
Xia, Y., Yang, Y., Ge, F., Zhang, S., & Yu, H. (2011). An integrated approach for information extraction. In Proceedings - 5th International Conference on New Trends in Information Science and Service Science, NISS 2011 (Vol. 1, pp. 122–127).
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.