Pattern discovery has become a fundamental technique for modern information extraction tasks. This paper presents a new two-phase pattern (2PP) discovery technique for information extraction. 2PP consists of orthographic pattern discovery (OPD) and semantic pattern discovery (SPD). The OPD determines the structural features from an identified region of a document and the SPD discovers a dominant semantic pattern for the region via inference, apposition and analogy. 2PP applies discovered pattern back into the region to extract required data items through pattern matching. Experimental evaluation on a large number of identified regions indicates that our 2PP technique achieves effective results. © Springer-Verlag Berlin Heidelberg 2004.
CITATION STYLE
Ma, L., & Shepherd, J. (2004). Information extraction via automatic pattern discovery in identified region. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 3180, 232–242. https://doi.org/10.1007/978-3-540-30075-5_23
Mendeley helps you to discover research relevant for your work.