Automating metadata extraction: Genre classification

  • Kim Y
  • Ross S
N/ACitations
Citations of this article
12Readers
Mendeley users who have this article in their library.

Abstract

A problem that frequently arises in the management and integration of scientific data is the lack of context and semantics that would link data encoded in disparate ways. To bridge the discrepancy, it often helps to mine scientific texts to aid the understanding of the database. Mining relevant text can be significantly aided by the availability of descriptive and semantic metadata. The Digital Curation Centre (DCC) has undertaken research to automate the extraction of metadata from documents in PDF([22]). Documents may include scientific journal papers, lab notes or even emails. We suggest genre classification as a first step toward automating metadata extraction. The classification method will be built on looking at the documents from five directions; as an object of specific visual format, a layout of strings with characteristic grammar, an object with stylo-metric signatures, an object with meaning and purpose, and an object linked to previously classified objects and external sources. Some results of experiments in relation to the first two directions are described here; they are meant to be indicative of the promise underlying this multi-faceted approach. https://pdfs.semanticscholar.org/bbfd/f993d75bb03e0eede40b54e161288951ac61.pdf?_ga=2.52189161.1316584424.1508588748-1907100061.1508588748

Cite

CITATION STYLE

APA

Kim, Y., & Ross, S. (2006). Automating metadata extraction: Genre classification. In S. J. Cox (Ed.), University of Glasgow (pp. 385–+).

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free