Automatic extraction of subcorpora based on subcategorization frames from a part-of-speech tagged corpus

13Citations
Citations of this article
85Readers
Mendeley users who have this article in their library.

Abstract

This paper presents a method for extracting subcorpora documenting different subcategorization frames for verbs, nouns, and adjectives in the 100 mio. word British National Corpus. The extraction tool consists of a set of batch files for use with the Corpus Query Processor (CQP), which is part of the IMS corpus workbench (cf. Christ 1994a,b). A macroprocessor has been developed that allows the user to specify in a simple input file which subcorpora are to be created for a given lemma. The resulting subcorpora can be used (1) to provide evidence for the subcategorization properties of a given lemma, and to facilitate the selection of corpus lines for lexicographic research, and (2) to determine the frequencies of different syntactic contexts of each lemma.

Cite

CITATION STYLE

APA

Gahl, S. (1998). Automatic extraction of subcorpora based on subcategorization frames from a part-of-speech tagged corpus. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 428–432). Association for Computational Linguistics (ACL). https://doi.org/10.3115/980845.980918

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free