Tomsk dialect corpus: Substantiation of the concept and prospects of development

5Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

The creation of a dialectal corpus is one of the topical problems for the Tomsk Dialectology School, the oldest research center for studying the folk speech culture of Siberia. The paper describes the general concept of the corpus, the substantiation of its purposes, the characteristics of the principles of meta-markup: the objectives of developing the new resource in the near and distant future are outlined. The concept of the Tomsk Dialect Corpus is developed taking into account the key directions of the school on the study of folk speech that correlate with the achievements of the modern science of language, and with the nature of the materials available to dialectologists. The orientation of the new electronic resource can be defined as lexis- and textcentric. The main form of representation of the Middle Ob dialects in the corpus is a text with an orthographic representation of separate features of oral speech. Reliance on this principle will allow to unify the representation of the diverse archive: from the first manuscript expedition notebooks to digital audio recordings of recent years. The chosen method of representation of the sounding dialect speech can be considered universal for lexicological, linguocultural, discursive and lexicographic research. Refusal from transcription is partly compensated by the possibility of accessing the existing audio records and scanned manual records of early expeditions. The main types of meta-markup of texts entered into the corpus are developed: passport, thematic and markup by type of text. Passport meta-markup includes extra-linguistic data about the texts entered in the corpus: instructions on the place and time of the recording, information about the informant, the type of recording (by hand / from the tape / recorder), the presence / absence of audio and video files, etc. Thematic meta-markup is made on the basis of an inductive analysis of the discursive practice of old-timers, with the identification of particular topics and their generalization to macro-topics. Each topic is three levels deep maximum. The principle of "soft" thematic division of the fixed speech stream is used with the possibility of overlapping the boundaries of the extracted texts and/or simultaneous attribution of one fragment of the text to several topics. Markup by type of text at this stage implies: a) indications of text varieties that differ in the degree of the spontaneity of speech manifestations (dialogues between dialect speakers, situational inclusions arising from deviations from a purposeful conversation with dialectologists, episodic metatexts, answers to questionnaires); b) the most frequent speech genres (autobiographical story, recollection, stories about other people, stories about an event, folklore genres). The first step on the way to lexical marking will be an opportunity to give an interpretation of the meaning of nonliterary units included in the differential dictionaries of the Middle Ob region. Prospects for the development of the corpus include development of the indicated types of meta-markup, introduction of lexical markup, the integration of its data with the created electronic library of dialect dictionaries and other auxiliary resources.

Cite

CITATION STYLE

APA

Ivantsova, E. V. (2017). Tomsk dialect corpus: Substantiation of the concept and prospects of development. Voprosy Leksikografii, (11), 54–70. https://doi.org/10.17223/22274200/11/4

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free