Parallel Corpus of Croatian-Italian Administrative Texts

3Citations
Citations of this article
67Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Parallel corpora constitute a unique resource for providing assistance to human translators. The selection and preparation of the parallel corpora also conditions the quality of the resulting MT engine. Since Croatian is a national language and Italian is officially recognized as a minority language in seven cities and twelve municipalities of Istria County, a large amount of parallel texts is produced on a daily basis. However, there have been no attempts in using these texts for compiling a parallel corpus. A domain-specific sentence-aligned parallel Croatian-Italian corpus of administrative texts would be of high value in creating different language tools and resources. The aim of this paper is, therefore, to explore the value of parallel documents which are publicly available mostly in pdf format and to investigate the use of automatically-built dictionaries in corpus compilation. The effects that a document format and, consequently sentence splitting, and the dictionary input have on the sentence alignment process are manually evaluated.

Cite

CITATION STYLE

APA

Bakaric, M. B., & Pacelat, I. L. (2019). Parallel Corpus of Croatian-Italian Administrative Texts. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 11–18). Incoma Ltd. https://doi.org/10.26615/issn.2683-0078.2019_002

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free