Abstract
This paper reports comparative authorship attribution results obtained on the Internet comments of the morphologically complex Lithuanian language. We have explored the impact of machine learning and similarity-based approaches on the different author set sizes (containing 10, 100, and 1,000 candidate authors), feature types (lexical, morphological, and character), and feature selection techniques (feature ranking, random selection). The authorship attribution task was complicated due to the used Lithuanian language characteristics, nonnormative texts, an extreme shortness of these texts, and a large number of candidate authors. The best results were achieved with the machine learning approaches. On the larger author sets the entire feature set composed of word-level character tetra-grams demonstrated the best performance.
Cite
CITATION STYLE
Kapociute-Dzikiene, J., Venckauskas, A., & Damasevicius, R. (2017). A comparison of authorship attribution approaches applied on the Lithuanian language. In Proceedings of the 2017 Federated Conference on Computer Science and Information Systems, FedCSIS 2017 (pp. 347–351). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.15439/2017F110
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.