Abstract
In practice, training language models for individual authors is often expensive because of limited data resources. In such cases, Neural Network Language Models (NNLMs), generally outperform the traditional non-parametric N-gram models. Here we investigate the performance of a feedforward NNLM on an authorship attribution problem, with moderate author set size and relatively limited data. We also consider how the text topics impact performance. Compared with a well-constructed N-gram baseline method with Kneser-Ney smoothing, the proposed method achieves nearly 2.5% reduction in perplexity and increases author classification accuracy by 3.43% on average, given as few as 5 test sentences. The performance is very competitive with the state of the art in terms of accuracy and demand on test data. The source code, preprocessed datasets, a detailed description of the methodology and results are available at https://github.com/zge/authorship-attribution.
Cite
CITATION STYLE
Ge, Z., Sun, Y., & Smith, M. J. T. (2016). Authorship attribution using a neural network language model. In 30th AAAI Conference on Artificial Intelligence, AAAI 2016 (pp. 4212–4213). AAAI press. https://doi.org/10.1609/aaai.v30i1.9924
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.