Classifying Industrial Sectors from German Textual Data with a Domain Adapted Transformer

N/ACitations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

For economics and sociological research, lists of industries and their branches are widely used in research to categorize data and get an overview on different types of industries. However, many different taxonomies and ordering schema exist, due to different research focus but also due to different national scenarios and interests. In this paper, we will focus without loss of generality on regional data from Germany. Manual annotation of textual data is time-consuming and tedious, naturally giving rise to our initial research question, also highly inspired by questions from computational social sciences: How can we automatically categorize textual data, e.g. job advertisements or business Profiles, by industrial sectors We will present an approach towards classification using a pre-trained domain-adapted Transformer model. We find that domain-adapted models generalize better and outperform state of the art non domain-adapted Transformer models on Out-Of-Distribution data. Additionally, we open source two novel data-sets mapping textual data to WZ2008 sections and divisions, enabling further research.

Cite

CITATION STYLE

APA

Fechner, R., Dorpinghaus, J., & Firll, A. (2023). Classifying Industrial Sectors from German Textual Data with a Domain Adapted Transformer. In Proceedings of the 18th Conference on Computer Science and Intelligence Systems, FedCSIS 2023 (pp. 463–470). Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.15439/2023F6694

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free