Compressor Fault Diagnosis Knowledge: A Benchmark Dataset for Knowledge Extraction from Maintenance Log Sheets Based on Sequence Labeling

16Citations
Citations of this article
27Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Compressor fault diagnosis requires expert knowledge. Using the sequence labeling technology, this expert knowledge can be automatically extracted from compressor maintenance log sheets. Previous studies indicate that sequence labeling methods often need a substantial amount of annotation data for knowledge extraction, Unfortunately, the annotation data are very scarce in the field of compressor fault diagnosis. In this paper, we introduce a benchmark dataset for extraction of knowledge suitable for air compressor fault diagnosis. First, we collected 11,418 pieces of information from air compressor maintenance log sheets. Fault description, service requests, causes and troubleshooting solutions were stored in a dataset for data preprocessing and masking. In addition, 6196 valid text pairs were developed after the 'noises' in the raw dataset were cleaned. Second, five kinds of entities and sequences, such as equipment, faults, service requests, causes and troubleshooting solutions, were annotated by three subject experts. The annotation consistency was assessed with F1 scores. Furthermore, our proposed baseline model (or the BERT-BI-LSTM-CRF model) was compared against other five sequence labeling models (BI-LSTM-CRF, Lattice LSTM, BERT NER, ZEN, and ERNIE). The BERT-BI-LSTM-CRF model gives superior performance in extracting expert knowledge from the subject dataset. Although the baseline model is not the most cutting-edge model in the sequence labeling and named entity recognition fields, it indeed presents a great potential for compressor fault diagnosis. The dataset is available at https://github.com/chentao1999/CFDK.

Cite

CITATION STYLE

APA

Chen, T., Zhu, J., Zeng, Z., & Jia, X. (2021). Compressor Fault Diagnosis Knowledge: A Benchmark Dataset for Knowledge Extraction from Maintenance Log Sheets Based on Sequence Labeling. IEEE Access, 9, 59394–59405. https://doi.org/10.1109/ACCESS.2021.3072927

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free