Establishment of Parallel Text Corpus of Equipment Manufacturing Industry Based on Data Mining Technology

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

In the era of language big data, traditional data analysis methods can't analyze semi-structured or unstructured data such as text, but all the contents in the equipment manufacturing corpus belong to text data. The equipment manufacturing corpus is a linguistic information base for legal activities and equipment manufacturing research, which aims to study equipment manufacturing and collect equipment manufacturing cases. At present, the construction of legal database in China is not perfect, and there are still many problems. In this paper, a method based on template transformation is proposed to automatically acquire parallel corpus on the Internet, and a method based on the number of transformation patterns and the retrieval and sorting of transformation patterns is adopted to verify bilingual parallel texts. This system can build a large-scale parallel corpus of equipment manufacturing industry by automatically acquiring a large number of parallel texts from the Internet.

Cite

CITATION STYLE

APA

Liu, D., Liu, J., Zhang, X., & Chen, M. (2021). Establishment of Parallel Text Corpus of Equipment Manufacturing Industry Based on Data Mining Technology. In Journal of Physics: Conference Series (Vol. 1881). IOP Publishing Ltd. https://doi.org/10.1088/1742-6596/1881/4/042091

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free