Abstract
This study aims to develop an automated system for processing scientific texts using advanced NLP techniques. The methodology integrates classical NLP methods with deep learning approaches, employing SciBERT for text classification, LDA for topic modeling, and a modified TextRank algorithm for keyword extraction. Results demonstrate high accuracy in document classification (F1-score of 0.92), effective topic identification, and precise keyword extraction. The developed web interface showcases the system's practical applicability. This research contributes to the field by presenting a comprehensive solution for scientific text analysis, combining state-of-the-art language models with established NLP techniques. The study's novelty lies in its tailored approach to scientific literature, addressing the unique challenges of domain-specific language and complex content structure in academic texts.
Cite
CITATION STYLE
Starukhin, Y., & Diukarev, V. (2024). AUTOMATION OF TEXT DATA PROCESSING USING NLP. The American Journal of Engineering and Technology, 6(7), 24–39. https://doi.org/10.37547/tajet/volume06issue07-04
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.