Abstract
Abstract With the exponential growth of textual data across diverse domains, the task of efficiently modelling and clustering large-scale text has emerged as a key challenge in natural language processing (NLP). Conventional text representation approaches, such as Term Frequency-Inverse Document Frequency (TF-IDF) and Bag-of-Words (BoW), often fall short in capturing semantic nuances. This limitation has encouraged the adoption of more advanced techniques, including word embeddings (e.g., Word2Vec, GloVe) and transformer-based models like BERT and GPT. Similarly, traditional clustering algorithms such as K-Means and Hierarchical Clustering often struggle with the high dimensionality and sparsity inherent in text data. Consequently, models like Latent Dirichlet Allocation (LDA), Non-negative Matrix Factorization (NMF), and deep learning-based clustering frameworks have gained popularity. This review paper presents a comprehensive overview of recent machine learning-based text representation and semantic clustering techniques, examining their performance, scalability, and relevance across applications. It also outlines persisting challenges such as interpretability, noise handling, and computational overhead, while identifying potential research directions to enhance semantic clustering in large-scale text environments. Keywords: Semantic Clustering, Text Representation, Word Embeddings, Transformer Models, Deep Learning in NLP, Text Mining.
Cite
CITATION STYLE
Khan, D. (2025). Modeling and Semantic Clustering in Large-scale Text Data: A Review of Machine Learning Techniques and Applications. INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT, 09(04), 1–9. https://doi.org/10.55041/ijsrem46510
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.