Multi-domain text classification via linguistic and semantic feature integration

0Citations
Citations of this article
28Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Text classification remains a fundamental task in natural language processing, with applications spanning sentiment analysis, spam detection, and hate speech identification. However, its performance is often limited when relying exclusively on either handcrafted linguistic features or semantic embedding representations in isolation. In real-world scenarios, text often exhibits high variability in style, structure, and context, making it challenging for single-representation approaches to capture both syntactic nuances and deeper semantic relationships. This limitation can lead to reduced robustness and generalization, particularly when models are deployed across various different tasks. This study proposes a hybrid feature fusion framework that integrates interpretable linguistic features extracted using the Linguistic Feature Toolkit with advanced semantic embeddings derived from Doc2Vec and transformer-based model. By combining syntactic structures with deep contextual representations, the approach aims to capture both surface-level and semantic nuances of textual data. The framework is evaluated on five benchmark datasets spanning three critical domains: Fake News Detection, Bloom’s Taxonomy Classification, and hate speech detection. Extensive experiments using multiple machine learning classifiers demonstrate that the fusion of linguistic and semantic features consistently outperforms single-feature baselines across all domains. The Bidirectional Encoder Representations from Transformer linguistic feature fusion approach achieved accuracies of up to 81% for Fake News Detection, 67% for Bloom’s Taxonomy classification, and 72% for HSD, with corresponding improvements in precision, recall, and F1-score. These findings confirm the effectiveness of integrating linguistic interpretability with deep semantic modeling, offering a robust and domain-agnostic solution for advancing text classification performance. While the study does not perform explicit cross-domain transfer experiments, it provides a comprehensive multi-domain benchmarking framework and quantifies domain shift across diverse datasets.

Cite

CITATION STYLE

APA

Hashmi, E., Shaikh, S., Yayilgan, S. Y., Abomhara, M., Akerkar, R., & Afzal, M. (2025). Multi-domain text classification via linguistic and semantic feature integration. Social Network Analysis and Mining, 15(1). https://doi.org/10.1007/s13278-025-01551-7

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free