Acoustic pornography recognition using fused pitch and mel-frequency cepstrum coefficients

2Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

The main objective of this paper is pornography recognition using audio features. Unlike most of the previous attempts, which have concentrated on the visual content of pornography images or videos, we propose to take advantage of sounds. Using sounds is particularly important in cases in which the visual features are not adequately informative of the contents (e.g., cluttered scenes, dark scenes, scenes with a covered body). To this end, our hypothesis is grounded in the assumption that scenes with pornographic content encompass audios with features specific to those scenes; these sounds can be in the form of speech or voice. More specifically, we propose to extract two types of features, (I) pitch and (II) mel-frequency cepstrum coefficients (MFCC), in order to train five different variations of the k-nearest neighbor (KNN) supervised classification models based on the fusion of these features. Later, the correctness of our hypothesis was investigated by conducting a set of evaluations based on a porno-sound dataset created based on an existing pornography video dataset. The experimental results confirm the feasibility of the proposed acoustic-driven approach by demonstrating an accuracy of 88.40%, an F-score of 85.20%, and an area under the curve (AUC) of 95% in the task of pornography recognition.

Cite

CITATION STYLE

APA

Banaeeyan, R., Karim, H. A., Lye, H., Fauzi, M. F. A., Mansor, S., & See, J. (2019). Acoustic pornography recognition using fused pitch and mel-frequency cepstrum coefficients. International Journal of Technology, 10(7), 1335–1343. https://doi.org/10.14716/ijtech.v10i7.3270

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free