Abstract
This paper addresses the issues of traditional authentication methods, such as the use of passwords, which often prove to be unreliable due to various vulnerabilities. The main drawbacks of these methods include the loss or theft of passwords, their weak resistance to various types of attacks, and the complexity of password management, especially in large systems. Biometric authentication methods, particularly those based on physical characteristics such as voice, present a promising alternative as they offer a higher level of security and user convenience. Biometric authentication systems have advantages over traditional methods because the voice is a unique characteristic for each person, making it substantially more challenging to forge or steal. However, there are challenges regarding the accuracy and reliability of such systems. Specifically, voice biometric systems can encounter issues related to changes in voice due to health, emotional state, or the surrounding environment. The primary objective of this paper is to compare contemporary deep learning models with traditional digital signal processing methods used for speaker recognition. For this study, text-dependent methods (Mel-Frequency Cepstral Coefficients — MFCC, Linear Predictive Coding — LPC) and text-independent methods (ECAPA-TDNN - Emphasized Channel Attention, Propagation and Aggregation in Time Delay Neural Network, ResNet - Residual Neural Network) were selected to compare their effectiveness in voice biometric authentication tasks. The experiment involved implementing biometric authentication systems based on each of the described methods and evaluating their performance on a specially collected dataset. Additionally, the paper provides a detailed examination of audio signal preprocessing methods used in voice authentication systems to ensure optimal performance in speaker recognition tasks, including noise reduction using spectral subtraction, energy normalization, enhancement filtering, framing, and windowing.
Cite
CITATION STYLE
Ruda, K., Sabodashko, D., Mykytyn, H., Shved, M., Borduliak, S., & Korshun, N. (2024). COMPARISON OF DIGITAL SIGNAL PROCESSING METHODS AND DEEP LEARNING MODELS IN VOICE AUTHENTICATION. Cybersecurity: Education, Science, Technique, 1(25), 140–160. https://doi.org/10.28925/2663-4023.2024.25.140160
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.