Perceptual Features Based Continuous Speech Recognition in Additive Noise Environment Using Various Modelling Techniques
Parole chiave:
Hidden Markov model (HMM), Neural network, Continuous density HMM (CHMM), Gaussian mixture model (GMM), Speech recognition, Vector quantization (VQ), Mel frequency perceptual linear predictive cepstrum (MFPLPC), Noise, Wavelet transform, Recursive least sAbstract
The main objective of this paper is to discuss the effectiveness of Mel frequency perceptual features and the noise reduction technique in evaluating the performance of multi speaker independent continuous speech recognition system in additive noise environment by using various modelling techniques. The proposed perceptual features are captured and trained using clustering technique, GMM, continuous density HMM and back propagation neural networks. Speech recognition system is evaluated on clean and noisy test speeches by using minimum distance criterion, maximum log likelihood criterion and minimum mean squared error criterion for these speech modelling techniques. Performance of these features is tested on speeches randomly chosen from “TIMIT” speech corpus. This algorithm provides 98.3, 90.8, 99.3 and 97.3% as accuracy for clean speech recognition system evaluated on GMM models and continuous density HMM models, Clustering models and neural network models respectively. Evaluation is done for 300 test speeches. System is also tested on noisy test speech by considering RLS and combination of RLS & wavelets as additional pre- processing techniques and the performance is found to be better than the testing without additional pre- processing. Noises such as babble, M109, white, destroyer engine, paper rustle, F16, factory and destroyerops are considered in this work and these are added to the test speeches at various levels and performance of the system is evaluated for 300 test speeches. Among the various modelling techniques, clustering technique and continuous density HMM technique are performing better for clean speech recognition and noisy speech recognition respectively. GMM and continuous density HMM along with additional pre-processing technique are performing better for noisy speech recognition even if the noise energy is greater than the signal energy. Keywords: Hidden Markov model (HMM), Neural network, Continuous density HMM (CHMM), Gaussian mixture model (GMM), Speech recognition, Vector quantization (VQ), Mel frequency perceptual linear predictive cepstrum (MFPLPC), Noise, Wavelet transform, Recursive least sPubblicato
Fascicolo
Sezione
Licenza
Declaration and Copyright Transfer Form
(to be completed by authors)
I/ We, the undersigned author(s) of the submitted manuscript, hereby declare, that the above manuscript which is submitted for publication in the STM Journals(s), is not published already in part or whole (except in the form of abstract) in any journal or magazine for private or public circulation, and, is not under consideration of publication elsewhere.
· I/We will not withdraw the manuscript after 1 week of submission as I have read the Author Guidelines and will adhere to the guidelines.
· I/We Author(s ) have niether given nor will give this manuscript elsewhere for publishing after submitting in STM Journal(s).
· I/ We have read the original version of the manuscript and am/ are responsible for the thought contents embodied in it. The work dealt in the manuscript is my/ our own, and my/ our individual contribution to this work is significant enough to qualify for authorship.
· I/We also agree to the authorship of the article in the following order:
Author’s name
1. ________________
2. ________________
3. ________________
_______________
We Author(s) tick this box and would request you to consider it as our signature as we agree to the terms of this Copyright Notice, which will apply to this submission if and when it is published by this journal. |