ACCENT-BASED SPEECH RECOGNITION MODEL FOR THE MAJOR NIGERIAN ETHNICS ENGLISH SPEAKERS

₦ 2,000.00
i h

ABSTRACT

The introduction of accent-based automatic speech recognition (ASR) by the educational institutions to increase the quality of electronics learning (e-learning) programmes by removing the accent variations among the e-learning participants from different accents background has been considered a milestone; however, the major problem is that a generalised ASR system could not take care of the needs of all the participants stressing the need for specialised accentbased speech recognition systems that is capable of understanding and processing local or regional accents. Several accents-based speech recognition systems have been developed by researchers in America, Asia and Europe to promote accents independent e-learning participation, unfortunately, these systems are not designed to cater for African speakers’ accents variation, particularly, that of the Nigerian major ethnic groups of Yoruba, Ibo and Hausa. This study proposed an accent-based automatic speech recognition model for the three Major Nigerian Ethnics English Speakers. As part of the study’s methodology, sixty-five (65) speakers across the three major Nigerian ethnic groups were used, twenty-five (25) speakers from Igbo ethnic group, twenty (20) from the Hausa ethnic group and twenty (20) from the Yoruba ethnic group with 33 males and 32 females across different professions. The data (speakers’ speeches) were collected using specialized audio recording devices in Dutse , Jigawa State; Abakaliki, Ebonyi State; Ibadan, Oyo State and Lekki in Lagos State. Mel-Frequency Cepstral Coefficients (MFCC) feature extraction algorithm was used for the data pre-processing to extract the audio features while ensemble Restricted Boltzmann Machine-Long Short Term Memory (RBM-LSTM) model architecture was designed for the data (audio) processing using python-based open-source platform (TensorFlow). An error xiii back-propagation through time training technique was adopted for the model training using eighty percent (80%) of the study’s dataset. The performance of the model was evaluated with the remaining twenty percent (20%) of the study’s dataset to enable a robust evaluation of the model using the automatic speech recognition standard evaluation metrics of word error rate (WER) and character error rate (CER). The results of the evaluation showed that the model recorded high words accuracy rate of 76.352568% and 74.195622% on Igbo and Yoruba accents respectively, and 34.027875% words accuracy rate on Hausa accent. The model was compared with the Upadhyay and Lui (2018) model and the results showed that our proposed model performed better on all the three considered accents.

0.0 0
Write your own review Close
  • Only registered users can write reviews
*
*
  • Bad
  • Excellent
*
*
*
*
Only registered users can write reviews