Comparison of feature extraction for accent dependent Thai speech recognition system

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

This paper aims to compare the feature extraction methods for accent dependent Thai speech from three regions including central, southern and northeastern regions. We investigate four frequency analysis methods: i.e., Energy Spectral Density (ESD), Power Spectral Density (PSD), Mel-Frequency Cepstral Coefficients (MFCC) and Spectrogram (SPT). Radial basis function kernel based on support vector machine is used as a classifier with 5-fold cross validation. The isolated speech data sets are recorded from 30 male and 30 female participants speaking the 10 Thai digits from 0 to 9. The MFCC-based feature gives better accuracy than ESD, PSD and SPT respectively. For within the same region, the MFCC-based feature provides average accuracy of 94.9% and 99.1% for male and female voices respectively. For the three regions, the MFCC-based feature provides average accuracy of 89.34% and 93.81% for male and female voices, respectively.

Description

Keywords

Accents classification, Energy spectral density, Feature extraction, Mel-frequency cepstral coefficients, Power spectral density, Spectrogram, Support vector machine

Citation

2018 IEEE 7th International Conference on Communications and Electronics Icce 2018, 322-325, 2018

Collections

Endorsement

Review

Supplemented By

Referenced By