Comparison of feature extraction for accent dependent Thai speech recognition system

dc.contributor.authorTantisatirapong, Suchada
dc.contributor.authorPrasoproek, Chalisa
dc.contributor.authorPhothisonothai, Montri
dc.date.accessioned2026-08-06T10:21:41Z
dc.date.available2026-08-06T10:21:41Z
dc.date.issued2018-09-13
dc.description.abstractThis paper aims to compare the feature extraction methods for accent dependent Thai speech from three regions including central, southern and northeastern regions. We investigate four frequency analysis methods: i.e., Energy Spectral Density (ESD), Power Spectral Density (PSD), Mel-Frequency Cepstral Coefficients (MFCC) and Spectrogram (SPT). Radial basis function kernel based on support vector machine is used as a classifier with 5-fold cross validation. The isolated speech data sets are recorded from 30 male and 30 female participants speaking the 10 Thai digits from 0 to 9. The MFCC-based feature gives better accuracy than ESD, PSD and SPT respectively. For within the same region, the MFCC-based feature provides average accuracy of 94.9% and 99.1% for male and female voices respectively. For the three regions, the MFCC-based feature provides average accuracy of 89.34% and 93.81% for male and female voices, respectively.
dc.identifier.citation2018 IEEE 7th International Conference on Communications and Electronics Icce 2018, 322-325, 2018
dc.identifier.doi10.1109/CCE.2018.8465705
dc.identifier.other2-s2.0-85057580394
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/9067
dc.source2018 IEEE 7th International Conference on Communications and Electronics Icce 2018
dc.subjectAccents classification
dc.subjectEnergy spectral density
dc.subjectFeature extraction
dc.subjectMel-frequency cepstral coefficients
dc.subjectPower spectral density
dc.subjectSpectrogram
dc.subjectSupport vector machine
dc.titleComparison of feature extraction for accent dependent Thai speech recognition system
dc.typeConference Paper

Files

Collections