Repository logo
Communities & Collections
Research Outputs
Fundings & Projects
People
Statistics
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. KMITL
  3. Publication
  4. Deep Residual Local Feature Learning for Speech Emotion Recognition
Loading...
Thumbnail Image

Deep Residual Local Feature Learning for Speech Emotion Recognition

Author(s)
Singkul, Sattaya
Chatchaisathaporn, Thakorn
Suntisrivaraporn, Boontawee
Woraratpanya, Kuntpong
Date Issued
January 1, 2020
Type
Conference Paper
DOI
10.1007/978-3-030-63830-6_21
Abstract
Speech Emotion Recognition (SER) is becoming a key role in global business today to improve service efficiency, like call center services. Recent SERs were based on a deep learning approach. However, the efficiency of deep learning depends on the number of layers, i.e., the deeper layers, the higher efficiency. On the other hand, the deeper layers are causes of a vanishing gradient problem, a low learning rate, and high time-consuming. Therefore, this paper proposed a redesign of existing local feature learning block (LFLB). The new design is called a deep residual local feature learning block (DeepResLFLB). DeepResLFLB consists of three cascade blocks: LFLB, residual local feature learning block (ResLFLB), and multilayer perceptron (MLP). LFLB is built for learning local correlations along with extracting hierarchical correlations; DeepResLFLB can take advantage of repeatedly learning to explain more detail in deeper layers using residual learning for solving vanishing gradient and reducing overfitting; and MLP is adopted to find the relationship of learning and discover probability for predicted speech emotions and gender types. Based on two available published datasets: EMODB and RAVDESS, the proposed DeepResLFLB can significantly improve performance when evaluated by standard metrics: accuracy, precision, recall, and F1-score.
Citation
Lecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics, 12532 LNCS, 241-252, 2020
Subjects

Chromagram

CNN network

Log-Mel spectrogram

Residual feature lear...

Speech Emotion Recogn...

Metrics
Get Involved!
  • Source Code
  • Documentation
  • Slack Channel
Make it your own

DSpace-CRIS can be extensively configured to meet your needs. Decide which information need to be collected and available with fine-grained security. Start updating the theme to match your Institution's web identity.

Need professional help?

The original creators of DSpace-CRIS at 4Science can take your project to the next level, get in touch!

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback