Classification of Depression Audio Data by Deep Learning

Phanomkorn Homsiang; Treesukon Treebupachatsakul; Komsan Kiatrungrit; Suvit Poomrittigul

doi:10.1109/bmeicon56653.2022.10012102

Classification of Depression Audio Data by Deep Learning

dc.contributor.author	Phanomkorn Homsiang
dc.contributor.author	Treesukon Treebupachatsakul
dc.contributor.author	Komsan Kiatrungrit
dc.contributor.author	Suvit Poomrittigul
dc.date.accessioned	2026-05-08T19:19:55Z
dc.date.issued	2022-11-10
dc.description.abstract	Due to many factors such as anxiety from contracting the disease and concern about the socioeconomic impacts, Thai people have accumulated stress and are at risk of depression. The diagnosis of depression can be primarily assessed by testing the assessments such as PHQ8, PHQ-9, and CES-D. The applied deep learning technology in medicine has received research interest and has been developing. In this research, we tried the classification of depression and non-depression audio datasets with the implementation of 4 model architectures: 1D CNN, 2D CNN, LSTM, and GRU. By converting wave audio format (WAV) of Daic-woz database to the Melfrequency cepstrum (MFC). We have done the training and evaluated the 4 model architectures and compared the results between non-augmented and augmented datasets. The highest accuracy was obtained from 1D CNN with a non-data augmentation of 95%, and a 2D CNN with a data augmentation of 75%. These results confirm that human voices can differentiate between depression and non-depression.
dc.identifier.doi	10.1109/bmeicon56653.2022.10012102
dc.identifier.uri	https://dspace.kmitl.ac.th/handle/123456789/17272
dc.subject	Phonocardiography and Auscultation Techniques
dc.subject	Emotion and Mood Recognition
dc.subject	ECG Monitoring and Analysis
dc.title	Classification of Depression Audio Data by Deep Learning
dc.type	Article

Collections

All

Classification of Depression Audio Data by Deep Learning

Files

Collections