KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 2 of 2
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Highlight Detection in Podcasts: A Multimodal Deep Learning Approach
    (2025-01-01)
    Phuengpanyaloet, Wongsapat
    ;
    Boonruengkhao, Nonpipat
    ;
    Anchutin, Viktor
    ;
    Pasupa, Kitsuchart
    ;
    Loo, Chu Kiong
    Podcasts have become a pervasive form of digital media, offering diverse content that often spans long hours. However, the vast volume of podcast episodes can make it challenging for listeners to locate the most engaging segments. Speech Emotion Recognition (SER) has witnessed remarkable advancements with the integration of deep learning techniques. This work proposes utilizing deep learning techniques employed in SER to discern emotional cues within podcasts, thereby enabling the detection of highlights. The task is framed as a binary classification problem, where the positive class contains examples of speech segments with high emotional activation. Transfer learning techniques from computer vision and speech recognition domains are applied, utilizing pre-trained models such as ConvNeXt, Vision Transformer, and wav2vec 2.0, which are compared with a baseline Convolutional Neural Network-Transformer hybrid. Additionally, multimodal models are introduced that learn from two distinct modalities: log mel-spectrograms and high-dimensional vector embeddings, both extracted from the raw audio data. The two modalities are combined using (i) a Simple Concatenated and (ii) CentralNet models. Experimental results demonstrate the effectiveness of combining two modalities over a single modality, achieving F<inf>1</inf>-scores of 0.6111 and 0.6270 for the Simple Concatenated and CentralNet models, respectively.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Enhancing Thai Food Recognition Through Multimodal Fusion of Image and Fourier Spectrum
    (2024-01-01)
    Pasupa, Kitsuchart
    ;
    Woraratpanya, Kuntpong
    Recognizing food is a challenging task in artificial intelligence research because food items may undergo deformations during cooking or serving, and they can be partially or fully occluded, making it difficult for recognition systems to analyze their complete visual information. Therefore, in addition to evaluating object detection effectiveness, consideration must be given to the texture of the food. However, Convolutional Neural Networks may fall short in capturing textural information. In this research, we propose a method to enhance the efficiency of Thai food recognition by employing the concept of multi-modal fusion, incorporating Fourier Spectrum images to take texture representation into account and improve the model’s performance. In the fusion process, we employed the CentralNet framework and compared it with baselines (using only images and conventional concatenation fusion) on three datasets: THFOOD-50, FoodyDudy, and our newly proposed food dataset called iTHFOOD-200. This new dataset encompasses a more diverse range of food types. The experimental results demonstrate that fusion with the CentralNet framework yielded better performance than the baselines using only images (7.1% Top-1 Accuracy and 3.9% Top-5 Accuracy on average) and conventional concatenation fusion (5.0% Top-1 Accuracy and 2.4% Top-5 Accuracy on average).