KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
2 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Towards Sentic-Aware Multimodal Models for Cyberbullying Detection in Thai Memes(2026-01-01) ;Weradechtaweewon, Nattawat ;Boondamnoen, MongkolPasupa, KitsuchartCyberbullying has become an increasingly urgent issue in online communities. Memes, a popular form of online expression, often blend text and imagery in emotionally charged, sarcastic, or offensive ways–posing unique challenges for automatic harmful content detection. This work explores sentic-aware multimodal models for cyberbullying detection in Thai memes, with a focus on integrating affective commonsense knowledge through SenticNet-based features that emphasize conceptual reasoning and structured emotion representation. To enable this, we propose ThaiSenticNet 7, a resource adapted for the Thai language by translating from SenticNet 7, which supports the generation of sentic features. We investigate three representations–sentic vectors, sentic spectrograms, and sentic mel-spectrograms–and their integration with various sequential models to form sentic embeddings. These embeddings are fused with textual and visual information, extracted via a fine-tuned WangchanBERTa and a Swin Transformer, respectively, forming a unified multimodal pipeline. Experiments on a curated Thai meme dataset show that incorporating sentic features significantly enhances classification performance, with the best configuration–combining all three modalities–achieving an F<inf>1</inf>-score of 0.8044. Notably, the mel-spectrogram transformation proves particularly effective, suggesting that frequency-domain encoding helps capture subtle affective transitions in text-derived emotional signals. Our findings highlight the value of affective knowledge and multimodal modeling in tackling harmful content in memes. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Highlight Detection in Podcasts: A Multimodal Deep Learning Approach(2025-01-01) ;Phuengpanyaloet, Wongsapat ;Boonruengkhao, Nonpipat ;Anchutin, Viktor ;Pasupa, KitsuchartLoo, Chu KiongPodcasts have become a pervasive form of digital media, offering diverse content that often spans long hours. However, the vast volume of podcast episodes can make it challenging for listeners to locate the most engaging segments. Speech Emotion Recognition (SER) has witnessed remarkable advancements with the integration of deep learning techniques. This work proposes utilizing deep learning techniques employed in SER to discern emotional cues within podcasts, thereby enabling the detection of highlights. The task is framed as a binary classification problem, where the positive class contains examples of speech segments with high emotional activation. Transfer learning techniques from computer vision and speech recognition domains are applied, utilizing pre-trained models such as ConvNeXt, Vision Transformer, and wav2vec 2.0, which are compared with a baseline Convolutional Neural Network-Transformer hybrid. Additionally, multimodal models are introduced that learn from two distinct modalities: log mel-spectrograms and high-dimensional vector embeddings, both extracted from the raw audio data. The two modalities are combined using (i) a Simple Concatenated and (ii) CentralNet models. Experimental results demonstrate the effectiveness of combining two modalities over a single modality, achieving F<inf>1</inf>-scores of 0.6111 and 0.6270 for the Simple Concatenated and CentralNet models, respectively.
