KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
69 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Item, A Persona-based Automated Evaluation Framework for Intelligent AI Teaching Assistants(2026-06-16) ;Limcharoen, Chananyu ;Sripinta, KeetaphatPasupa, KitsuchartTo address global teacher shortages and the need for personalized learning, we develop an intelligent teaching assistant leveraging multimodal large language models, agentic retrieval-augmented generation, and function calling. Unlike standard models, our system provides verifiable pedagogical guidance by reasoning across heterogeneous resources, including lecture slides and instructional videos. To overcome the scarcity of real-world datasets and high human evaluation costs, we propose a scalable, human-annotation-free framework utilizing simulated learner agents grounded in the Big Five personality theory. This allows for systematic assessment across 18 distinct student personas. We introduce a novel process-level metric, the dialogue recovery rate, and a dynamic adaptive policy to mitigate conversational deadlocks. Experimental results across 1,800 simulated dialogues using the Qwen3 family (8B, 14B, 32B) and its Thai-variant (Typhoon 2.5) reveal that Qwen3-14B attains the highest robustness and aggregate tutor-performance score within the proposed framework (72.6%). Analysis demonstrates significant correlations between learner traits and performance: Conscientiousness positively correlates with success (r = 0.49), while Extraversion is negatively associated with structural adherence (r = -0.54). This work establishes a reproducible benchmarking protocol for persona-aware, adaptive intelligent tutoring systems. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Self-attention hierarchical kernel reservoir state network for inland water level prediction(2026-01-15) ;Liu, Zongying ;Xu, Xiaohan ;Pasupa, Kitsuchart ;Loo, Chu KiongWei, YangWaterway transportation sustainably facilitates global trade through eco-efficient cargo movement, where accurate water level forecasting is critical for ensuring navigational safety and operational continuity. To develop a highly accurate prediction model, it is essential to consider the periodic characteristics of water level data, which often emerge in real-world datasets. This study introduces a novel reservoir state structure based on reservoir computing theory, the Self-attention Hierarchical Kernel Reservoir State Network (SHK-RSN). It employs three primary mechanisms. First, a hierarchical feature extraction method groups training data and extracts high-dimensional features from these groups using the kernel trick in a hierarchical manner. Second, a self-attention weight selection approach is introduced to replace the random weights in the Hierarchical Kernel Reservoir State Network (HK-RSN), improving the rationale for hidden neuron connections and enhancing the interpretability of weight selection. Third, a novel reservoir state structure is proposed to capture periodic information and extract temporal features across periods, enabling the model to capture richer temporal information and identify relationships among periods. Experiments are conducted on one artificial and five real-world time series datasets, with forecast performance evaluated over 1–7 steps. Our proposed model, SHK-RSN, is compared with models based on randomization, the kernel trick, and deep learning. The experimental results demonstrate that SHK-RSN exhibits superior forecasting ability relative to the baselines. It achieves the best Symmetric Mean Absolute Percentage Error (SMAPE) across all datasets in the 1–7 period average among baseline methods, demonstrating a relative improvement of 25.7% to 46.9% over the conventional Echo State Network. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Towards Sentic-Aware Multimodal Models for Cyberbullying Detection in Thai Memes(2026-01-01) ;Weradechtaweewon, Nattawat ;Boondamnoen, MongkolPasupa, KitsuchartCyberbullying has become an increasingly urgent issue in online communities. Memes, a popular form of online expression, often blend text and imagery in emotionally charged, sarcastic, or offensive ways–posing unique challenges for automatic harmful content detection. This work explores sentic-aware multimodal models for cyberbullying detection in Thai memes, with a focus on integrating affective commonsense knowledge through SenticNet-based features that emphasize conceptual reasoning and structured emotion representation. To enable this, we propose ThaiSenticNet 7, a resource adapted for the Thai language by translating from SenticNet 7, which supports the generation of sentic features. We investigate three representations–sentic vectors, sentic spectrograms, and sentic mel-spectrograms–and their integration with various sequential models to form sentic embeddings. These embeddings are fused with textual and visual information, extracted via a fine-tuned WangchanBERTa and a Swin Transformer, respectively, forming a unified multimodal pipeline. Experiments on a curated Thai meme dataset show that incorporating sentic features significantly enhances classification performance, with the best configuration–combining all three modalities–achieving an F<inf>1</inf>-score of 0.8044. Notably, the mel-spectrogram transformation proves particularly effective, suggesting that frequency-domain encoding helps capture subtle affective transitions in text-derived emotional signals. Our findings highlight the value of affective knowledge and multimodal modeling in tackling harmful content in memes. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Unified Multimodal-Multitask Learning for Vehicle Damage Assessment in Insurance Applications(2026-01-01) ;Phuengpanyaloet, Wongsapat ;Pasupa, Kitsuchart ;Angsarawanee, Thanatwit ;Chetprayoon, PanumateSakdejayont, TheeratAutomated vehicle damage assessment requires both precise localization and clear textual reporting. While existing methods typically treat these as separate tasks, the trade-offs of unified multimodal-multitask learning in this domain remain underexplored. This paper conducts a comparative study between a unified vision-language framework, Generative Region-to-Text Transformer (GRiT), and single-task baselines derived from GRiT by isolating the detection and captioning components. We adapt GRiT to the insurance domain using a dataset enriched with vehicle part annotations and structured damage descriptions. Experimental results demonstrate that the unified model achieves competitive detection performance (F<inf>1</inf>-score: 0.54), slightly outperforming the detection baseline model. Crucially, it significantly surpasses the caption baseline model in description quality (METEOR: 0.75, ROUGE: 0.70, BLEU: 0.46), confirming that object-level visual grounding is essential for accurate reporting. These findings indicate that unified multimodal learning enhances semantic interpretation without compromising localization accuracy, offering a promising direction for automated insurance workflows. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters(2025-10-21) ;Das, Sarmistha ;Mathur, Priya ;Sharma, Ishani ;Saha, SriparnaPasupa, KitsuchartThe exponential technological breakthrough of the FinTech industry has significantly enhanced user engagement through sophisticated advisory chatbots. However, large-scale fine-tuning of LLMs can occasionally yield unprofessional or flippant remarks, such as 'With that money, you're going to change the world,' which, though factually correct, can be contextually inappropriate and erode user trust. The scarcity of domain-specific datasets has led previous studies to focus on isolated components, such as reasoning-aware frameworks or the enhancement of human-like response generation. To address this research gap, we present Fin-Solution 2.O, an advanced solution that 1) introduces the multi-turn financial conversational dataset, Fin-Vault, and 2) incorporates a unified model, Fin-Ally, which integrates commonsense reasoning, politeness, and human-like conversational dynamics. Fin-Ally is powered by COMET-BART-embedded commonsense context and optimized with a Direct Preference Optimization (DPO) mechanism to generate human-aligned responses. The novel Fin-Vault dataset, consisting of 1,417 annotated multi-turn dialogues, enables Fin-Ally to extend beyond basic account management to provide personalized budgeting, real-time expense tracking, and automated financial planning. Our comprehensive results demonstrate that incorporating commonsense context enables language models to generate more refined, textually precise, and professionally grounded financial guidance, positioning this approach as a next-generation AI solution for the FinTech sector. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Optimizing colorectal polyp detection and localization: Impact of RGB color adjustment on CNN performance(2025-06-01) ;Jamrasnarodom, Jirakorn ;Rajborirug, Pharuj ;Pisespongsa, PisesPasupa, KitsuchartColorectal cancer, arising from adenomatous polyps, is a leading cause of cancer-related mortality, making early detection and removal crucial for preventing cancer progression. Machine learning is increasingly used to enhance polyp detection during colonoscopy, the gold standard for colorectal cancer screening, despite its operator-dependent miss rates. This study explores the impact of RGB color adjustment on Convolutional Neural Network (CNN) models for improving polyp detection and localization in colonoscopic images. Using datasets from Harvard Dataverse for training and internal validation, and LDPolypVideo-Benchmark for external validation, RGB color adjustments were applied, and YOLOv8s was used to develop models. Bayesian optimization identified the best RGB adjustments, with performance assessed using mean average precision (mAP) and F<inf>1</inf>-scores. Results showed that RGB adjustment with 1.0 R-1.0 G-0.8 B improved polyp detection, achieving an mAP of 0.777 and an F<inf>1</inf>-score of 0.720 on internal test sets, and localization performance with an F<inf>1</inf>-score of 0.883 on adjusted images. External validation showed improvement but with a lower F<inf>1</inf>-score of 0.556. While RGB adjustments improved performance in our study, their generalizability to diverse datasets and clinical settings has yet to be validated. Thus, although RGB color adjustment enhances CNN model performance for detecting and localizing colorectal polyps, further research is needed to verify these improvements across diverse datasets and clinical settings. • RGB Color Adjustment: Applied RGB color adjustments to colonoscopic images to enhance the performance of Convolutional Neural Network (CNN) models. • Model Development: Used YOLOv8s for polyp detection and localization, with Bayesian optimization to identify the best RGB adjustments. • Performance Evaluation: Assessed model performance using mAP and F<inf>1</inf>-scores on both internal and external validation datasets. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Detecting Cyberbullying in Thai Memes: A Multimodal Approach Using Deep Learning(2025-01-01) ;Phosit, Salisa ;Kongsamlit, SawarodPasupa, KitsuchartToday, people tend to communicate more through the Internet due to its speed and convenience. However, there are hidden drawbacks, such as the misuse of online social media and the occurrence of cyberbullying, which can lead to negative feelings or even mental health issues. Therefore, it is necessary to develop cyberbullying detection methods. While these methods have been studied in various languages, more research is still needed in Thai due to its status as a low-resource language. This research aims to develop a model for detecting cyberbullying in online social media, specifically focusing on memes. The primary objective is to develop a Thai meme dataset collected from Facebook using deep learning techniques for model development. The developed model utilizes multimodal concepts, incorporating both text and images. In addition to detecting cyberbullying, it also considers other aspects, such as harmfulness, offensiveness, sarcasm, sentiment, emotions, and topics. Our findings indicate that multimodal models outperform unimodal models, with the combination of WangchanBERTa and ViT-B/32 achieving the highest performance across most tasks. Compared to the best-performing unimodal models, the multimodal approach resulted in a 1.53% improvement in the F<inf>1</inf> -Score on the Bully task. This underscores the significant value of integrating textual and visual information for a more robust and nuanced understanding of cyberbullying content. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Highlight Detection in Podcasts: A Multimodal Deep Learning Approach(2025-01-01) ;Phuengpanyaloet, Wongsapat ;Boonruengkhao, Nonpipat ;Anchutin, Viktor ;Pasupa, KitsuchartLoo, Chu KiongPodcasts have become a pervasive form of digital media, offering diverse content that often spans long hours. However, the vast volume of podcast episodes can make it challenging for listeners to locate the most engaging segments. Speech Emotion Recognition (SER) has witnessed remarkable advancements with the integration of deep learning techniques. This work proposes utilizing deep learning techniques employed in SER to discern emotional cues within podcasts, thereby enabling the detection of highlights. The task is framed as a binary classification problem, where the positive class contains examples of speech segments with high emotional activation. Transfer learning techniques from computer vision and speech recognition domains are applied, utilizing pre-trained models such as ConvNeXt, Vision Transformer, and wav2vec 2.0, which are compared with a baseline Convolutional Neural Network-Transformer hybrid. Additionally, multimodal models are introduced that learn from two distinct modalities: log mel-spectrograms and high-dimensional vector embeddings, both extracted from the raw audio data. The two modalities are combined using (i) a Simple Concatenated and (ii) CentralNet models. Experimental results demonstrate the effectiveness of combining two modalities over a single modality, achieving F<inf>1</inf>-scores of 0.6111 and 0.6270 for the Simple Concatenated and CentralNet models, respectively. - Some of the metrics are blocked by yourconsent settings
Item type:Item, MS-PatchTST: Leveraging Multi-Scale Temporal Features for Water Level Forecasting(2025-01-01) ;Zhang, Dong ;Pasupa, Kitsuchart ;Liu, ZongyingPan, MingyangAccurate water level forecasting is essential for navigation, enabling safe sailing, effective drought management, optimized route planning, and efficient port operations. However, traditional statistical approaches and conventional machine learning models often struggle to capture adaptive, multi-scale temporal features, thereby limiting forecasting accuracy. In recent years, patch-based forecasting methods have demonstrated strong capabilities in modeling consecutive temporal features. Building on this foundation, we propose Multi-Scale PatchTST (MS-PatchTST), a framework designed to enhance the perception of multi-scale information. The model incorporates a newly developed multi-scale parallel convolutional network (Multi-Scale ConvNet) to extract interaction features across different time scales. These features are then fused through a Transformer Encoder with relative positional encoding to capture temporal dependencies more effectively. Finally, the kernel mean squared error loss function is employed in place of the conventional mean squared error loss, improving the optimization process and enhancing overall training performance. Experiments on four real-world water level datasets demonstrate that MS-PatchTST consistently outperforms state-of-the-art baselines, achieving an average reduction of approximately 13% in both MAE and SMAPE compared with PatchTST. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports(2025-01-01) ;Singh, Punit Kumar ;Kumar, Nishant ;Ghosh, Akash ;Pasad, KunalSoni, KhushiLanguage Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce CultSportQA, a benchmark designed to assess LMs' understanding of traditional sports across 60 countries and 6 continents, encompassing four distinct cultural categories. The dataset features 33,000 multiple-choice questions (MCQs) across text and image modalities, each of which is categorized into three key types: history-based, rule-based, and scenario-based. To evaluate model performance, we employ zero-shot, few-shot, and chain-of-thought (CoT) prompting across a diverse set of Large Language Models (LLMs), Small Language Models (SLMs), and Multimodal Large Language Models (MLMs). By providing a comprehensive multilingual and multicultural sports benchmark, CultSportQA establishes a new standard for assessing AI's ability to understand and reason about traditional sports.
