KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 10 of 27
  • Some of the metrics are blocked by your 
    Item type:Item,
    Improved Naive RAG by Integrated Advanced Techniques: A Comprehensive Framework Using Parent-Child Architecture, Hybrid Retrieval, and Contextual Compression
    (2026-07-01)
    Aromsuk, Tinnarat
    ;
    Netisopakul, Ponrudee
    ;
    Nootyaskool, Supakit
    This paper presents a novel approach to enhance Retrieval-Augmented Generation (RAG) systems through the integration of three advanced techniques: Parent-Child Architecture, Hybrid Retrieval, and Contextual Compression with cross-encoder re-ranking. We implement this framework using Langchain and FAISS vector search, with Anthropic’s Claude as the foundation model. Our comprehensive evaluation across 14 diverse Wikipedia-based knowledge domains employs the RAGAS framework to measure multiple performance dimensions. Results demonstrate that our advanced framework yields significant improvements in key metrics: Context Precision 10.11%, Context Recall 2.25%,and BLEU scores 1.14% compared to Naive RAG implementations. Domain analysis reveals particularly strong performance in Medicine 8.0% BLEU, Science 4.9%, and specialized Technology domains 4.9%. While, some technical domains such as Cybersecurity (−2.4%) and Biology (−6.6%) show performance degradation. Our framework achieves these improvements with minimal computational overhead by 1.89%,offering a practical approach to implementing domain-adaptive RAG systems that optimize context quality for improved generation performance.
  • Some of the metrics are blocked by your 
    Item type:Item,
    A Comparative Study of Deep Reinforcement Learning Agents for Gold Trading with Technical Indicators and LLM-Filtered News Sentiment
    (2026-06-16)
    Thanasarn, Thanapong
    ;
    Anuntachai, Anuntapat
    ;
    Netisopakul, Ponrudee
    Integrating macroeconomic news into deep reinforcement learning (DRL) for daily gold (XAU/USD) trading remains challenging. This study implements a natural language processing pipeline using majority voting from three large language models (Llama-3.1-8B-Instruct, Qwen3-8B, and Gemma-3-12B-it) to filter articles from The New York Times from 2010 to 2024, yielding 592 relevant articles that are subsequently assigned sentiment scores using FinBERT. We evaluate A2C, PPO, and SAC agents in a FinRL and Stable-Baselines3 framework using a four-fold walk-forward expanding-window protocol and Optuna hyperparameter tuning. The models are compared under two feature settings: (1) six technical indicators only and (2) technical indicators combined with sentiment features. Results show that incorporating LLM-filtered sentiment can modestly improve trading performance for some agents. PPO with sentiment achieves the best average cumulative return (10.05%) and Sharpe ratio (0.74), compared with its indicator-only version (9.80%, 0.72), and slightly outperforms the Buy-and-Hold (B&H) baseline (9.58%, 0.71). A2C also improves with sentiment (9.50% to 9.90%), while SAC shows no improvement in this setting. These findings suggest that LLM-filtered sentiment provides a modest benefit for some DRL agents in daily gold trading in our experiments.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Personalized Health Insurance Recommendation with Retrieval-Augmented Generation: A Study of Lexical, Dense, and Hybrid Retrieval
    (2026-06-16)
    Rujireksareekul, Phakon
    ;
    Anuntachai, Anuntapat
    ;
    Netisopakul, Ponrudee
    Health insurance information in Thailand is available on many company websites and is presented in different formats, making it difficult for users to compare coverage, conditions, and benefits. In this study, we propose a personalized health insurance recommendation system using Retrieval-Augmented Generation (RAG). The system retrieves information using three methods: BM25, dense retrieval, and hybrid retrieval with Reciprocal Rank Fusion. The language model analyzes the retrieved insurance plans and recommends suitable options based on the user query. Experimental results show that dense retrieval provides the best overall performance, while hybrid retrieval performs better than lexical search. The proposed RAG system also maintains practical response latency, making it suitable for interactive applications.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Machine Learning Models for Multi-Horizon Classification of Bitcoin Future Price Movements
    (2026-06-16)
    Boonpai, Sirawat
    ;
    Netisopakul, Ponrudee
    ;
    Anuntachai, Anuntapat
    This study presents a comprehensive framework for multi-horizon classification of Bitcoin futures price movements using machine learning. Five models of Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost were systematically evaluated across three timeframes (4h, 12h, 1d) and four classification schemes: binary (up/down) and multi-class (up/down/stable) with thresholds of 0.5%, 1.0%, and 1.5%, totaling 60 distinct experimental configurations. Rather than pursuing a strong predictive performance, this study prioritizes a rigorous comparative analysis across multiple dimensions to identify which combinations of model, timeframe, and classification scheme are most effective for Bitcoin futures. The results demonstrate that binary classification achieves the best predictive performance, with shorter timeframes yielding significantly better results, confirming the effectiveness of technical indicators in capturing the rapid price dynamics of Bitcoin futures. Notably, CatBoost achieved the highest F1-score for binary classification, while Random Forest proved the most robust model across diverse configurations. Feature importance analysis revealed that momentum-based indicators are the dominant predictors of price direction, while volatility features play a critical role in distinguishing sideways movements from directional ones in multi-class settings. Furthermore, the study demonstrates that narrower classification thresholds (0.5%) introduce noisier class boundaries and degrade performance, whereas wider thresholds (1.0%-1.5%) yield more stable results. These findings provide actionable guidelines for algorithmic trading in Bitcoin futures and establish a reproducible benchmark for future research.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Stock Price Prediction from Multi Data Sources Using LSTM, FinBERT and BERTweet
    (2026-06-16)
    Auensupa, Suchart
    ;
    Netisopakul, Ponrudee
    ;
    Chotipant, Supannada
    This research presents a stock price forecasting approach for technology sector companies Apple, Amazon and Tesla. The study begins by comparing a statistical model (ARIMA) with a deep learning model (LSTM) to identify the best-performing model, which is then used as the baseline for stock price forecasting. The approach integrates data from multiple sources, including numerical time series data and textual data. Numerical inputs consist of open, high, and low prices, which are used to forecast the closing price. Technical indicator features are subsequently added to enhance the model's predictive capability. Finally, textual data from economic news and Twitter social media reflecting market sentiment are incorporated. News sentiment is analyzed using the FinBERT model, while sentiment from social media data is evaluated using the BERTweet model. The resulting sentiment features are then combined with the numerical data and all inputs are processed using the LSTM model. Experimental results show that incorporating technical indicator features improves forecasting accuracy by an average of 17%. Furthermore, integrating textual data from news improves accuracy by an additional 6%, resulting in an overall performance improvement of up to 23%. These findings demonstrate the value of integrating multi-source data and highlight the important role of textual information in enhancing stock price forecasting performance.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Virtue-Based Thai Folktale Recommendation with Ensemble LLMs
    (2026-06-16)
    Daeng-Am, Wassana
    ;
    Anuntachai, Anuntapat
    ;
    Netisopakul, Ponrudee
    Moral learning plays an important role in early childhood education, yet teachers often spend considerable time identifying moral lessons in stories and deciding which virtues they represent. This study investigates whether a multi-model ensemble strategy can produce more teacher-aligned moral extractions from Thai folktales than a single-LLM baseline. We propose an LLM-based ensemble framework that extracts concise moral statements and classifies them into eight core virtues promoted by the Thai Ministry of Education: diligence, frugality, honesty, discipline, politeness, cleanliness, unity, and kindness. The framework combines outputs from Google Gemini 2.0 Flash, OpenAI GPT-4o-mini, and Anthropic Claude 3.5 Sonnet through a semantic consensus mechanism using BGE-M3 embeddings and majority voting. The system is evaluated on a corpus of 200 Thai folktales annotated by three experienced early childhood educators using multi-label metrics including Hamming Loss, Jaccard Similarity, and Exact Match. The results show that the ensemble approach achieves a Hamming Loss of 0.208, Jaccard Similarity of 0.564, and Exact Match of 0.175, consistently outperforming all single-model baselines. These findings suggest that consensus-driven ensemble inference provides a more robust and teacher-aligned foundation for automated moral education tools in Thai NLP.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Benchmarking Next-Generation Frontier Models for Document-Level Arabic-Thai Medical Translation: A Reliability Study of LLM-as-a-Judge
    (2026-01-01)
    Lertsuksakda, Rathawut
    ;
    Netisopakul, Ponrudee
    ;
    Boonsang, Siridech
    This research targets Arabic-to-Thai medical translation, a low-resource and linguistically distant pair critical for public health. We utilize the OpenWHO dataset, a resource comprising 293 documents. This dataset is shielded from webcrawling to minimize training data contamination and ensure rigorous zero-shot evaluation. We conduct a comparative analysis between 2026-era frontier models including GPT-5.2, Gemini 3 Pro, and Claude Opus 4.5 and industry-standard NMT represented by Google Translate. Using an LLM-as-a-Judge framework with an AI jury of efficient reasoning models such as GPT-5.1, Gemini 3 Flash, and Claude Sonnet 4.5, we performed 3,516 evaluations of fidelity, fluency, and cultural appropriateness. Results demonstrate that frontier models outperform traditional NMT across all dimensions. While high exact agreement suggests LLM-as-a-Judge frameworks are promising for scalable evaluation, human validation reveals persistent opportunities for improving AI-human alignment, necessitating targeted oversight for safety-critical applications.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Application of Large Language Models for Aspect-Based Sentiment Analysis on Social Media Data: A Case Study of the Thai Telecommunications Industry
    (2026-01-01)
    Limseesawan, Krittapas
    ;
    Netisopakul, Ponrudee
    ;
    Chotipant, Supannada
    ;
    Voravuthikunchai, Winn
    ;
    Sirivorachodphokin, Thitirat
    Social media is a valuable source of consumer opinion data, particularly in Thailand's highly competitive telecommunications market among AIS, TRUE, and DTAC. This study compares three LLM-based approaches for Aspect-Based Sentiment Analysis (ABSA) on 3,508 Thai-language messages from Pantip.com and YouTube.com: (1) Prompt Engineering, (2) Domain Adaptation & Fine-Tuning, and (3) Multi-Agent Debate Framework. Results show that Domain Adaptation & Fine-Tuning achieves the best accuracy-latency trade-off, with Qwen2.5-1.5B-Instruct exceeding 87% average accuracy in under 4 seconds per message. Multi-Agent Debate achieves the highest Category accuracy (82.82%) but at a latency cost of 10-16 seconds per message. Gemini-2.5-Flash-Lite provides the best speed-accuracy balance for Prompt Engineering without additional training. Critically, small open-source models (1B-1.5B parameters) subjected to domain-specific fine-tuning can approach or surpass proprietary large models on this task, suggesting domain alignment may outweigh raw parameter scale for narrow, well-defined ABSA tasks.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Baseline Performance of Pre-trained Models on Movie Genre Classification from Spectrograms
    (2025-04-01)
    Visutsak, Porawat
    ;
    Treeraphapkajondet, Kavin
    ;
    Sakphet, Visaroot
    ;
    Nitinuntatip, Wachirawit
    ;
    Satthong, Pawwinkan
    This study investigates the use of deep learning for classifying movie genres based on audio spectrograms. We construct a dataset of movie trailers, transform them into spectrograms, and label them by genre. Then, we utilize MATLAB's pre-trained convolutional neural networks (CNNs) for classification, comparing the performance of 9 different architectures, including MobileNet-v2, RestNet-18, DenseNet-201, Places365-GoogLeNet, VGG-16, VGG-19, Inception-RestNet-v2, Inception-v3, and NASANet-Mobile. We evaluated all models based on their ability to classify movie trailers into five genres: action, romance, drama, comedy, and thriller. Our results, based on accuracy and F1-score across genres, indicate that VGG16 achieves the highest overall performance with an accuracy of 86.27%, an F1-score of 86.69%, a recall of 86.87%, and a precision of 87.28%. This research demonstrates the potential of leveraging pre-trained CNNs, particularly VGG-16, for effcient and effective audio-based genre classification in movie trailers.
  • Some of the metrics are blocked by your 
    Item type:Item,
    COMPARING THE PERFORMANCE OF QUESTION ANSWERING BY LLMS USING QUANTIZATION AND RETRIEVAL AUGMENTED GENERATION TECHNIQUES
    (2025-03-01)
    Netisopakul, Ponrudee
    ;
    Haruethaipree, Sira
    The development of large language models (LLMs) like ChatGPT and Google Bard has led to the creation of intelligent chatbots and question-answering systems that are gaining widespread popularity. However, there are still limitations in using LLMs to develop applications, including the substantial computational resources required for fine-tuning and deployment. This paper studies and experiments with two techniques to reduce the computing resources required for developing a question-answering system using LLMs. A quantization technique is employed to compress the model’s size, and the application of Retrieval Augmented Generation (RAG) techniques is utilized for information retrieval. The study compares the performance of compressed-size models using Quantization and RAG against the original-sized models. The results show that quantizing the model can compress the VRAM resources used in the GPU between 38% to 57% while still achieving 68.9% accuracies compared to 70% in the non-compress model.