KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 3 of 3
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Thai Generalized Alignment Task (TGAT): A Corpus and Comparative Study for Hallucination Detection in Thai
    (2025-01-01)
    Wangwon, Sivakorn
    ;
    Pitijaroonpong, Rapepong
    ;
    Chuangkrud, Piyawat
    ;
    Damrongrat, Chaianun
    ;
    Kongyoung, Sarawoot
    Large Language Models (LLMs) face critical reliability challenges due to hallucination-the generation of factually inaccurate content. While hallucination detection has advanced for major languages, Thai remains underserved, lacking specialized alignment models and datasets. We introduce the Thai Generalized Alignment Task (TGAT), a seven-subtask corpus tailored for Thai hallucination detection, spanning Fact Verification, Natural Language Inference, Information Retrieval, Question Answering, Summarization, Semantic Textual Similarity, and Paraphrase Identification. We conduct a comparative study across transformer families-encoder-only, encoder-decoder, and decoder-only-evaluated on the test set of each sub-task to analyze architectural compatibility and generalization for Thai hallucination detection. We also perform an ablation study to quantify how the presence of each subtask dataset affects the overall average performance, clarifying the role of dataset composition. Our experiments show that decoder-only models consistently outperform encoder-decoder and encoder-only alternatives under a standardized zero-shot prompting setup, establishing strong baselines and offering guidance for model selection and dataset design in low-resource settings.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Pantip Multi-turn Datasets Generating from Thai Large Social Platform Forum Using Sentence Similarity Techniques
    (2024-01-01)
    Sae-Oueng, Anon
    ;
    Kerdthaisong, Kun
    ;
    Sukhantharat, Kittisak
    ;
    Phasook, Pakawat
    ;
    Chuangkrud, Piyawat
    Fine-tuning Large Language Models (LLMs) for specific domains is crucial. However, lack of Thai open dialogues presents a major challenge. For the major challenge, this study proposes a novel methodology for extracting and constructing multi-turn conversational data from existing Thai large social platform, named Pantip. Our approach implements semantic matching algorithms to identify and compile both single-turn and multi-turn dialogues. By employing a cosine similarity threshold ≥ 0.3, we yields contextually coherent conversation pairs directly from the source data. The outcome dataset represents real Thai conversation styles, which could improve how accurately fine-tuned language models reflect Thai social and linguistic norms. Our approach introduces a new way to use existing publicly available data to create training datasets, which is especially valuable for languages with limited resources. Chaotic Pantip datasets can be contributed to the development of more culturally attuned and linguistically precise Thai language models, potentially advancing culturally-specific natural language processing.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    ThaiBKD: Effective of Continual Pre-Training LLM in Thai Language Based on Knowledge Dataset
    (2024-01-01)
    Phasook, Pakawat
    ;
    Pranee, Jessada
    ;
    Limcharoen, Chananyu
    ;
    Sukhantharat, Kittisak
    ;
    Saeoueng, Anon
    Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing (NLP) tasks. Existing works emphasize the importance of continual pretraining with high-quality knowledge data to enhance LLM performance. This paper investigates the comparative effectiveness of using meticulously curated, high-quality knowledge datasets - comprising scholarly articles, journalism, medical expertise, financial data, legal documents, and verified library resources - against general data sourced from social media and diverse domain sources, particularly from the popular Thai blog platform Pantip, as pre-training data for LLMs.The evaluation results show that ThaiBKD-based LLM, trained on clean and domain-specific data, consistently outperforms Pantip-based LLM in most aspects of Thai language evaluation tasks, demonstrating more effective reasoning capabilities and achieving higher scores across various benchmarks. Notably, the Thai Language Based on Knowledge Dataset (ThaiBKD) LLM outperforms or higher than the performance of GPT-3.5 Turbo in a versatile evaluation set, particularly in tasks requiring specialized knowledge.However, the Pantip-based LLM exhibits substantial strengths in culturally nuanced tasks, such as XCOPA, XNLI, and Belebele, where its understanding of informal and diverse language structures from social media provides a competitive edge. These findings highlight the nuanced trade-offs between data quality, domain specificity, and cultural relevance, underscoring the need for strategic data selection in the pre-training of advanced Large language Models and Continual-pretraining Large Language Models.