KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Pantip Multi-turn Datasets Generating from Thai Large Social Platform Forum Using Sentence Similarity Techniques(2024-01-01) ;Sae-Oueng, Anon ;Kerdthaisong, Kun ;Sukhantharat, Kittisak ;Phasook, PakawatChuangkrud, PiyawatFine-tuning Large Language Models (LLMs) for specific domains is crucial. However, lack of Thai open dialogues presents a major challenge. For the major challenge, this study proposes a novel methodology for extracting and constructing multi-turn conversational data from existing Thai large social platform, named Pantip. Our approach implements semantic matching algorithms to identify and compile both single-turn and multi-turn dialogues. By employing a cosine similarity threshold ≥ 0.3, we yields contextually coherent conversation pairs directly from the source data. The outcome dataset represents real Thai conversation styles, which could improve how accurately fine-tuned language models reflect Thai social and linguistic norms. Our approach introduces a new way to use existing publicly available data to create training datasets, which is especially valuable for languages with limited resources. Chaotic Pantip datasets can be contributed to the development of more culturally attuned and linguistically precise Thai language models, potentially advancing culturally-specific natural language processing. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, ThaiBKD: Effective of Continual Pre-Training LLM in Thai Language Based on Knowledge Dataset(2024-01-01) ;Phasook, Pakawat ;Pranee, Jessada ;Limcharoen, Chananyu ;Sukhantharat, KittisakSaeoueng, AnonLarge Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing (NLP) tasks. Existing works emphasize the importance of continual pretraining with high-quality knowledge data to enhance LLM performance. This paper investigates the comparative effectiveness of using meticulously curated, high-quality knowledge datasets - comprising scholarly articles, journalism, medical expertise, financial data, legal documents, and verified library resources - against general data sourced from social media and diverse domain sources, particularly from the popular Thai blog platform Pantip, as pre-training data for LLMs.The evaluation results show that ThaiBKD-based LLM, trained on clean and domain-specific data, consistently outperforms Pantip-based LLM in most aspects of Thai language evaluation tasks, demonstrating more effective reasoning capabilities and achieving higher scores across various benchmarks. Notably, the Thai Language Based on Knowledge Dataset (ThaiBKD) LLM outperforms or higher than the performance of GPT-3.5 Turbo in a versatile evaluation set, particularly in tasks requiring specialized knowledge.However, the Pantip-based LLM exhibits substantial strengths in culturally nuanced tasks, such as XCOPA, XNLI, and Belebele, where its understanding of informal and diverse language structures from social media provides a competitive edge. These findings highlight the nuanced trade-offs between data quality, domain specificity, and cultural relevance, underscoring the need for strategic data selection in the pre-training of advanced Large language Models and Continual-pretraining Large Language Models.
