Repository logo
Communities & Collections
Research Outputs
Fundings & Projects
People
Statistics
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. KMITL
  3. Publication
  4. DataDecon: Data Cleansing Tools for Large Language Model with Efficient Decontamination Techniques
Loading...
Thumbnail Image

DataDecon: Data Cleansing Tools for Large Language Model with Efficient Decontamination Techniques

Author(s)
Yuenyong, Sumeth
Buppodom, Norapat
Sangkaew, Koravich
Boonmeeprakob, Konthee
Boonkwan, Prachya
Jaroenkantasima, Jillaphat
Khlaisamniang, Pitikorn
Lertpiya, Anuruth
Piyatumrong, Apivadee
Rojratchadakorn, Peerawat
Rugsujarit, Thaweewat
Saengsukhiran, Teerapol
Saetan, Kriangkrai
Sukprapa, Isada
Thavornmongkol, Thanachot
Thongthungwong, Nucharee
Triamamornwooth, Patteera
Utupon, Chanon
Viriyayudhakorn, Kobkrit
Witchutanon, Phoochit
Wongprayon, Sadit
Supnithi, Thepchai
Date Issued
January 1, 2024
Type
Conference Paper
DOI
10.1109/iSAI-NLP64410.2024.10799278
Abstract
Large language models (LLMs) play an important role in modern NLP technology as they are versatile for a wide array of NLP tasks. However, constructing an LLM is challenging due to concealed construction pipelines, the lack of cleansed datasets, and hyperparameter settings, making it almost irreproducible. This paper presents an efficient pipeline for constructing an LLM tailored to a low-to-medium-sourced language with a high level of data contamination and tools to cleanse the dataset. Following our pipeline, we constructed OpenThaiGPT, an LLM for Thai, with only open-sourced datasets such as CC100, OSCAR, and mC4, and achieved the state-of-the-art accuracies on our downstream tasks. Here, we disclosed the data statistics and all hyperparameter settings for reproducibility.
Citation
19th International Joint Symposium on Artificial Intelligence and Natural Language Processing Isai Nlp 2024, 2024
Subjects

corpus construction

data cleansing

large language models...

text deduplication

Thai language

Metrics
Get Involved!
  • Source Code
  • Documentation
  • Slack Channel
Make it your own

DSpace-CRIS can be extensively configured to meet your needs. Decide which information need to be collected and available with fine-grained security. Start updating the theme to match your Institution's web identity.

Need professional help?

The original creators of DSpace-CRIS at 4Science can take your project to the next level, get in touch!

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Accessibility settings
  • Privacy policy
  • End User Agreement
  • Send Feedback