Loading...
Comparison of Data Augmentation Techniques for Thai Text Sentiment Analysis
Author(s)
Rongsawad, Kanda
Chatwiriya, Watchara
Date Issued
January 1, 2023
Type
Conference Paper
Abstract
Many text corpora e.g., Thai text, have problems with low amounts of data or imbalanced data. Data augmentation is one technique used to solve this problem by generating synthetic data. Various data augmentation techniques are applied to text, each method has different algorithms to generate new text. This paper will compare the performance of three well-known data augmentation techniques: Easy Data Augmentation, Embedding Replacement, and Replacement by the masked language model. We apply these data augmentation techniques to the Thai text sentiment corpus to generate synthetic text, then the synthetic text will be used to train the sentiment classifier. After that, the performances of classifiers are analyzed and compared. Three sentiment classifiers used in this paper are Convolutional Neural Network (CNN), Bidirectional Long-Short Term Memory (Bi-LSTM), and Gated Recurrent Unit (GRU). From the results, we found that data augmentation techniques can improve sentiment classifiers’ performance. The EDA technique outperformed the others in the Cosmetics Drink and All topics sets but in Restaurant sets, MLM is the best technique that can improve Thai sentiment classification performance. The results show that EDA can be used as a baseline for further study of data augmentation.
Citation
Lecture Notes in Networks and Systems, 679 LNNS, 131-139, 2023
