KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 4 of 4
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Text analytics of IMDB reviews using latent Dirichlet allocation
    (2025-08-29)
    Wiriya, Surapong
    ;
    Thongbai, Tanajak
    ;
    Kamolsin, Chiranun
    ;
    Chomjan, Suchada
    ;
    Visutsak, Porawat
    This study utilizes a seven-step text analytics methodology in MATLAB to analyze IMDB review data. The methodology includes: 1) Data Loading and Preprocessing, involving loading the data and preprocessing text using functions such as converting text to lowercase, tokenization, punctuation removal, stop word removal, short and long word removal, and lemmatization. 2) Exploratory Data Analysis, using word clouds to visualize the most frequent words. 3) Bag-of-Words Model Creation, including splitting the data into training and validation sets and removing infrequent words and empty documents. 4) Topic Modeling with LDA, testing different solvers and evaluating performance using perplexity. 5) Optimal Topic Number Selection, optimizing the number of topics by comparing validation perplexities. 6) Final Topic Model Training, training a final LDA model and assessing its performance. and 7) Topic Interpretation and Analysis, involving visualizing word clouds and finding relevant reviews for specific words.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Visualizing Political Communication Trends across Generations on X (Twitter): Insights Through Topic Modeling and Word Clouds
    (2024-01-01)
    Udomwisanpat, Prinwat
    ;
    Jitkajornwanich, Kulsawasd
    ;
    Kraishan, Obada
    ;
    Srestasatheirn, Panu
    ;
    Lawawirojwong, Siam
    This study examines the interests and significance of words on Twitter (or X) across different generational groups: Baby Boomers, Generation X, Generation Y, and Generation Z. Using Topic Modeling with Latent Dirichlet Allocation (LDA), the research explores relationships and word importance within each group. As the results of topic modeling are not always easy to interpret, we used word cloud visualization to help make sense of the results for each generation. The findings reveal distinct patterns: Baby Boomers frequently mention print media, news websites, and prominent Thai political figures; Generation X emphasizes individuals and local political issues in Bangkok; Generation Y discusses political and social events; and Generation Z uniquely questions political and social activities. This research methodology is applicable across languages and tasks, offering insights into generational behaviors and interests.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    When are Latent Topics Useful for Text Mining?: Enriching Bag-of-Words Representations with Information Extraction in Thai News Articles
    (2023-01-01)
    Kanungsukkasem, Nont
    ;
    Chuangkrud, Piyawat
    ;
    Pitichotchokphokhin, Pimpitcha
    ;
    Damrongrat, Chaianun
    ;
    Leelanupab, Teerapong
    The Bag-of-Words (BOW) model is simple but one of the successful representations of text documents. This model, however, suffers from the sparse matrix, in which most of the elements are zero. Topic modeling is an unsupervised learning method that can represent text documents in a low-dimensional space. Latent Dirichlet Allocation (LDA) is a topic modeling technique used for topic extraction and data exploration, with interpretable output. This paper presents a thorough study of potential benefits of applying LDA, as a feature extraction, to topic discovery and document classification in Thai news articles, comparing with TF–IDF and Word2Vec. We also studied how much of the top Thai terms extracted from LDA with the different numbers of topics can be interpretable and meaningful, and can be a representative of the corpus. Besides, a set of Topic Coherence measures were included in our study to estimate the degree of semantic similarity of extracted topics. To compare the performance and optimization time of classification of features from the different feature extraction methods, various types of classifiers, e.g., Logistic Regression, Random Forest, XGBoosting, etc., were experimented.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Discover Underlying Topics in Thai News Articles: A Comparative Study of Probabilistic and Matrix Factorization Approaches
    (2020-06-01)
    Pitichotchokphokhin, Pimpitcha
    ;
    Chuangkrud, Piyawat
    ;
    Kalakan, Kongkan
    ;
    Suntisrivaraporn, Boontawee
    ;
    Leelanupab, Teerapong
    Topic modeling is an unsupervised learning approach, which can automatically discover the hidden thematic structure in text documents. For text mining, topic modeling is a language-independent technique that disregards grammar and word order. Apart from semantic and structural issues, Thai language is typically considered more complex than others. Due to the lack of word delimiter and a surfeit of composite words. Errors from word tokenization can create significant problems for any post processes of text, such as document retrieval, sentiment analysis, machine translation, etc., adversely decreasing the performance of text applications. Despite a strong correlation between word ordering and semantic meaning, topic modeling has been widely reported that it can extract latent information, aka. latent topic or latent semantic, encoded in documents. Although there were few previous research works on studying topic modeling in Thai language, they mostly focused on upstream processes of Natural Language Processing (NLP) in, for example, applying a refined stop-word list to, or adding N-gram on a single specific topic modeling method. To our knowledge, this paper is the first to explore different topic modeling approaches, i.e., Latent Dirichlet Allocation (LDA) and Nonnegative Metrix Factorization (NMF), in Thai Language to compare their coherence. We also employ and compare a set of state-of-the-art evaluation metrics based on Topic Coherence.