KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
39 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Item, Predicting the Risk of Ultimate Low Ratings in Online Courses Using Machine Learning: Analyzing Engagement and Complexity for AI-Driven Early Instructional Communication Interventions(2026-01-01) ;Kraishan, Obada ;Jitkajornwanich, Kulsawasd ;Vijaranakul, NattadetKee, KerkOnline asynchronous courses have become a common way of learning, especially after the COVID-19 epidemic. Despite their convenience and flexibility, these courses often suffer from low engagement and/or poor course quality, ultimately reflected by a low rating. A low rating is often noted at the end of a course, by when it is often too late to intervene or improve the course. The goal of this study is twofold. First, we aim to identify key factors influencing the ultimate ratings for online asynchronous courses. Second, we then build an AI model to predict the risk of ultimate low ratings before a course is fully completed. Inspired by Moore’s model of interaction, which includes two main factors: learner-content and learner-instructor interactions, we derived two features from the dataset: course complexity and course engagement. We used a dataset of 191,849 Udemy courses collected from May 2024 to August 2024. To predict the ultimate low ratings in advance so mid-course interventions are possible, we applied machine learning models, including Logistic Regression, Random Forest, and XGBoost. In our experiments, the XGBoost model achieved the highest performance with 71% accuracy, with course complexity and the number of reviews left by learners being the two most influential factors in predicting a low course rating. Constantly monitoring course complexity and the number of course reviews can help with early detection. The findings can provide valuable, actionable insights and guidance for online education platforms, especially for instructors to adjust their courses before the courses are complete. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Early Diagnosis of Knee Osteoarthritis With a Natural Language Processing–Driven Approach Based on Clinician Notes: Development and Validation Study(2025-01-01) ;Thanyakunsajja, Narathip ;Jitkajornwanich, Kulsawasd ;Xu, Shan ;Shin, DongheeCharoenporn, PattamaBackground: Knee osteoarthritis (OA) is a common form of knee arthritis that can cause significant disability and affect a patient’s quality of life. Although this disease is chronic and irreversible, the patient’s condition can be improved and the progression of the disease can be prevented if the disease is diagnosed early and the patient receives appropriate treatment immediately. Therefore, the prediction of knee OA is considered one of the essential steps to effectively diagnose and prevent further severe OA conditions. Knee OA is commonly diagnosed by medical experts or physicians, and the diagnosis of OA is mainly based on patients’ laboratory results and medical images, including x-ray and magnetic resonance images. However, diagnosis through such data is often time-consuming. Moreover, the diagnosis results can vary among physicians depending on their expertise. Previous studies mostly focused on using approaches, such as those involving artificial intelligence, to automatically detect knee OA through such data. However, these studies did not incorporate clinicians’ or doctors’ notes (text data) into the analysis, although these data involving reported symptoms and behaviors are already available and easier to collect and access than laboratory data and image data. Objective: We propose a novel natural language processing–driven approach based on clinicians’ or doctors’ notes of patient-reported symptoms (text data only) for diagnosing knee OA. Methods: The textual information from clinicians’ or doctors’ notes was first preprocessed using text analysis algorithms with respect to natural language processing. We then incorporated deep learning models, including convolutional neural networks, bidirectional long short-term memory (BiLSTM), and gated recurrent units. Lastly, a disease-specific standard questionnaire called WOMAC (Western Ontario and McMaster Universities Arthritis Index) was taken into account to improve the overall performance of the models. Results: Our experiment included 5849 records (OA: 3455; non-OA: 2394). Before applying our WOMAC-based processing approach, the best-performing model was BiLSTM (area under the curve, 0.85; accuracy, 0.87; precision, 0.85; sensitivity, 0.95; specificity, 0.76; F<inf>1</inf>-score, 0.90), and there was an improvement in the results with BiLSTM after applying our approach (area under the curve, 0.91; accuracy, 0.91; precision, 0.91; sensitivity, 0.94; specificity, 0.87; F<inf>1</inf>-score, 0.93). Conclusions: Our proposed method for predicting the occurrence of knee OA showed better performance than other conventional methods that use image data and statistical laboratory data. The findings indicate the feasibility of using text data (symptom descriptions reported by patients and recorded by doctors) to predict knee OA. Medical notes of symptom reports can be considered a valuable data source for predicting whether a particular knee is likely to experience OA progression. - Some of the metrics are blocked by yourconsent settings
Item type:Item, A longitudinal examination of collaboration diversity among communication scholars: 1990–2023(2024-12-01) ;Xu, Shan ;Jitkajornwanich, Kulsawasd ;David, Prabu ;Park, Hye JungZhao, YaniThis study examines racial diversity in co-authorship in articles published in communication journals and its association with citations accrued over time. We analyzed 76,217 publications from 73 communication journals, spanning from 1990 to 2023, with a focus on racial diversity in authorship as an indicator of collaboration diversity. Our results reveal that diversity is positively associated with the number of citations received, with this positive effect increasing over time. In addition, non-White lead authors collaborated more diversely, whereas White authors exhibited a faster increase in collaboration diversity over the years. Furthermore, the positive association between collaboration diversity and citations was more pronounced when the lead author was non-White than when White. Additional analyses show a concerning disparity: While non-White first authors are equally likely as their White counterparts to publish in top journals, they receive significantly fewer citations. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Enhancing risk communication and environmental crisis management through satellite imagery and AI for air quality index estimation(2024-06-01) ;Jitkajornwanich, Kulsawasd ;Vijaranakul, Nattadet ;Jaiyen, Saichon ;Srestasathiern, PanuLawawirojwong, SiamDue to climate change, the air pollution problem has become more and more prominent [23]. Air pollution has impacts on people globally, and is considered one of the leading risk factors for premature death worldwide; it was ranked as number 4 according to the website [24]. A study, ‘The Global Burden of Disease,’ reported 4,506,193 deaths were caused by outdoor air pollution in 2019 [22,25]. The air pollution problem is become even more apparent when it comes to developing countries [22], including Thailand, which is considered one of the developing countries [26]. In this research, we focus and analyze the air pollution in Thailand, which has the annual average PM2.5 (particulate matter 2.5) concentration falls in between 15 and 25, classified as the interim target 2 by 2021′s WHO AQG (World Health Organization's Air Quality Guidelines) [27]. (The interim targets refer to areas where the air pollutants concentration is high, with 1 being the highest concentration and decreasing down to 4 [27,28]). However, the methodology proposed here can also be adopted in other areas as well. During the winter in Thailand, Bangkok and its surrounding metroplex have been facing the issue of air pollution (e.g., PM2.5) every year. Currently, air quality measurement is done by simply implementing physical air quality measurement devices at designated—but limited number of locations. In this work, we propose a method that allows us to estimate the Air Quality Index (AQI) on a larger scale by utilizing Landsat 8 images with machine learning techniques. We propose and compare hybrid models with pure regression models to enhance AQI prediction based on satellite images. Our hybrid model consists of two parts as follows: • The classification part and the estimation part, whereas the pure regressor model consists of only one part, which is a pure regression model for AQI estimation. • The two parts of the hybrid model work hand in hand such that the classification part classifies data points into each class of air quality standard, which is then passed to the estimation part to estimate the final AQI. From our experiments, after considering all factors and comparing their performances, we conclude that the hybrid model has a slightly better performance than the pure regressor model, although both models can achieve a generally minimum R<sup>2</sup> (R<sup>2</sup> > 0.7). We also introduced and tested an additional factor, DOY (day of year), and incorporated it into our model. Additional experiments with similar approaches are also performed and compared. And, the results also show that our hybrid model outperform them. Keywords: climate change, air pollution, air quality assessment, air quality index, AQI, machine learning, AI, Landsat 8, satellite imagery analysis, environmental data analysis, natural disaster monitoring and management, crisis and disaster management and communication. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Visualizing Political Communication Trends across Generations on X (Twitter): Insights Through Topic Modeling and Word Clouds(2024-01-01) ;Udomwisanpat, Prinwat ;Jitkajornwanich, Kulsawasd ;Kraishan, Obada ;Srestasatheirn, PanuLawawirojwong, SiamThis study examines the interests and significance of words on Twitter (or X) across different generational groups: Baby Boomers, Generation X, Generation Y, and Generation Z. Using Topic Modeling with Latent Dirichlet Allocation (LDA), the research explores relationships and word importance within each group. As the results of topic modeling are not always easy to interpret, we used word cloud visualization to help make sense of the results for each generation. The findings reveal distinct patterns: Baby Boomers frequently mention print media, news websites, and prominent Thai political figures; Generation X emphasizes individuals and local political issues in Bangkok; Generation Y discusses political and social events; and Generation Z uniquely questions political and social activities. This research methodology is applicable across languages and tasks, offering insights into generational behaviors and interests. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Unleashing Hidden Business Insights: Harnessing Unstructured Big Data through Text Analysis, NLP, and Visualizations for Budgetary Decisions in Governmental Organizations(2024-01-01) ;Kongthong, Chanwit ;Jitkajornwanich, KulsawasdIntakosum, SarunProcessing Thai language texts can be a challenge due to the complexities of the language, particularly texts from social media and online platforms. This paper introduces an analysis and visualization framework specifically designed to tackle the intricacies associated with processing the Thai language data within the context of online textual content, by utilizing natural language processing (NLP) and visualization techniques. The objectives of this study were to develop an effective Thai text data analysis and visualization framework that allows us to effectively and automatically get a better understanding of the content embedded in Thai textual data. The methodology initiated with a review of existing analysis frameworks and visualization techniques with a specific focus on Thai. The data collection phase encompassed a diverse corpus of Thai text data gathered from online sources. The selected data underwent preprocessing to address language-specific challenges. The proposed Thai analysis and visualization framework consists of multiple stages. Each stage is tailored to accommodate the intricacies of the Thai language, facilitating improved information extraction and text comprehension. The proposed visualization techniques utilize interactive graphs, such as bar charts, line charts, pie charts and donut charts, to offer intuitive and insightful representations of the processed data. Results from our case study show the effectiveness of our Thai analysis framework and visualization techniques in capturing crucial information from online contents written in Thai from governmental organizations. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Improving OpenAI's Whisper Model for Transcribing Homophones in Legal News(2024-01-01) ;Siriket, Lattapon ;Jitkajornwanich, Kulsawasd ;Jaiyen, SaichonIntakosum, SarunThe 'Whisper' model provides a tool for those who require transcription of human voice. It equips with opensource features and diverse functionalities. The model is capable of effectively deciphering messages in multiple languages, including support for the Thai language. This paper focuses on improving the transcription process of Thai homophones using the Whisper model in reducing the word error rate (WER). We focus on words in the legal news category and identify factors that lead to Whisper's incorrect sound predictions. We examined homophones using snippets of legal news video clips and compiled them into a homophone dictionary. We compare words extracted from the Whisper model by determining the word error rate and spelling of words. Based on the initial results obtained from the original Whisper model and the created homophone dictionary, 48 % of the words were incorrectly transcribed out of a total of 94 words. Then, we propose a methodology by which the performance of the Whisper is improved. That way, the automatic speech recognition of Thai language using the Whisper model can fully be utilized and used in other applications. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Leveraging Race Prediction Algorithms to Enhance Team Composition in Big Data Science Teams(2024-01-01) ;Chumthong, Thanathip ;Jitkajornwanich, Kulsawasd ;Kraishan, Obada ;Kee, Kerk F.Narabin, AkanAs big data science projects scale in complexity, optimizing team composition has become vital for improving creativity, productivity, and project success. We explore the possibility of incorporating race prediction algorithms for enhancing racial diversity in team composition in big data science projects. This paper evaluates five race prediction algorithms - wru, ethnicolr, ethnicolr2, pyethnicity, and rethnicity - and then discuss their potential in supporting racially diverse team assembly in big data projects. Utilizing three datasets, we assess algorithm performance and applicability, emphasizing their role in building balanced teams that enhance agility, inclusivity, and bias mitigation. We present an actionable methodology for integrating demographic insights into team management. In addition, we propose ethical safeguards to ensure responsible race prediction use, recommending data privacy measures, aggregate-only data handling, and transparency in communication. We argue that when used within ethical constraints, race prediction can support robust team processes, reduce reliance on less diverse teams, and ultimately facilitate more creative and equitable big data project outcomes. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Question Classification for Thai Conversational Chatbots Using Artificial Neural Networks and Multilingual BERT Models(2023-01-01) ;Thananukhun, Kit ;Jaiyen, Saichon ;Jitkajornwanich, KulsawasdHanskunatai, AnantapornQuestion-Answering (QA) models are part of Natural Language Processing (NLP) field used for ensuring questions match the answers appropriately. QA consists of several steps, one of which is called Question Classification, which is to classify the context of communication. In this step, it categorizes group of questions based on what users need to know in order to combine answers within the same category and respond accurately. It helps saving us time to search for answers as well. In this paper, we present a question classification model for Thai Conversational Chatbot using Artificial Neural Network and Multilingual Bidirectional Encoder Representations from Transformer (BERT) models using BERT-base multilingual cased combined with Multilayer Perceptron (MLP). The method yields the highest accuracy of 92.57%, compared to the BERT-base multilingual cased combined with other classification models, including Support Vector Machine (SVM), Naive Bayes (NB), K-Nearest Neighbors (KNN) and Decision Trees (DTs) with the accuracy scores of 88.57%, 80.00%, 78.57% and 60.29%, respectively. In addition, we also compare the performance of our proposed BERT model with another well-known Thai word embedding model, called Thai2Vec, which also combines with other classification models including MLP, SVM, NB, KNN and DTs, and their results of accuracies are: 85.71%, 85.71%, 75.71%, 75.71% and 58.86%, respectively. From the experiments, the BERT model combined with MLP can achieve the highest performance in term of accuracy among other methods. - Some of the metrics are blocked by yourconsent settings
Item type:Item, A Performance Comparison between GIS-based and Neuron Network Methods for Flood Susceptibility Assessment in Ayutthaya Province(2022-01-15) ;Vajeethaveesin, Thanat ;Panboonyuen, Teerapong ;Lawawironjwong, Siam ;Srestasathiern, PanuJaiyen, SaichonFlooding has been a long withstanding issue in Thailand. Due to its geographical setup, mitigation and management of floods are challenging and hard to execute. One of the tools used in managing the events is “flood susceptibility mapping,” in which an incident probability as well as a rescue path is estimated and planned. To create one, the traditional GIS method called FRAM (flood risk assessment model), combined with AHP (analytical hierarchy process), is used and implemented on ArcGIS software. In this method, we first created a comparison table to compute weights for each of the selected factors. Then the computed weights were used in the FRAM model in ArcGIS to create a flood susceptibility map for each region. Each region was then classified as very high, high, medium, low, and very low risk. On the other hand, in computer science, machine learning and AI are prevalent and being adopted to various domains, promising the effectiveness of the method, potentially beat the forementioned traditional method. Therefore, ANN (artificial neural network) is adopted in this work to create the flood susceptibility map. The ANN technique is developed by using causal factors. The ANN classifies areas as either flood areas or flood-free areas. The 2 methods from different disciplines (GIS and Computer Science) are applied and described in this paper with the intention to prove whether the machine learning is really efficient and can outperform the traditional GIS approach. Data on Thailand’ s Ayutthaya Province is used in this work as a case study-in order to assess flood prone areas and compared for performance evaluation. Both of which use the 6 selected factors according to the literature: (i) flow accumulation, (ii) elevation, (iii) land use, (iv) rainfall intensity, (v) slope and (vi) soil types. The results from the 2 methods were verified with historical flood data and compared. The results showed that ANN (obtained via sensitivity analysis) outperformed the FRAM with precision of 79.90 %, recall of 79.04 %, F1-score of 79.08 % and accuracy of 79.31 %. In addition, we found that (according to our ANN experiments) the main causal factors related to flood susceptibility map only included 3 factors: flow accumulation, elevation, and soil types. Therefore, the proposed methodology for assessment of flood susceptibility areas using these 3 factors could be considered sufficient and applied to other regions in related applications, when needed.
