KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 10 of 18
  • Some of the metrics are blocked by your 
    Item type:Item,
    Multi-Label Classification of Foreign Tourists' Opinions on Thailand Tourism Development
    (2024-10-03)
    Suramanka, Lalita
    ;
    Hanskunatai, Anantaporn
    The enhancement of tourism quality in Thailand through the understanding and utilization of foreign tourists' opinions presents challenges due to the extensive volume of data involved. This research proposes a two-fold approach to address this issue: (1) the development of an opinion classification model, and (2) the analysis of tourists' opinions through a dashboard. A dataset, compiled by the Tourism Authority of Thailand (TAT) and consisting of opinions from foreign tourists regarding areas for improvement in Thai tourism, was utilized. A total of 2,249 comments were collected. Experimental results demonstrate that the use of data augmentation, feature selection, Multi-label transformation using Classifier Chains, and the Random Forest classification model on the training dataset yields promising results with an accuracy rate of 80%, precision of 90%, recall of 81%, F1-score of 85%, and Hamming loss of 0.04. Analysis from the dashboard revealed the top three key areas for improvement: communication/language, traffic/public transportation, and cleanliness/hygiene.
  • Some of the metrics are blocked by your 
    Item type:Item,
    A deep neural network-correlation phase sensitive mask based estimation to improve speech intelligibility
    (2023-09-01)
    Sivapatham, Shoba
    ;
    Kar, Asutosh
    ;
    Bodile, Roshan
    ;
    Mladenovic, Vladimir
    ;
    Sooraksa, Pitikhate
    General masking-based speech enhancement using a deep learning architecture (DNN) approach focuses on the spectral values of the speech in order to show improvement in intelligibility. But, the residual noise present in the phase spectrum and inter-channel correlation dependency between noise and noisy speech can impact the results of intelligibility in speech enhancement. This research work proposes a correlation phase-sensitive novel mask which contains phase, magnitude spectral and inter-channel correlation for a tangible improvement in speech intelligibility. The correlation parameter finds the dependency between the signals and phase spectrum factor eliminates the residual noise. In addition, selecting the prior features from the feature combination also helps in reducing the dimensionality and increases the accuracy of the enhanced speech. This work also aims to decrease the complexity of the DNN by analysing the network with different parameters. The performance of the mask is evaluated with various intelligibility factors. The proposed mask has been compared with different mask estimators. The proposed mask has generated estimated speech with an increase in the intelligibility of 0.02-0.001 over six different noises and four different signal-to-noise (SNR) levels.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Efficient distributed SNP selection by a Modified Binary Flower Pollination Algorithm
    (2020-07-01)
    Rathasamuth, Wanthanee
    ;
    Pasupa, Kitsuchart
    Porcine Single Nucleotide Polymorphisms (SNPs-certain pieces of nucleotide in a DNA sequence) can be indirectly associated with traits of an individual pig, like its meat quality or resistance to common diseases. It is most desirable to obtain a smallest number of most significant SNPs in genomics research, and several computer classification algorithms have been used to do so. For instance, for breed classification, one needs to obtain a set of a much smaller number of significant SNPs than that of the entire SNP data set. This study attempted to find such significant porcine SNPs by using computational feature selection and classification methods. In a preliminary trial, a binary flower pollination algorithm (BFPA) was used and shown not to able to reduce the number of selected SNPs to a sufficiently low number. Therefore, to achieve our objective, we developed a vertically distributed feature selection method incorporating a modified BFPA and a support vector machine classifier for selecting significant porcine SNPs. The developed method was evaluated and compared against four baseline methods. It provided the smallest average number of significant SNPs (128.40) that resulted in 94.57% classification accuracy. This and other findings in this study may directly benefit researchers in the bioinformatics field in their effort to map SNPs.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Discovery of significant porcine SNPs for swine breed identification by a hybrid of information gain, genetic algorithm, and frequency feature selection technique
    (2020-05-26)
    Pasupa, Kitsuchart
    ;
    Rathasamuth, Wanthanee
    ;
    Tongsima, Sissades
    Background: The number of porcine Single Nucleotide Polymorphisms (SNPs) used in genetic association studies is very large, suitable for statistical testing. However, in breed classification problem, one needs to have a much smaller porcine-classifying SNPs (PCSNPs) set that could accurately classify pigs into different breeds. This study attempted to find such PCSNPs by using several combinations of feature selection and classification methods. We experimented with different combinations of feature selection methods including information gain, conventional as well as modified genetic algorithms, and our developed frequency feature selection method in combination with a common classification method, Support Vector Machine, to evaluate the method's performance. Experiments were conducted on a comprehensive data set containing SNPs from native pigs from America, Europe, Africa, and Asia including Chinese breeds, Vietnamese breeds, and hybrid breeds from Thailand. Results: The best combination of feature selection methods - information gain, modified genetic algorithm, and frequency feature selection hybrid - was able to reduce the number of possible PCSNPs to only 1.62% (164 PCSNPs) of the total number of SNPs (10,210 SNPs) while maintaining a high classification accuracy (95.12%). Moreover, the near-identical performance of this PCSNPs set to those of bigger data sets as well as even the entire data set. Moreover, most PCSNPs were well-matched to a set of 94 genes in the PANTHER pathway, conforming to a suggestion by the Porcine Genomic Sequencing Initiative. Conclusions: The best hybrid method truly provided a sufficiently small number of porcine SNPs that accurately classified swine breeds.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Efficient breadth-first reduct search
    (2020-05-01)
    Boonjing, Veera
    ;
    Chanvarasuth, Pisit
    This paper formulates the problem of determining all reducts of an information system as a graph search problem. The search space is represented in the form of a rooted graph. The proposed algorithm uses a breadth-first search strategy to search for all reducts starting from the graph root. It expands nodes in breadth-first order and uses a pruning rule to decrease the search space. It is mathematically shown that the proposed algorithm is both time and space efficient.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Feature Selection Method Based on Correlation Tree
    (2020-01-01)
    Yapila, Prajak
    ;
    Threepak, Thanunchai
    Machine learning is one of techniques adapted to detect intrusion for cyber security. One of importance techniques to find anomaly is classification. But classification with huge dataset has the resources and time consumption. Feature selection is choice to reduce the data dimension to improve processing performance. In this paper, we introduce the new feature selection method that selects some fields of data set using position of each feature in correlation tree. Then, the result from the correlation tree feature selection of KDDCUP’99 data set are compared with two feature selection technique, correlation of coefficient (CC-type) and BFS by using three reference classifier, Decision Tree (DT), Random Forest (RF), and Naive Bayes (NB).
  • Some of the metrics are blocked by your 
    Item type:Item,
    A Modified Binary Flower Pollination Algorithm: A Fast and Effective Combination of Feature Selection Techniques for SNP Classification
    (2019-10-01)
    Rathasamuth, Wanthanee
    ;
    Pasupa, Kitsuchart
    Single nucleotide polymorphism (SNP) is a genetic trait responsible for the differences in the characteristics of individuals of a living species. Machine learning has been brought in to classify swine breed according to their SNPs. However, since the number of samples (number of pigs sampled) is usually much smaller than the number of features (SNPs) to classify, there may occur an overfitting problem. Therefore, some feature selection techniques were applied to the entire SNPs to reduce them to a much smaller number of most significant SNPs to be used in the classification. In this study, we used information gain in combination with binary flower pollination algorithm for feature selection as well as a cut-off-point-finding threshold for specifying a 0 or 1 value for a position in the solution vector and a GA bit-flip mutation operator. We called it Modified-BFPA. The classifier was SVM. Evaluated against a few other feature selection techniques, our combination of techniques was, at the very least, competitive to those. It selected only 1.76 % of most significant SNPs from the entire set of 10,210 SNPs. The SNPs that it selected provided 95.12 % classification accuracy. Moreover, it was fast: an average of 1.60 iterations in combination with SVM to find a set of best SNPs that provided the highest classification accuracy.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Cervical cancer detection and classification from pap smear images
    (2019-09-16)
    Win, Kyi Pyar
    ;
    Kitjaidure, Yuttana
    ;
    Paing, May Phyu
    ;
    Hamamoto, Kazuhiko
    In this paper, we propose a framework for detection and classification of cervical cancer from pap smear images. Early detection and accurate diagnosis of cervical cancer can reduce the death rate of cervical cancer patients. Pap smear or pap test is the most popular technique for early detection of cervical cancer. However, the manual analysis is labor intensive and time consuming process which relies on expert cytologist. Hence, it is needed to develop a computer aided diagnosis system to make the pap smear test more accurate and reliable. The objective of this paper is to present an innovative idea of applying random forest algorithm (RF) as a feature selection method using proposed bagging ensemble classifier for improving the predictive performance. The four basic steps of cervical cancer detection and classification system, image enhancement, segmentation, feature extraction and classification were used. K-means clustering combining with morphology operations obtained good segmentation for cell nuclei and cytoplasm. The most important features, shape, color and texture of nuclei and cytoplasm were applied to detect cervical cancer. To improve the accuracy of prediction results, random forest (RF) algorithm was used as a feature selection method. In classification stage, bagging ensemble classifier was applied which aggregated the results of five classifiers, linear discriminant (LD), support vector machine (SVM), weighted k-nearest neighbor (KNN), boosted trees and bagged trees. Herlev data set was used to prove the effectiveness of our proposed method. According to the experimental results, the high classification accuracy was achieved with top10 features using our proposed combined classifier. The accuracy was 97.83% in two class problem and 81.54% in seven class problem. When the results were compared with five classifiers, our proposed method was significantly better in two class and seven class problems.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Music genre classification of audio signals using particle swarm optimization and stacking ensemble
    (2019-03-01)
    Leartpantulak, Krittika
    ;
    Kitjaidure, Yuttana
    Genre classification is a process of grouping similarities, such as patterns, styles, or objectives with management data as already in the music (e.g. pop and rock). It is used along with the classification of topics. This paper will classify songs from audio signal to a hierarchy of musical genre by using feature extraction. Trimbral texture, rhythmic content and pitch content are used as the main feature sets. Feature selection is selected by using Particle Swarm Optimization (PSO) and sent selected feature to classification. The result in classification has low accuracy. Thus, using stacking ensemble method is to improve the prediction. In this paper, the purpose is to improve the prediction by using stacking ensemble method. Stacking ensemble that have the second level is base classifier and meta-classifier. In base classifier consists of 5 classification; K-Nearest Neighbors (k-NN), Decision Tree (DT), Random Forest, Support Vector Machines (SVM); and Naïve Bayes. This process is generated to build multiple classifier predictors and sent it to meta-classifier. In the process of meta-classifier will create new model to predict test data. The new model has been created from neural network which train data is the output of base classifier.
  • Some of the metrics are blocked by your 
    Item type:Item,
    Selection of a minimal number of significant porcine snps by an information gain and genetic algorithm hybrid model
    (2019-01-01)
    Rathasamuth, Wanthanee
    ;
    Pasupa, Kitsuchart
    ;
    Tongsima, Sissades
    A panel of a large number of common Single Nucleotide Polymorphisms (SNPs) distributed across an entire porcine genome has been widely used to represent genetic variability of pigs. With the advent of SNP-array technology, a genome-wide genetic profile of a specimen can be easily observed. Among the large number of such variations, there exists a much smaller subset of the SNP panel that could equally be used to correctly identify the corresponding breed. This work presents a SNP selection heuristic that can still be used effectively in the breed classification. The features were selected by combining a filter method and a wrapper method-information gain method and genetic algorithm-plus a feature frequency selection step, while classification used a support vector machine. We were able to reduce the number of significant SNPs to 0.86 % of the total number of SNPs in a swine dataset with 94.80 % classification accuracy.