KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
4 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Improved iterative pruning principal component analysis with graph-theoretic hierarchical clustering(2012-10-02) ;Amornbunchornvej, C. ;Limpiti, T. ;Assawamakin, A. ;Intarapanich, A.Tongsima, S.Various unsupervised clustering algorithms have been used to infer population structure in genetic data. The goals are to separate individuals of similar genetic characteristics into clusters and to estimate the number of clusters within each dataset. Among them, a framework called iterative pruning principal component analysis (ipPCA) have been developed. It performs PCA iteratively on subsets of data samples and clusters them using fuzzy c-mean. We believe that the choice of model-based clustering method affects the individual assignments and cluster quality, as well as the estimated number of clusters. Thus, in this paper we introduce a hierarchical tree clustering concept from graph theory, whose performance is independent of cluster shapes, into the ipPCA framework. We also add a PCA-based feature selection technique as a data pre-processing step to reduce data dimension and increase computational efficiency. The resulting algorithm is called HiClust-ipPCA. We illustrate the improved clustering results of the HiClust-ipPCA algorithm using 47-breed bovine and 28-breed sheep datasets. © 2012 IEEE. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Iterative Neighbor-Joining tree clustering algorithm for genotypic data(2012-01-01) ;Amornbunchornvej, C. ;Limpiti, T. ;Assawamakin, A. ;Intarapanich, A.Tongsima, S.Issues to explore in genotypic datasets include the number and characteristic patterns of subpopulations and, possibly, relationships among them. Model-based clustering methods have been adopted to find a number of clusters and the individual assignments. However, they cannot infer genetic relationships among subpopulations the way phylogenetic trees, e.g., the widely-used Neighbor-Joining (NJ) tree, can. In this paper we propose an unsupervised, iterative clustering framework called iNJclust. It performs clustering on an NJ tree with a graph-based partitioning technique. The iterative process enhances the zooming ability and corrects the topology of the final NJ trees. Inference on genetic similarities between subpopulations is also possible. As final outputs, the iNJclust algorithm provides an estimate of the number of clusters, individual assignments, a population tree, as well as sub-trees of the terminal nodes. We illustrate the superior clustering performance of the proposed algorithm using Human 27 populations, bovine 47 breeds, and sheep 28 breeds datasets. © 2012 ICPR Org Committee. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Time-frequency analysis for cancer detection using proteomic MS-spectra(2011-12-01) ;Limpiti, T. ;Assawamakin, A. ;Intarapanich, A.Tongsima, S.Mass spectrum data is proven useful in cancer detection and biomarker discovery. Nevertheless, existing methods which analyze peaks of mass spectrum data still have some limitations, including variation in peak locations among individual samples, noisy data, irreproducibility of peak profiles, and computational burden. We introduce a simple algorithm in this paper which alleviate these drawbacks. Our approach is to analyze the mass spectrum data using time-frequency analysis. The data is transformed to features in the time-frequency domain. Informative features are then selected and subsequently used for detection or classification. To assess the efficacy of the proposed algorithm, we apply our algorithm to cancer detection problem. The performance of the algorithm is evaluated on real ovarian and prostate cancer datasets. The promising detection results with high sensitivity and specificity confirm the potential of our method in cancer detection. The algorithm is also applicable to multi-class classification and biomarker identification problems. © 2011 IEEE. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Iterative PCA for population structure analysis(2011-08-18) ;Limpiti, T. ;Intarapanich, A. ;Assawamakin, A. ;Wangkumhang, P.Tongsima, S.An extension of principal component analysis called ipPCA has been proposed earlier for analyzing structure in genetic data. This non-parametric framework iteratively classifies individuals into subpopulations. However, it is prone to false positives when dealing with large datasets and mixed-type genetic markers. We address these shortcomings by introducing a unified encoding scheme and suggesting a new terminating criterion for ipPCA. To validate the improvements, simulated datasets as well as real bovine and large human genetic datasets are analyzed. It is observed that the estimation of the number of subpopulations and the individual assignment accuracy have been improved. Furthermore, the structure resolved by this approach can be used to identify subset of individuals for further parametric population structure analysis. © 2011 IEEE.
