KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
3 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Discovery of significant porcine SNPs for swine breed identification by a hybrid of information gain, genetic algorithm, and frequency feature selection technique(2020-05-26) ;Pasupa, Kitsuchart ;Rathasamuth, WanthaneeTongsima, SissadesBackground: The number of porcine Single Nucleotide Polymorphisms (SNPs) used in genetic association studies is very large, suitable for statistical testing. However, in breed classification problem, one needs to have a much smaller porcine-classifying SNPs (PCSNPs) set that could accurately classify pigs into different breeds. This study attempted to find such PCSNPs by using several combinations of feature selection and classification methods. We experimented with different combinations of feature selection methods including information gain, conventional as well as modified genetic algorithms, and our developed frequency feature selection method in combination with a common classification method, Support Vector Machine, to evaluate the method's performance. Experiments were conducted on a comprehensive data set containing SNPs from native pigs from America, Europe, Africa, and Asia including Chinese breeds, Vietnamese breeds, and hybrid breeds from Thailand. Results: The best combination of feature selection methods - information gain, modified genetic algorithm, and frequency feature selection hybrid - was able to reduce the number of possible PCSNPs to only 1.62% (164 PCSNPs) of the total number of SNPs (10,210 SNPs) while maintaining a high classification accuracy (95.12%). Moreover, the near-identical performance of this PCSNPs set to those of bigger data sets as well as even the entire data set. Moreover, most PCSNPs were well-matched to a set of 94 genes in the PANTHER pathway, conforming to a suggestion by the Porcine Genomic Sequencing Initiative. Conclusions: The best hybrid method truly provided a sufficiently small number of porcine SNPs that accurately classified swine breeds. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Selection of a minimal number of significant porcine snps by an information gain and genetic algorithm hybrid model(2019-01-01) ;Rathasamuth, Wanthanee ;Pasupa, KitsuchartTongsima, SissadesA panel of a large number of common Single Nucleotide Polymorphisms (SNPs) distributed across an entire porcine genome has been widely used to represent genetic variability of pigs. With the advent of SNP-array technology, a genome-wide genetic profile of a specimen can be easily observed. Among the large number of such variations, there exists a much smaller subset of the SNP panel that could equally be used to correctly identify the corresponding breed. This work presents a SNP selection heuristic that can still be used effectively in the breed classification. The features were selected by combining a filter method and a wrapper method-information gain method and genetic algorithm-plus a feature frequency selection step, while classification used a support vector machine. We were able to reduce the number of significant SNPs to 0.86 % of the total number of SNPs in a swine dataset with 94.80 % classification accuracy. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, PCA-based informative SNP selection for analyzing population structure(2016-12-19) ;Limpiti, Tulaya ;Intarapanich, ApichartTongsima, SissadesPhenotypic differences among individuals of the same species are the result of a set of genetic variations which can be observed in the DNA sequence. To conduct a population genetic study, a high throughput genotyping platform such as Single Nucleotide Polymorphism (SNP) array is popularly used to obtain a large set of SNPs for each individual. However, analyzing today's genotypic data can be computationally expensive due to its large size and complexity. Faulty substructure may also be detected if the data is noisy from redundant or non-informative SNPs. Considerable efforts have been done to extract a smaller informative SNP subset that still represents the same intrinsic structure of populations within a data set as the full panel of SNPs. This work describes a foundation of a PCA-based informative marker selection technique. The proposed technique is simple and efficient. It improves upon another spectral analysis technique called PCA-correlated SNPs. A new informativeness score based on a basis function expansion of the SNP variation patterns across individuals is introduced. Such score is computed for each SNP to select a subset of SNPs with the best scores. Using a bovine data set, we demonstrate that our technique is superior to the PCAcorrelated SNPs method, which requires accurate rank estimation to perform well. In contrast, our method is robust to the assumed rank of the data. High data representation accuracy is also achieved after a significant reduction of the number of SNPs while retaining information about the underlying population structure from the original data.
