KMITL

Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1

Browse

Search Results

Now showing 1 - 9 of 9
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Mutual information rough sets feature selection and classification for microarray data analysis
    (2014-01-01)
    Ounsrimuang, Pimolrat
    ;
    Boonjing, Veera
    The feature selection (FS) techniques aim to reduce the subset size of an original data set, which are retained in the most useful information by selecting the most informative feature instead of irrelevant or redundant features. The benefits of FS for classification analysis can reduce the input data, improved predictive accuracy, learned knowledge is that easily understood, and reduced execution time. Many approaches based on rough set theory up to now, have operated the dependency function for measuring the goodness of the feature. However, there is not tolerance to noisy or inconsistency data, especially on high dimensional data microarray data sets. Moreover, mostly relevant information could be invisible by using only information from a positive region but neglecting a boundary region, mostly relevant may be invisible. Therefore, this paper proposes the maximal positive region and minimal boundary region criterion, based on rough set and mutual information, which use the different values among the information contained in the positive region, and the information contained in the boundary region. The experimental results indicate that our proposed method can increase the classification accuracy. © 2014 Pushpa Publishing House, Allahabad, India.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    The proposed algorithm for feature selection based on rough set and mutual information
    (2014-01-01)
    Ounsrimuang, Pimolrat
    ;
    Boonjing, Veera
    The feature selection approaches based on rough set theory aim to reduce the input data for improvement classification accuracy. Most existing approaches have concerned the discernibility relation to find the features, and have employed the dependency function for measuring the goodness of feature. The most relevant information cannot be visible by using information from discernibility relation only, so that neglecting indiscernibility relation, mostly relevant may be invisible. Moreover, their results are not tolerant to noisy or inconsistency data. Therefore, this paper proposes new algorithm based on rough set theory, which concerned both the discernibility and indiscernibility relations. The experimental results show that our approach gives higher classification accuracy than existing approaches. © 2014 Pushpa Publishing House, Allahabad, India.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Rough-mutual feature selection based on min-uncertainty and max-certainty
    (2012-01-01)
    Foitong, Sombut
    ;
    Pinngern, Ouen
    ;
    Attachoo, Boonwat
    Feature selection (FS) plays an important role in pattern recognition and machine learning. FS is applied to dimensionality reduction and its purpose is to select a subset of the original features of a data set which is rich in the most useful information. Most existing FS methods based on rough set theory focus on dependency function, which is based on lower approximation as for evaluating the goodness of a feature subset. However, by determining only information from a positive region but neglecting a boundary region, most relevant information could be invisible. This paper, the maximal lower approximation (Max Certainty) minimal boundary region (Mm Uncertainty) criterion, focuses on feature selection methods based on rough set and mutual infonnation which use different values among the lower approximation information and the information contained in the boundary region. The use of this idea can result in higher predictive accuracy than those obtained using the measure based on the positive region (certainty region) alone. This demonstrates that much valuable information can be extracted by using this idea. Experimental results are illustrated for discrete, continuous, and microarray data and compared with other FS methods in terms of subset size and classification accuracy. key words: rough sets, mutual information, feature selection, boundary region, classification accuracy. Copyright © 2012 The Institute of Electronics, Information and Communication Engineers.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Novel matrix forms of rough set flow graphs with applications to data integration
    (2010-11-01)
    Chitcharoen, Doungrat
    ;
    Pattaraintakorn, Puntip
    Pawlak's flow graphs have attracted both practical and theoretical researchers because of their ability to visualize information flow. In this paper, we invent a new schema to represent throughflow of a flow graph and three coefficients of both normalized and combined normalized flow graphs in matrix form. Alternatively, starting from a flow graph with its throughflow matrix, we reform Pawlak's formulas to calculate these three coefficients in flow graphs by using matrix properties. While traditional algorithms for computing these three coefficients of the connection are exponential in l, an algorithm using our matrix representation is polynomial in l, where l is the number of layers of a flow graph. The matrix form can simplify computation, improve time complexity, alleviate problems due to missing coefficients and hence help to widen the applications of flow graphs. Practically, data sets often reside at different sources (heterogeneous data sources). Their individual analysis at each source is inadequate and requires special treatment. Hence, we introduce a composition method for flow graphs and corresponding formulas for calculating their coefficients which can omit some data sharing. We provide a real-world experiment on the Promotion of Academic Olympiads and Development of Science Education Foundation (POSN) data set which illustrates a desirable outcome and the advantages of the proposed matrix forms and the composition method. © 2010 Elsevier Ltd. All rights reserved.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Estimating optimal feature subsets using mutual information feature selector and rough sets
    (2009-01-01)
    Foitong, Sombut
    ;
    Rojanavasu, Pornthep
    ;
    Attachoo, Boonwat
    ;
    Pinngern, Ouen
    Mutual Information (MI) is a good selector of relevance between input and output feature and have been used as a measure for ranking features in several feature selection methods. Theses methods cannot estimate optimal feature subsets by themselves, but depend on user defined performance. In this paper, we propose estimation of optimal feature subsets by using rough sets to determine candidate feature subset which receives from MI feature selector. The experiment shows that we can correct nonlinear problems and problems in situation of two or more combined features are dominant features, maintain an improve classification accuracy. © Springer-Verlag Berlin Heidelberg 2009.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Rough sets based approach to reduct approximation: RSBARA
    (2008-12-01)
    Foitong, Sombut
    ;
    Srinil, Phaitoon
    ;
    Pinngern, Ouen
    Attribute reduction is the process of choosing a subset of attributes from the original set of attributes forming patterns in a given dataset. The subset should be necessary and sufficient to describe target concepts. Rough set theory has been used as an attribute reduction method with impressive success, but current methods are inadequate at finding optimal reductions. On the other hand, the optimal reducts can be obtained by using the stochastic approaches, but it is not easy because its computational complexity for computing reducts is at least O (NxM<sup>2</sup>), where N is the number of attributes and M is the total number of objects. In this paper, we propose an algorithm which uses rough set theory to approximate the reduct and reduces the required computational effort to O(N<sup>2</sup>xM). Experimentation is carried out, using UCI data, which compares with a particle swarm approach and other deterministic rough set reduction algorithms. The experimental results show that the purposed method is more efficient both accuracy and attribute reduction. © 2008 IEEE.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    A foundation of rough sets theoretical and computational hybrid intelligent system for survival analysis
    (2008-10-01)
    Pattaraintakorn, Puntip
    ;
    Cercone, Nick
    What do we (not) know about the association between diabetes and survival time? Our study offers an alternative mathematical framework based on rough sets to analyze medical data and provide epidemiology survival analysis with risk factor diabetes. We experiment on three data sets: geriatric, melanoma and Primary Biliary Cirrhosis. A case study reports from 8547 geriatric Canadian patients at the Dalhousie Medical School. Notification status (dead or alive) is treated as the censor attribute and the time lived is treated as the survival time. The analysis result illustrates diabetes is a very significant risk factor to survival time in our geriatric patients data. This paper offers both theoretical and practical guidelines in the construction of a rough sets hybrid intelligent system, for the analysis of real world data. Furthermore, we discuss the potential of rough sets, artificial neural networks (ANNs) and frailty index in predicting survival tendency. © 2008 Elsevier Ltd. All rights reserved.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Hybrid rough sets intelligent system architecture for survival analysis
    (2007-01-01)
    Pattaraintakorn, Puntip
    ;
    Cercone, Nick
    ;
    Naruedomkul, Kanlaya
    Survival analysis challenges researchers because of two issues. First, in practice, the studies do not span wide enough to collect all survival times of each individual patient. All of these patients require censor variables and cannot be analyzed without special treatment. Second, analyzing risk factors to indicate the significance of the effect on survival time is necessary. Hence, we propose "Enhanced Hybrid Rough Sets Intelligent System Architecture for Survival Analysis" (Enhanced HYRIS) that can circumvent these two extra issues. Given the survival data set, Enhanced HYRIS can analyze and construct a life time table and Kaplan-Meier survival curves that account for censor variables. We employ three statistical hypothesis tests and use the p-value to identify the significance of a particular risk factor. Subsequently, rough set theory generates the probe reducts and reducts. Probe reducts and reducts include only a risk factor subset that is large enough to include all of the essential information and small enough for our survival prediction model to be created. Furthermore, in the rule induction stage we offer survival prediction models in the form of decision rules and association rules. In the validation stage, we provide cross validation with ELEM2 as well as decision tree. To demonstrate the utility of our methods, we apply Enhanced HYRIS to various data sets: geriatric, melanoma and primary biliary cirrhosis (PBC) data sets. Our experiments cover analyzing risk factors, performing hypothesis tests and we induce survival prediction models that can predict survival time efficiently and accurately. © Springer-Verlag Berlin Heidelberg 2007.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Rule analysis with rough sets theory
    (2006-11-22)
    Pattaraintakorn, Puntip
    ;
    Cercone, Nick
    ;
    Naruedomkul, Kanlaya
    Postprocessing is a significant step in the data analysis process which is often ignored or glossed over. Once we have a large set of generated rules, how can we elicit the sufficient and necessary rules? In this paper, we propose an alternative approach for decision rule learning with rough sets theory in the postprocessing step called 'ROSERULE'. Essentially, we introduce rule reducts, a sufficient and necessary part which preserves classification of the rule universe, as a rough sets tool for rule analysis. ROSERULE learns and analyzes from the rule set to generate rule reducts which can be used to reduce the number of the rules. This is in contrast to common rule analysis which simply performs rule selection. We illustrate the performance of ROSERULE with several case studies; melanoma, primary biliary cirrhosis, pneumonia and a real-world case study, geriatric data sets. ROSERULE is run on these data sets and the result are a reduced number of rules that successfully preserve the original classification. © 2006 IEEE.