A new feature selection based on class dependency and feature dissimilarity

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Feature selection method is an important task for data preprocessing in data mining. Before a classifier learns the training data, there are a lot of features in each data set that makes the learning process slower. It is not appropriated for big data analytics. This paper proposes feature selection method based on the class dependency and feature dissimilarity (CDFD) using mutual information and Euclidean distance. The mutual information is applied to determine the dependency between the feature and the class if the dataset contains discrete data. If the dataset contains continuous data, the correlation between the feature and the class is used instead. The Euclidean distance is used for reducing the duplicated features based on dissimilarity between features. The experiments are conducted on five datasets. From the experimental results, the propose feature selection method can reduce the number of features in the data set and reduce the classification error of classifiers. Furthermore, it can be applied to discrete and continuous data and it can help classifiers improving their classification accuracies and reducing the computational times for learning.

Description

Keywords

Correlation, Euclidean Distance, Feature Selection, Mutual Information

Citation

Icaicta 2015 2015 International Conference on Advanced Informatics Concepts Theory and Applications, 2015

Collections

Endorsement

Review

Supplemented By

Referenced By