A new feature selection based on class dependency and feature dissimilarity
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
Feature selection method is an important task for data preprocessing in data mining. Before a classifier learns the training data, there are a lot of features in each data set that makes the learning process slower. It is not appropriated for big data analytics. This paper proposes feature selection method based on the class dependency and feature dissimilarity (CDFD) using mutual information and Euclidean distance. The mutual information is applied to determine the dependency between the feature and the class if the dataset contains discrete data. If the dataset contains continuous data, the correlation between the feature and the class is used instead. The Euclidean distance is used for reducing the duplicated features based on dissimilarity between features. The experiments are conducted on five datasets. From the experimental results, the propose feature selection method can reduce the number of features in the data set and reduce the classification error of classifiers. Furthermore, it can be applied to discrete and continuous data and it can help classifiers improving their classification accuracies and reducing the computational times for learning.
