A novel technique for feature subset selection based on Cosine similarity

Loading...
Thumbnail Image

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Nowadays, data mining has been playing an important role in the various disciplines of sciences and technologies. Data mining is composed of many tasks but one of the essential procedures of data mining is feature selection, which is the technique mostly based on the machine learning for selecting a subset consisted of significant features, building a stronger learning model, and enhancing the efficiency of prediction rate. Normally, a processing of building a learning model from the huge amount of collected data needs high computation cost. Therefore, with feature selection, the computation cost can be reduced by selecting relevant features. In the previous researches on feature selection, the criteria and algorithms for selecting the features from the raw data are mostly complicated and difficult to implement. Therefore, this paper presents a novel method by applied Cosine similarity to feature selection method. The proposed algorithm begins with selecting a robust feature subset using the Cosine similarity. This method is the simple algorithm using smallerstorage space, reducing computation time and gaining higher predictive performance. During the evaluation phase, the tendata sets from UCI benchmark data sets are used to evaluate the performance of proposed approach by using the C5.0, CARTand Neural Networks classifiers. Experimental results show that the method based on the Cosine similarity can improve the performance of accuracy detection rate with less error rate.

Description

Keywords

Accuracy detection rate, Classification, Cosine similarity, Selection

Citation

Applied Mathematical Sciences, 6(133-136), 6627-6655, 2012

Collections

Endorsement

Review

Supplemented By

Referenced By