Comparison of Sampling Methods for Imbalanced Data Classification in Random Forest

dc.contributor.authorPaing, May Phu
dc.contributor.authorPintavirooj, C.
dc.contributor.authorTungjitkusolmun, Supan
dc.contributor.authorChoomchuay, Somsak
dc.contributor.authorHamamoto, Kazuhiko
dc.date.accessioned2026-08-06T10:23:55Z
dc.date.available2026-08-06T10:23:55Z
dc.date.issued2019-01-10
dc.description.abstractImbalanced data classification is a serious and challenging task for most of the medical image diagnosis applications. They usually produce a larger number of false samples compared to the actual ones. That is the number of samples for the class of interest (minority) is significantly fewer than other types of class (majority). The classification performed using such data is called imbalanced data classification. As a consequence, the learning model bias towards the majority class and fails the classification of the minority class. Data sampling and ensemble methods are common ways to compensate for this issue. Random forest (RF), an ensemble of multiple decision trees, is very famous in both of the classification and regression problems because of its robust and accurate predictions. However, it also suffers class bias in the imbalanced data classification problems. This paper proposes and compares different sampling methods to solve the imbalanced data classification in RF.
dc.identifier.citationBmeicon 2018 11th Biomedical Engineering International Conference, 2019
dc.identifier.doi10.1109/BMEiCON.2018.8609946
dc.identifier.other2-s2.0-85062088168
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/9683
dc.sourceBmeicon 2018 11th Biomedical Engineering International Conference
dc.subjectImbalanced data classification
dc.subjectRandom forest
dc.subjectSampling
dc.titleComparison of Sampling Methods for Imbalanced Data Classification in Random Forest
dc.typeConference Paper

Files

Collections