Now showing 1 - 2 of 2
  • Some of the metrics are blocked by your 
    Item type:Publication,
    A Classification Study in High-Dimensional Data of Linear Discriminant Analysis and Regularized Discriminant Analysis
    (2023-01-01) ;
    Banditvilai, Somsri
    The objective of this work is to compare linear discriminant analysis (LDA) and regularized discriminant analysis (RDA) for classification in high-dimensional data. This dataset consists of the response variable as a binary or dichotomous variable and the explanatory as a continuous variable. The LDA and RDA methods are well-known in statistical and probabilistic learning classification. The LDA has created the decision boundary as a linear function where the covariance of two classes is equal. Then the RDA is extended from the LDA to resolve the estimated covariance when the number of observations exceeds the explanatory variables, or called high-dimensional data. The explanatory dataset is generated from the normal distribution, contaminated normal distribution, and uniform distribution. The binary of the response variables is computed from the logit function depending on the explanatory variable. The highest average accuracy percentage evaluates to propose the performance of the classification methods in several situations. Through simulation results, the LDA was successful when using large sample sizes, but the RDA performed when using the most sample sizes.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Comparison of Logistic Regression and Discriminant Analysis for Classification of Multicollinearity Data
    (2023-01-01)
    The objective of this study is to concentrate on the classification method of the logistic regression and the discriminant analysis by using the simulation dataset and the liver patients as the actual data. These datasets are used the binary dependent variable depending on the correlated independent variables or called multicollinearity data. The standard classification method is logistic regression, which uses the logit function's probability to conduct the dichotomous dependent variable. The iteration process can be solved to estimate logit function parameters and explain the relationship between a dependent binary variable and independent variables. Discriminant analysis is a powerful classification based on linear discriminant analysis (LDA), quadratic discriminant analysis (QDA), and regularized discriminant analysis (RDA). These methods consider the decision boundaries by building a classifier model on the multivariate normal distribution. LDA defines the standard covariance matrix, but QDA has an individual covariance matrix. RDA extends from QDA by setting the regularized parameter to estimate the covariance matrix. In the case of the simulation study, the independent variables are generated by defining the constant correlation on the multivariate normal distribution that made the multicollinearity problem. Then the binary response variable can be approximated from the logit function. For application to actual data, we expressed the classification of type liver and non-liver patients as the dependent variables and obtained patient personal information on the nine independent variables. The highest average percentage of accuracy determines the performance of these methods. The results have shown that the logistic regression was successful when using small independent variables, but the RDA performed when using large independent variables.