Now showing 1 - 10 of 12
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Thai traffic sign detection and recognition for driver assistance
    (2018-11-05)
    Promlainak, Sakan
    ;
    ;
    Kuengwong, Jirapat
    ;
    Kuroki, Yoshimitsu
    Nowadays, driver assistance systems are embedded with some expensive cars, but more importantly, those systems are not able to recognize Thai traffic signs. This paper proposes a Thai traffic sign detection and recognition system. The proposed system is implemented with two main processes: Thai traffic sign detection and recognition. For the former process, a cascade classifier trained with histogram of oriented gradient (HOG) features is used to generate a trained model for a sign detector, and then Viola-Jones cascade detector is used to classify sign and non-sign objects of the input image. For the latter process, a linear support vector machine (SVM) learner trained with HOG features is used to generate the trained model for sign symbol recognition, and then a SVM class prediction is applied for recognizing the HOG features of the detected sign. Based on a real world data-set, the proposed system can correctly detectand recognize Thai traffic signs in near real time.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Text-background decomposition for thai text localization and recognition in natural scenes
    (2014-01-12) ; ;
    Suttapakti, Ungsumalee
    ;
    Boonchukusol, Pimlak
    ;
    Thai text localization and recognition in natural scenes is still a grand challenge in current applications. However, the efficiency of recognition rates depends on text localization, i.e., the higher purity of text-background decomposition leads to the higher accuracy rate of character recognition. In order to achieve this purpose, the text-background decomposition methods, namely adaptive boundary clustering (ABC) and n-point boundary clustering (n-PBC), are proposed to improve a precision of text localization. These methods are evaluated by self-en-tropy for purity measure. Based on 300 test images, the experimental results demonstrate that the ABC method achieves the very low self-entropy, i.e., the low self-entropy implies the good decomposition of text and background. Furthermore, based on 8,077 characters in natural scene test images, the ABC method helps increase the precision of text localization and improves the accuracy rate of character recognition, when compared to the conventional methods.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Improved Thai text detection from natural scenes
    (2013-01-01) ;
    Boonchukusol, Pimlak
    ;
    Kuroki, Yoshimitsu
    ;
    Kato, Yasushi
    Thai text detection from natural scenes is still a challenging task for language translation applications, since there are many unsolved issues. Furthermore, the existing related works cannot completely detect Thai text. The main reason is that Thai text layout has vowels and tonal marks that differ from other languages. This paper proposes an approach to detect Thai text from natural scenes. The approach consists of two main procedures. (i) Fast boundary clustering algorithm decomposes scene features into multilayers, so that it is faster and easier to analyze Thai text characters. (ii) Modified connected component analysis method is applied to such scene features in order to detect Thai text boundaries. Based on 150 test images with 4,920 characters, the experimental results demonstrate that the proposed approach achieves the high average precision and recall, 0.80 and 0.90. © 2013 IEEE.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Vision Transformer with Fractal Dimension Transformation: Effects of Resolution and Patch Size
    (2025-01-01)
    Ngamkham, Woramat
    ;
    ;
    Kuroki, Yoshimitsu
    Vision Transformer (ViT) achieves strong performance in computer vision but requires substantial computational resources, particularly with high-resolution data. A key challenge lies in the quadratic complexity of self-attention with respect to the number of image patches, which is jointly determined by input size and patch size. Conventional resizing is a common strategy to reduce resolution and thus the number of patches, but it risks discarding structural details that may be important for prediction. To address this issue, this study investigates how input size, patch size, and dimensionality reduction influence ViT training time and prediction accuracy. Using the NIH Chest X-ray dataset, we compared two preprocessing methods: conventional resizing and a Fractal Dimension (FD)-based transformation. Results show that the FD-based method consistently reduced training time across all settings, demonstrating its effectiveness in lowering computational costs. In terms of accuracy, conventional resizing generally performed slightly better overall; however, the differences were not uniform, as smaller patches improved AUROC mainly at higher resolutions but not consistently at lower ones. These findings highlight a tradeoff between efficiency and accuracy, positioning FD-based representations as a practical complement to conventional resizing when computational resources are limited.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Distributed compressed video sensing with a pre-learned consensus convolutional dictionary
    (2026-02-27)
    Muta, Ibuki
    ;
    ;
    Kuroki, Yoshimitsu
    Distributed Compressed Video Sensing (DCVS) is a video coding framework well suited to low-power, low-complexity encoding environments. In conventional DCVS with Convolutional Sparse Representation (CSR), convolutional dictionary filters are learned from a key frame in each group of pictures (GOP), and the remaining non-key frames are reconstructed by solving a Convolutional Sparse Coding (CSC) problem using that dictionary. In this work, we investigate a CSR-based DCVS framework that instead employs a pre-learned convolutional dictionary trained offline on multiple video datasets via a consensus-based dictionary learning framework. Using this fixed dictionary, every frame in a sequence is reconstructed independently as if it were a key frame, i.e., without referencing other frames in the same sequence. We evaluate the proposed pre-learned-dictionary DCVS on the Foreman, Akiyo, and Coastguard sequences under two configurations that differ in the choice of data-fidelity term (L1 or L2) with symmetric boundary handling. Experimental results show that all test sequences can be successfully reconstructed using the pre-learned dictionary, indicating that sequence-specific key-frame-based dictionary learning at the decoder is not necessary. Moreover, the L1 data-fidelity term consistently yields better reconstruction quality than the L2 term in terms of PSNR and SSIM.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Distributed compressed video sensing with multiple key frames
    (2026-02-27)
    Nomaguchi, Mizuki
    ;
    Inoue, Ryota
    ;
    ;
    Kuroki, Yoshimitsu
    Distributed Compressed Video Sensing is a video compression method utilizing Compressed Sensing and Distributed Video Coding. With this method, compressed frames are reconstructed with information obtained by applying Convolutional Sparse Coding to a non-compressed frame. In this study, we aim to increase the reconstruction accuracy by selecting multiple non-compressed frames. In addition, we use symmetric convolution in order to solve a high computational optimization problem. The experimental results show our proposed method outperforms the conventional method.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    An improved 2DPCA for face recognition under illumination effects
    (2015-01-01) ;
    Sornnoi, Monmorakot
    ;
    Leelaburanapong, Savita
    ;
    ;
    Varakulsiripunth, Ruttikorn
    Principal component analysis (PCA) is one of the successful techniques for applying to face recognition, but its challenge still remains for solving an illumination effect condition. This paper proposes an improved 2DPCA (I-2DPCA) for overwhelming the illumination effect in face recognition. The proposed method is based on two assumptions. The first assumption is to create the covariance matrix that can effectively decompose the components of illumination effects from the eigenfaces. This avoids the illumination effect problem. The second assumption is to select the suitable eigenvectors that can significantly improve the recognition rate. Based on the Extended Yale Face Database B+ containing 60 illumination conditions, the experimental results show that not only does the proposed method decrease the computing time, but it also improves the recognition rate up to 95.93%.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    L1-L1 norm-based convolutional sparse coding via Anderson-accelerated Douglas-Rachford splitting
    (2026-02-27)
    Take, Hiroto
    ;
    Furusho, Riku
    ;
    ;
    Kuroki, Yoshimitsu
    Convolutional Sparse Coding (CSC) represents a signal through the convolution of dictionary filters and sparse coefficients. While the Alternating Direction Method of Multipliers (ADMM) has conventionally been used to solve CSC problems, recent studies have demonstrated that Douglas-Rachford (DR) splitting can achieve faster convergence. In this study, we propose an accelerated CSC algorithm by applying Anderson Acceleration to the DR splitting method. Experimental results demonstrate that the proposed method significantly improves convergence speed compared to standard DR splitting.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Local variance image-based for scene text binarization under illumination effects
    (2017-07-18)
    Peuwnuan, Kittipop
    ;
    ; ;
    Kuroki, Yoshimitsu
    Illumination effects, especially shadow and lighting condition, are grand challenges for scene-text localization. With these effects, text localization faces a difficult task to discriminate text regions from a nature scene due to edge and detail of characters affected by surrounding environments. To improve effectiveness of the scene-text localization, this paper proposes a local variance image technique to enhance character's edge for easily segmenting the scene text from the back-ground under illumination effects. In this method, the local variance image plays an important role in indicating how high complexity is in each local area. Then the proposed adaptive kernel size thresholding method is applied to such a local variance image to segment the scene text from the background. When the proposed method is tested with a Thai text dataset, the experimental results show the scene-text binarization is better than those of the state-of-the-art methods.
  • Some of the metrics are blocked by your 
    Item type:Publication,
    Texture analysis assessment for images
    (2017-02-23) ;
    Kaewaramsri, Yothin
    ;
    Suttapakti, Ungsumalee
    ;
    ;
    Kuroki, Yoshimitsu
    Commonly, the existing metrics such as mean square error (MSE), peak signal-To-noise ratio (PSNR), quality index (QI), structural similarity index metric (SSIM), and quality index based on local variance (QILV) use the image intensity-based statistics approach to assess the quality of distorted images. These metrics are successful in discriminating the quality of distorted images, such as de-noising, JPEG compressed, and blur images. However, they are unsuccessful in discriminating the quality of channel decomposition images. Therefore, this paper proposes the texture analysis assessment (TAA) to measure the quality of both normally distorted images and channel decomposition images. The proposed metric uses image intensity statistics in conjunction with texture analysis for quality discrimination of slightly different distorted and channel decomposition images. The texture analysis based on edge orientation is an important part employed to measure precise image errors. The experimental results illustrate that the TAA metric can evidently discriminate the quality of normally distorted images and channel decomposition images, when compared with state-of-The-Art metrics. Furthermore, the perceived visual quality and the quality value of TAA are corresponding; the lower visual quality human-eye perceives, the lower quality value TAA measures, and vice versa.