KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
4 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Item, Distributed compressed video sensing with multiple key frames(2026-02-27) ;Nomaguchi, Mizuki ;Inoue, Ryota ;Woraratpanya, KuntpongKuroki, YoshimitsuDistributed Compressed Video Sensing is a video compression method utilizing Compressed Sensing and Distributed Video Coding. With this method, compressed frames are reconstructed with information obtained by applying Convolutional Sparse Coding to a non-compressed frame. In this study, we aim to increase the reconstruction accuracy by selecting multiple non-compressed frames. In addition, we use symmetric convolution in order to solve a high computational optimization problem. The experimental results show our proposed method outperforms the conventional method. - Some of the metrics are blocked by yourconsent settings
Item type:Item, L1-L1 norm-based convolutional sparse coding via Anderson-accelerated Douglas-Rachford splitting(2026-02-27) ;Take, Hiroto ;Furusho, Riku ;Woraratpanya, KuntpongKuroki, YoshimitsuConvolutional Sparse Coding (CSC) represents a signal through the convolution of dictionary filters and sparse coefficients. While the Alternating Direction Method of Multipliers (ADMM) has conventionally been used to solve CSC problems, recent studies have demonstrated that Douglas-Rachford (DR) splitting can achieve faster convergence. In this study, we propose an accelerated CSC algorithm by applying Anderson Acceleration to the DR splitting method. Experimental results demonstrate that the proposed method significantly improves convergence speed compared to standard DR splitting. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Distributed compressed video sensing with a pre-learned consensus convolutional dictionary(2026-02-27) ;Muta, Ibuki ;Weraratpanya, KuntpongKuroki, YoshimitsuDistributed Compressed Video Sensing (DCVS) is a video coding framework well suited to low-power, low-complexity encoding environments. In conventional DCVS with Convolutional Sparse Representation (CSR), convolutional dictionary filters are learned from a key frame in each group of pictures (GOP), and the remaining non-key frames are reconstructed by solving a Convolutional Sparse Coding (CSC) problem using that dictionary. In this work, we investigate a CSR-based DCVS framework that instead employs a pre-learned convolutional dictionary trained offline on multiple video datasets via a consensus-based dictionary learning framework. Using this fixed dictionary, every frame in a sequence is reconstructed independently as if it were a key frame, i.e., without referencing other frames in the same sequence. We evaluate the proposed pre-learned-dictionary DCVS on the Foreman, Akiyo, and Coastguard sequences under two configurations that differ in the choice of data-fidelity term (L1 or L2) with symmetric boundary handling. Experimental results show that all test sequences can be successfully reconstructed using the pre-learned dictionary, indicating that sequence-specific key-frame-based dictionary learning at the decoder is not necessary. Moreover, the L1 data-fidelity term consistently yields better reconstruction quality than the L2 term in terms of PSNR and SSIM. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Vision Transformer with Fractal Dimension Transformation: Effects of Resolution and Patch Size(2025-01-01) ;Ngamkham, Woramat ;Woraratpanya, KuntpongKuroki, YoshimitsuVision Transformer (ViT) achieves strong performance in computer vision but requires substantial computational resources, particularly with high-resolution data. A key challenge lies in the quadratic complexity of self-attention with respect to the number of image patches, which is jointly determined by input size and patch size. Conventional resizing is a common strategy to reduce resolution and thus the number of patches, but it risks discarding structural details that may be important for prediction. To address this issue, this study investigates how input size, patch size, and dimensionality reduction influence ViT training time and prediction accuracy. Using the NIH Chest X-ray dataset, we compared two preprocessing methods: conventional resizing and a Fractal Dimension (FD)-based transformation. Results show that the FD-based method consistently reduced training time across all settings, demonstrating its effectiveness in lowering computational costs. In terms of accuracy, conventional resizing generally performed slightly better overall; however, the differences were not uniform, as smaller patches improved AUROC mainly at higher resolutions but not consistently at lower ones. These findings highlight a tradeoff between efficiency and accuracy, positioning FD-based representations as a practical complement to conventional resizing when computational resources are limited.
