KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
5 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, n-LIPO: Framework for Diverse Cooperative Agent Generation Using Policy Compatibility(2025-01-01) ;Charakorn, Rujikorn ;Manoonpong, PoramateDilokthanakul, NatDiverse training partners in multiagent tasks are crucial for training a robust and adaptable cooperative agent. Prior methods often rely on state-action information to diversify partners’ behaviors, but this can lead to minor variations instead of diverse behaviors and solutions. We address this limitation by introducing a novel training objective based on “policy compatibility.” Our method learns diverse behaviors by encouraging agents within a team to be compatible with each other while being incompatible with agents from other teams. We theoretically prove that incompatible policies are inherently dissimilar, allowing us to use policy compatibility as a proxy for diversity. We call this method learning incompatible policies for n -player cooperative games (n-LIPO). We propose to further diversify individual policies by incorporating a mutual information objective using state-action information. We empirically demonstrate that n-LIPO effectively generates diverse joint policies in various two-player and multi-player cooperative environments. In a complex cooperative task, two-player multi-recipe Overcooked, we find that n-LIPO generates a population of behaviorally diverse partners. These populations are then used to train robust generalist agents that can generalize better than using baseline populations. Finally, we demonstrate that n-LIPO can be applied to a high-dimensional StarCraft multiagent challenge (SMAC) multiplayer cooperative environment to discover diverse winning strategies when only a single goal exists. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Diversity Is Not All You Need: Training A Robust Cooperative Agent Needs Specialist Partners(2024-01-01) ;Charakorn, Rujikorn ;Manoonpong, PoramateDilokthanakul, NatPartner diversity is known to be crucial for training a robust generalist cooperative agent. In this paper, we show that partner specialization, in addition to diversity, is crucial for the robustness of a downstream generalist agent. We propose a principled method for quantifying both the diversity and specialization of a partner population based on the concept of mutual information. Then, we observe that the recently proposed cross-play minimization (XP-min) technique produces diverse and specialized partners. However, the generated partners are overfit, reducing their usefulness as training partners. To address this, we propose simple methods, based on reinforcement learning and supervised learning, for extracting the diverse and specialized behaviors of XP-min generated partners but not their overfitness. We demonstrate empirically that the proposed method effectively removes overfitness, and extracted populations produce more robust generalist agents compared to the source XP-min populations. This result highlights the importance of considering both the diversity and specialization of training partners while carefully managing their overfitness for training robust cooperative generalists. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, BAMS: Binary Sequence-Augmented Spectrogram with Self-Attention Deep Learning for Human Activity Recognition(2024-01-01) ;Sricom, Natchaya ;Charakorn, Rujikorn ;Manoonpong, PoramateLimpiti, TulayaHuman Activity Recognition (HAR) has rapidly gained interest over the years due to its wide range of applications in AI-based systems, particularly healthcare monitoring. HAR methods typically involve extracting relevant features from data provided by wearable sensors, smartphone sensors, cameras, or their combinations to classify different activities. Nevertheless, a major challenge lies in achieving high classification accuracy with limited data samples, particularly when distinguishing between activities with similar signal attributes. To address this challenge, we propose a novel HAR method called BinAry sequence-augmented spectrograM with Self-attention deep learning (BAMS). Our proposed method leverages only basic wearable sensor data. It utilizes short-time Fourier transform spectrograms to extract spatio-temporal sensor information. The spectrogram is integrated with a binary sequence that captures movement direction. We integrate a scaled dot-product self-attention mechanism into the model to prioritize data from wearable sensors, thereby enhancing the model's performance. The proposed method is evaluated on a public dataset using leave-one-subject-out cross-validation for efficacy and robustness. The method is found to achieve significant improvement over other state-of-the-art methods with the classification accuracy percentage and weighted F-1 scores of 88.06±5.11 and 87.36±5.96, respectively, for a twelve-activity classification. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Hybrid learning mechanisms under a neural control network for various walking speed generation of a quadruped robot(2023-10-01) ;Zhang, Yanbin ;Thor, Mathias ;Dilokthanakul, Nat ;Dai, ZhendongManoonpong, PoramateLegged robots that can instantly change motor patterns at different walking speeds are useful and can accomplish various tasks efficiently. However, state-of-the-art control methods either are difficult to develop or require long training times. In this study, we present a comprehensible neural control framework to integrate probability-based black-box optimization (PI<sup>BB</sup>) and supervised learning for robot motor pattern generation at various walking speeds. The control framework structure is based on a combination of a central pattern generator (CPG), a radial basis function (RBF) -based premotor network and a hypernetwork, resulting in a so-called neural CPG-RBF-hyper control network. First, the CPG-driven RBF network, acting as a complex motor pattern generator, was trained to learn policies (multiple motor patterns) for different speeds using PI<sup>BB</sup>. We also introduce an incremental learning strategy to avoid local optima. Second, the hypernetwork, which acts as a task/behavior to control parameter mapping, was trained using supervised learning. It creates a mapping between the internal CPG frequency (reflecting the walking speed) and motor behavior. This map represents the prior knowledge of the robot, which contains the optimal motor joint patterns at various CPG frequencies. Finally, when a user-defined robot walking frequency or speed is provided, the hypernetwork generates the corresponding policy for the CPG-RBF network. The result is a versatile locomotion controller which enables a quadruped robot to perform stable and robust walking at different speeds without sensory feedback. The policy of the controller was trained in the simulation (less than 1 h) and capable of transferring to a real robot. The generalization ability of the controller was demonstrated by testing the CPG frequencies that were not encountered during training. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, GENERATING DIVERSE COOPERATIVE AGENTS BY LEARNING INCOMPATIBLE POLICIES(2023-01-01) ;Charakorn, Rujikorn ;Manoonpong, PoramateDilokthanakul, NatTraining a robust cooperative agent requires diverse partner agents. However, obtaining those agents is difficult. Previous works aim to learn diverse behaviors by changing the state-action distribution of agents. But, without information about the task's goal, the diversified agents are not guided to find other important, albeit sub-optimal, solutions: the agents might learn only variations of the same solution. In this work, we propose to learn diverse behaviors via policy compatibility. Conceptually, policy compatibility measures whether policies of interest can coordinate effectively. We theoretically show that incompatible policies are not similar. Thus, policy compatibility-which has been used exclusively as a measure of robustness-can be used as a proxy for learning diverse behaviors. Then, we incorporate the proposed objective into a population-based training scheme to allow concurrent training of multiple agents. Additionally, we use state-action information to induce local variations of each policy. Empirically, the proposed method consistently discovers more solutions than baseline methods across various multi-goal cooperative environments. Finally, in multi-recipe Overcooked, we show that our method produces populations of behaviorally diverse agents, which enables generalist agents trained with such a population to be more robust.
