KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
4 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Item, Improving the state-of-the-art in Thai semantic similarity using distributional semantics and ontological information(2021-02-01) ;Netisopakul, Ponrudee ;Wohlgenannt, Gerhard ;Pulich, AlekseiHlaing, Zar ZarResearch into semantic similarity has a long history in lexical semantics, and it has applications in many natural language processing (NLP) tasks like word sense disambiguation or machine translation. The task of calculating semantic similarity is usually presented in the form of datasets which contain word pairs and a human-assigned similarity score. Algorithms are then evaluated by their ability to approximate the gold standard similarity scores. Many such datasets, with different characteristics, have been created for English language. Recently, four of those were transformed to Thai language versions, namely WordSim-353, SimLex-999, SemEval-2017-500, and R&G-65. Given those four datasets, in this work we aim to improve the previous baseline evaluations for Thai semantic similarity and solve challenges of unsegmented Asian languages (particularly the high fraction of out-of-vocabulary (OOV) dataset terms). To this end we apply and integrate different strategies to compute similarity, including traditional word-level embeddings, subword-unit embeddings, and ontological or hybrid sources like WordNet and ConceptNet. With our best model, which combines self-trained fastText subword embeddings with ConceptNet Numberbatch, we managed to raise the state-of-the-art, measured with the harmonic mean of Pearson on Spearman ρ, by a large margin from 0.356 to 0.688 for TH-WordSim-353, from 0.286 to 0.769 for TH-SemEval-500, from 0.397 to 0.717 for TH-SimLex-999, and from 0.505 to 0.901 for TWS-65. - Some of the metrics are blocked by yourconsent settings
Item type:Item, Word Similarity Datasets for Thai: Construction and Evaluation(2019-01-01) ;Netisopakul, Ponrudee ;Wohlgenannt, GerhardPulich, AlekseiDistributional semantics in the form of word embeddings are an essential ingredient to many modern natural language processing systems. The quantification of semantic similarity between words can be used to evaluate the ability of a system to perform semantic interpretation. To this end, a number of word similarity datasets have been created for the English language over the last decades. For Thai language few such resources are available. In this work, we create three Thai word similarity datasets by translating and re-rating the popular WordSim-353, SimLex-999 and SemEval-2017-Task-2 datasets. The three datasets contain 1852 word pairs in total and have different characteristics in terms of difficulty, domain coverage, and notion of similarity (relatedness vs. similarity). These features help to gain a broader picture of the properties of an evaluated word embedding model. We include baseline evaluations with existing Thai embedding models, and identify the high ratio of out-of-vocabulary words as one of the biggest challenges in the evaluation process. All datasets, evaluation results, and a tool for easy evaluation of new Thai embedding models are available to the NLP community online. - Some of the metrics are blocked by yourconsent settings
Item type:Item, A survey of Thai knowledge extraction for the semantic web research and tools(2018-04-01) ;Netisopakul, PonrudeeWohlgenannt, GerhardAs the manual creation of domain models and also of linked data is very costly, the extraction of knowledge from structured and unstructured data has been one of the central research areas in the Semantic Web field in the last two decades. Here, we look specifically at the extraction of formalized knowledge from natural language text, which is the most abundant source of human knowledge available. There are many tools on hand for information and knowledge extraction for English natural language, for written Thai language the situation is different. The goal of this work is to assess the state-of-The-Art of research on formal knowledge extraction specifically from Thai language text, and then give suggestions and practical research ideas on how to improve the state-of-The-Art. To address the goal, first we distinguish nine knowledge extraction for the SemanticWeb tasks defined in literature on knowledge extraction from English text, for example taxonomy extraction, relation extraction, or named entity recognition. For each of the nine tasks, we analyze the publications and tools available for Thai text in the form of a comprehensive literature survey. Additionally to our assessment, we measure the self-Assessment by the Thai research community with the help of a questionnaire-based survey on each of the tasks. Furthermore, the structure and size of the Thai community is analyzed using complex literature database queries. Combining all the collected information we finally identify research gaps in knowledge extraction from Thai language. An extensive list of practical research ideas is presented, focusing on concrete suggestions for every knowledge extraction task - which can be implemented and evaluated with reasonable effort. Besides the task-specific hints for improvements of the state-of-The-Art, we also include general recommendations on how to raise the efficiency of the respective research community. - Some of the metrics are blocked by yourconsent settings
Item type:Item, The State of Knowledge Extraction from Text for Thai Language(2017-11-15) ;Netisopakul, PonrudeeWohlgenannt, GerhardWith the emergence of the Semantic Web (or Linked Data), increased efforts have been made to automatically extract formalized semantic knowledge from natural language text. Most research work and tools for knowledge extraction are focusing on text in English language. In this work, wepresent our research-in-progress on evaluating the state-of-the-art in knowledge extraction from text for Thai language. For this purpose, we investigate the existing knowledge extraction literature and group the available research work and tools into eight knowledge extraction tasks. Our preliminary results from the survey of the state-of-the-art show that there exist large research gaps and therefore future research opportunities in Thaiknowledge extraction, and we provide hints for available research directions.
