Loading...
Longest Matching and Rule-based Techniques for Khmer Word Segmentation
Author(s)
Long, Pakrigna
Boonjing, Veera
Date Issued
August 6, 2018
Type
Conference Paper
Abstract
Word boundaries are the essential assignment to be done in natural language processing research. In most Asian languages, as well as Khmer language, many studies involved with word segmentation have been investigated. In Khmer Word Segmentation, several approaches related to segmenting words based on dictionary have been studied. There are only few researches about solving unknown word problem. This matter is a quite challenge task in word separation. In this research, Maximum Matching algorithm (MMA) together with Rule-based technique has been proposed. First, MMA and a Khmer manual corpus were used to make word boundaries in each sentence. Then the unknown words were then defined and solved by using 21 grammar rules created. We tested the segmentation with 2018 sentences from agriculture, magazine, newspaper, technology, health and history. With Maximum Matching alone, we could achieve the accuracy of 88.55% and along with Rule-based, the accuracy increased to 92.81%.
Citation
2018 10th International Conference on Knowledge and Smart Technology Cybernetics in the Next Decades Kst 2018, 80-83, 2018
