KMITL
Permanent URI for this communityhttps://dspace.kmitl.ac.th/handle/123456789/1
Browse
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, E-commerce web page classification based on automatic content extraction(2015-08-24) ;Petprasit, WaridJaiyen, SaichonCurrently, There are many E-commerce websites around the internet world. These E-commerce websites can be categorized into many types which one of them is C2C (Customer to Customer) websites such as eBay and Amazon. The main objective of C2C websites is an online market place that everyone can buy or sell anything at any time. Since, there are a lot of products in the E-commerce websites and each product are classified into its category by human. It is very hard to define their categories in automatic manner when the data is very large. In this paper, we propose the method for classifying E-commerce web pages based on their product types. Firstly, we apply the proposed automatic content extraction to extract the contents of E-commerce web pages. Then, we apply the automatic key word extraction to select words from these extracted contents for generating the feature vectors that represent the E-commerce web pages. Finally, we apply the machine learning technique for classifying the E-commerce web pages based on their product types. The experimental results signify that our proposed method can classify the E-commerce web pages in automatic fashion. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Web content extraction based on subject detection and node density(2015-02-27) ;Petprasit, WaridJaiyen, SaichonCurrently, very large data have been transferred from everywhere through World Wide Web. Consequently, the information extraction systems have been arising and many researches have been focusing on those data for utilizing them. These systems are very useful for data pre-processing and cleaning for real-time applications. Moreover, these systems can make other analyzing systems to analyze the data in real time such as social network mining, web mining, data mining, or even special tasks such as false advertisement detection, demand forecasting, and comment extraction on product and service reviews. In this paper, we focus on extracting the content data of web pages in e-commerce web sites based on subject detection and node density. In the experimental results, it can signify that our proposed method is appropriated to extract the data rich region in data-intensive pages in an automatic fashion.
