Classification of file duplication by hierarchical clustering based on similarity relations

dc.contributor.authorPhankokkruad, Manop
dc.date.accessioned2026-08-06T10:20:13Z
dc.date.available2026-08-06T10:20:13Z
dc.date.issued2018-06-21
dc.description.abstractThis paper have proposed the classification of the duplicate file by measuring the similarity score between the couple of files. This work examined the distance between the pairwise of files by the Smith-Waterman algorithm. In addition, the make use of the Euclidean distance matrix could identify the relativity between the persons who often copies the files each other. Since the regularity of the duplication happens, this work could classify the proximity to the persons, and a group of person who positioned closely together by applying the hierarchical clustering. The result revealed that the Smith-Waterman algorithms could measure the similarity between files effectively. Also, this work could analyze the relativity of the persons, classifies the person who positioned closely together, and the person between nearest related members of the group. Finally, this work represented the amount of time that person duplicated the files.
dc.identifier.citationIcnc Fskd 2017 13th International Conference on Natural Computation Fuzzy Systems and Knowledge Discovery, 1598-1603, 2018
dc.identifier.doi10.1109/FSKD.2017.8393004
dc.identifier.other2-s2.0-85050248777
dc.identifier.urihttps://dspace.kmitl.ac.th/handle/123456789/8648
dc.sourceIcnc Fskd 2017 13th International Conference on Natural Computation Fuzzy Systems and Knowledge Discovery
dc.subjectAlgorithms
dc.subjectClustering
dc.subjectDetection
dc.subjectDistance
dc.subjectDuplication
dc.subjectEuclidean
dc.subjectHierarchical
dc.subjectProgramming
dc.subjectSimilarity
dc.titleClassification of file duplication by hierarchical clustering based on similarity relations
dc.typeConference Paper

Files

Collections