Applied Mathematics and Nonlinear Sciences
Journal license

Journal

Applied Mathematics and Nonlinear Sciences


Volume
& Issue

Volume 5, Issue 1


Published
on

January 20, 2020


Pages

1-10


DOI

Article

Improvement of the Fast Clustering Algorithm Improved by K-Means in the Big Data

Check for updates


Authors

Ting Xie Affiliation:
College of Science, Chongqing University of Technology, Chongqing 400054, China
, Ruihua Liu Affiliation:
College of Artificial Intelligence, Chongqing University of Technology, Chongqing 400054, China
and Zhengyuan Wei Affiliation:
College of Science, Chongqing University of Technology, Chongqing 400054, China


Abstract

Clustering as a fundamental unsupervised learning is considered an important method of data analysis, and K-means is demonstrably the most popular clustering algorithm. In this paper, we consider clustering on feature space to solve the low efficiency caused in the Big Data clustering by K-means. Different from the traditional methods, the algorithm guaranteed the consistency of the clustering accuracy before and after descending dimension, accelerated K-means when the clustering centeres and distance functions satisfy certain conditions, completely matched in the preprocessing step and clustering step, and improved the efficiency and accuracy. Experimental results have demonstrated the effectiveness of the proposed algorithm.


Keywords

Big Data, Clustering, K-means, Feature space, 62K86


Citation

Xie, T., Liu, R., & Wei, Z. (2020). Improvement of the fast clustering algorithm improved by k-means in the big data. Applied Mathematics and Nonlinear Sciences, 5(1), 1–10. https://doi.org/10.2478/amns.2020.1.00001

Published by: Engineering Journals

Engineering Journals Logo