Beyond Predefined Clusters: A Comprehensive Review of Clustering Methods for Unknown Cluster Numbers
IEEE Transactions on Knowledge and Data Engineering, vol.38, no.7, pp.4121-4138, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 38 Issue: 7
- Publication Date: 2026
- Doi Number: 10.1109/tkde.2026.3680286
- Journal Name: IEEE Transactions on Knowledge and Data Engineering
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC
- Page Numbers: pp.4121-4138
- Keywords: Automatic clustering, deep clustering, semi-supervised clustering, unknown number of clusters
- Dokuz Eylül University Affiliated: Yes
Abstract
Clustering is an unsupervised learning task that groups data points by their inherent similarities. Nonautomatic clustering algorithms face significant challenges when the true number of clusters is unknown or changes dynamically, as they require this number to be predefined. This paper provides a comprehensive review of automatic clustering algorithms specifically designed to handle such uncertainty. In this paper, these algorithms are systematically classified based on three key perspectives: clustering framework (classical vs. deep), clustering strategy (e.g., density-based, model based, graph-theoretic, subspace methods), and the use of labeled data (unsupervised vs. semi-supervised). We analyze each algorithm based on its core principles, key contributions, strengths, and limitations. Furthermore, we address the current challenges in this area and propose future research directions to enhance the scalability, robustness, and effectiveness of automatic clustering algorithms.