Clustering methods for high-dimensional data

Clustering methods for high-dimensional data
Citations

WEB OF SCIENCE

0

초록

High-dimensional data which is characterized by observations with thousands of features are prevalent in various fields such as gene expression data analysis, image processing, and natural language processing, etc. Despite of offering rich information, such data present substantial challenges for clustering due to the curse of dimensionality, degradation of similarity measures, and lack of interpretability. In response, a wide array of methodologies has been developed so far, including subspace clustering, dimension reduction, model-based approaches, and feature selection, etc. More recent advances incorporate regularization techniques and bootstrap-based procedures for estimating the number of clusters. This paper provides a comprehensive overview of contemporary clustering methods tailored for high-dimensional data, critically evaluating their strengths, weaknesses, and applicability. Furthermore, it outlines promising research directions aimed at developing efficient, robust, scalable and interpretable clustering algorithms.

키워드

cluster analysisdimension reductionfeature selectionhigh-dimensional datamodel-based clusteringsubspace clusteringDISCRIMINANT-ANALYSISFEATURE-SELECTIONMODELVARIABLES
제목
Clustering methods for high-dimensional data
제목 (타언어)
Clustering methods for high-dimensional data
저자
Lee, EunhaKim, JooyoungKim, Jaejik
DOI
10.5351/KJAS.2025.38.5.693
발행일
2025-10
유형
Article
저널명
응용통계연구
38
5
페이지
693 ~ 719