Publications
Sort:
Open Access Research Article Issue
Subspace and metric space learning for coqualitative data clustering
Electronic Research Archive 2026, 34(6): 4191-4215
Published: 19 May 2026
Abstract PDF (5.7 MB) Collect
Downloads:6

Cluster analysis of unlabeled categorical data is crucial in a wide range of practical applications, such as medical diagnosis, financial risk assessment, and recommendation systems. Unlike numerical data residing in explicit Euclidean spaces, categorical data consists of qualitative values without inherent ordering, making the definition of object similarity a critical yet challenging determinant of clustering success. Conventional approaches typically rely on single, predefined metrics (e.g., Hamming distance or context-based measures). However, these metrics are often constructed based on limited prior knowledge or specific statistical assumptions, failing to capture the complex, intrinsic structures of diverse datasets. Consequently, the mismatch between the defined metric space and the "true" data structure significantly hinders the performance of downstream clustering tasks. To address these limitations, this paper proposes a novel subspace and metric space co-learning framework named SBMS. Instead of relying on a static measure, SBMS introduces an adaptive learning paradigm that iteratively optimizes two coupled spaces: a metric space, where multiple complementary distance metrics are fused to provide a comprehensive similarity measure; and an attribute subspace, where attribute weights are dynamically adjusted based on cluster discrimination and compactness to identify the most relevant features for each cluster. Furthermore, we provide a theoretical analysis of the proposed method, discussing its computational complexity and demonstrating the convergence properties of the optimization algorithm. Extensive experiments on real-world public datasets from various domains illustrate that SBMS effectively bridges the gap between defined and true metric spaces, yielding superior clustering accuracy and stability compared to state-of-the-art baselines.

Total 1