Label scarcity and long-tailed distribution imbalance are significant challenges in industrial equipment monitoring. Currently, self-supervised learning methods are affected by sample quantity bias and semantic confusion under complex operating conditions, which limits their ability to represent sparse critical states. To address these issues, we propose a co-evolutionary prototypical contrastive learning (EPCL) framework. Through progressive learning from coarse-grained semantic discovery to fine-grained discriminative enhancement, this framework enables an in-depth analysis of the intrinsic structure of long-tailed data. Specifically, an adaptive prototype-based clustering algorithm based on optimal transport theory is introduced, thereby achieving unbiased representation learning through data-driven dynamic priors. Furthermore, a semantic-aware and hierarchical negative sample weighting scheme is designed to optimize discriminative boundaries while mitigating class imbalance by enforcing prototype consistency constraints and employing an adaptive weighting strategy. Extensive experiments were conducted on several public long-tailed visual benchmarks, including CIFAR10-LT, CIFAR100-LT, and ImageNet-100-LT, as well as the industrial fault diagnosis dataset. The results demonstrated that the EPCL achieved better performance than fifteen mainstream self-supervised methods (e.g., SimCLR and SwAV) in both linear evaluation and few-shot classification tasks. On the CIFAR100-LT dataset, the EPCL improved the tail-class accuracy by 4.56% compared to SimCLR. Ablation studies and visualization results verified the effectiveness and generalization ability of the framework. This work offers a promising insight and practical solution for representation learning from unlabeled long-tailed measurement data.
- Article type
- Year
Open Access
Issue
Open Access
Issue
Data collected in fields such as cybersecurity and biomedicine often encounter high dimensionality and class imbalance. To address the problem of low classification accuracy for minority class samples arising from numerous irrelevant and redundant features in high-dimensional imbalanced data, we proposed a novel feature selection method named AMF-SGSK based on adaptive multi-filter and subspace-based gaining sharing knowledge. Firstly, the balanced dataset was obtained by random under-sampling. Secondly, combining the feature importance score with the AUC score for each filter method, we proposed a concept called feature hardness to judge the importance of feature, which could adaptively select the essential features. Finally, the optimal feature subset was obtained by gaining sharing knowledge in multiple subspaces. This approach effectively achieved dimensionality reduction for high-dimensional imbalanced data. The experiment results on 30 benchmark imbalanced datasets showed that AMF-SGSK performed better than other eight commonly used algorithms including BGWO and IG-SSO in terms of F1-score, AUC, and G-mean. The mean values of F1-score, AUC, and G-mean for AMF-SGSK are 0.950, 0.967, and 0.965, respectively, achieving the highest among all algorithms. And the mean value of G-mean is higher than those of IG-PSO, ReliefF-GWO, and BGOA by 3.72%, 11.12%, and 20.06%, respectively. Furthermore, the selected feature ratio is below 0.01 across the selected ten datasets, further demonstrating the proposed method’s overall superiority over competing approaches. AMF-SGSK could adaptively remove irrelevant and redundant features and effectively improve the classification accuracy of high-dimensional imbalanced data, providing scientific and technological references for practical applications.
京公网安备11010802044758号