Robust motion similarity retrieval from monocular 2D pose sequences is challenged by body-scale variation, viewpoint inconsistency, translation drift, and temporal misalignment. Existing contrastive skeleton learning methods primarily address action recognition and rarely integrate explicit geometric canonicalization for retrieval-oriented metric learning. This paper proposes a spatial-temporal normalized contrastive embedding framework that unifies structured nuisance suppression with scalable similarity representation learning. A four-stage normalization pipeline—torso-scale normalization, pelvis-centered alignment, posture-axis alignment, and phase-synchronized temporal resampling—removes geometric and temporal distortions prior to embedding. The normalized sequences are encoded using an acausal dilated temporal convolutional network trained with a hybrid contrastive objective combining NT-Xent and semi-hard triplet loss, enabling both global separation and fine-grained stylistic discrimination. A prototype-based representation further supports interpretable amateur-to-professional style mapping. Experiments on a golf swing benchmark achieve a Top-1 accuracy of 91.3%, outperforming BiLSTM and Dynamic Time Warping baselines. The framework establishes an invariant and interpretable paradigm for motion similarity retrieval applicable to broader human movement analysis tasks.
Publications
- Article type
- Year
Article type
Year
Open Access
Article
Issue
Computers, Materials & Continua 2026, 88(3): 5
Published: 23 July 2026
Downloads:0
Total 1
京公网安备11010802044758号