Using sparse sensors to capture human movements and driving virtual humans in computers to reproduce these movements is one of the key technologies in fields such as virtual reality, among which, the relevant calculations for reproducing movements must simultaneously satisfy motion constraints and physical constraints. Currently, nonlinear optimization methods are primarily employed to address the physical constraints in such computations. However, these methods suffer from several drawbacks, including long computation time, high computational complexity, and the necessity of designing dedicated optimizers. To address these issues, a method based on a deep neural network model for solving the physical constraints optimization in sparse sensor motion capture is proposed. Firstly, the model effectively combines the multi-modal feature fusion network and the multi-path refinement network to form a post-fusion multi-level structure, which serves as the basic model of this research. Secondly, through a progressive fusion method, back-projection connections between the layers of the basic model are established, enabling the iteration of the basic model. This allows the information fused at deeper layers to be utilized by the shallower layers. Thirdly, a loss function in the form of combined weighting that incorporates physical constraints is proposed, which is suitable for fine-tuning the deep model in the presence of both implicit and explicit physical constraints. The experimental results demonstrate that the proposed method not only exhibits good feasibility but also improves computational efficiency by approximately 20 percentage points. Compared with other commonly used optimization methods, the proposed method performs better in terms of four mainstream evaluation indicators. Additionally, the method yields favorable results when applied to other datasets, demonstrating its strong generalization capability. This method provides a new perspective for the formation of an end-to-end deep model for sparse sensor motion capture.
- Article type
- Year
- Co-author
Open Access
Issue
Open Access
Issue
Detecting and tracking basketball in videos are helpful for coaches to review gameplays. In video streams of games, the You Only Look Once v5 (YOLOv5) algorithm exhibits low discriminative ability between basketball and other small circular targets due to the small size of the basketball target. To address this issue, we propose several improvements based on the YOLOv5. Firstly, we replace the original C3 module with the VoVNet C3 (V-C3) module to address the problem of limited basketball features and validate the effectiveness of this enhancement through Kullback-Leibler divergence. Secondly, we introduce the Bridge Path Aggregation Network (BPANet) to replace the Path Aggregation Network (PANet) for better detection of small basketball targets in the scene. Thirdly, a classification penalty mechanism is constructed to reduce false alarms between basketball and similar targets. Lastly, we explore the influence of various parameters on the performance of the basketball detection algorithm to determine optimal parameter values and model structures. Experimental results demonstrate that the improved algorithm improves recognition accuracy by approximately 3% over the original YOLOv5 algorithm, with an average precision increase of about 2.4% on the COCO dataset, and reduces the algorithm's parameter size by about 5.3%. The proposed four enhancement strategies of this study based on the YOLOv5 algorithm improve the detection accuracy of basketball targets in videos while reducing model complexity, thereby offering a new approach for similar object detection tasks.
Open Access
Issue
The accuracy issues of player tracking in basketball game video streams due to occlusions and identity loss, which are crucial for reconstructing players' movement trajectories throughout the game, are addressed. To overcome the limitations of the DeepSORT algorithm in handling occlusions and maintaining identity consistency, three improvements are proposed. Firstly, the integration of Soft-NMS and Focal E-IoU loss functions enhances detection performance during occlusions. Secondly, an occlusion trajectory matching mechanism is introduced to reduce identity confusion when players occlude each other. Finally, a module for extracting and linking number plate features on players' jerseys is added to resolve identity changes when players re-enter the field of view. On main popular datasets, the proposed algorithm demonstrates significant improvements over the original DeepSORT with an 11.21 percentage points and 6.63 percentage points increase in the HOTA metric, a 15.56 percentage points and 10.11 percentage points improvement in the IDF1 metric, and a reduction of 36% and 27% in the number of identity changes. Compared with the competitive OCSORT algorithm, this method further enhances the HOTA metric by 2.39 percentage points and 0.98 percentage points, and the MOTA metric by 7.48 percentage points and 7.92 percentage points. The enhancements effectively improve the tracking precision of basketball players in video streams, particularly in dealing with challenges such as identity switches and occlusions, thereby laying a solid foundation for accurately reconstructing the full-game movement trajectories of athletes.
京公网安备11010802044758号