Publications
Sort:
Open Access Issue
Multimodal human-object interaction detection based on knowledge retrieval
Journal of Beijing University of Chemical Technology (Natural Science Edition) 2025, 52(1): 113-121
Published: 20 January 2025
Abstract PDF (1.6 MB) Collect
Downloads:0

Human-object interation (HOI) plays a crucial role in understanding complex scenes. Recently, contrastive language image pre-training has shown great potential in providing prior knowledge about interactions in HOI detectors through knowledge extraction. However, this method typically relies on large-scale training data, and most methods directly map parameter interaction queries to a set of HOI predictions in a one-stage manner. This leads to a lack of sufficient exploration and utilization of the rich interaction structures. The use of multimodal data allows more dimensional information to be extracted and offers a more comprehensive understanding of the interaction behavior between human and object. In this study we designed a Transformer style HOI detector. This process scheme first retrieves comparative language image pre-training (CLIP) knowledge based on queries, then performs interactive suggestion generation, and finally converts non-parametric interactive suggestions into HOI predictions through a structure aware network. Structural awareness networks improve the accuracy of prediction results by encoding the overall semantic structure and local spatial structure with additional encoding. The accuracy of this model on the public dataset V-COCO reached 64.83%, and the accuracy on HICO-DET reached 28.78%. Compared with existing HOI detection algorithms, this algorithm exhibits superior performance, demonstrating its effectiveness.

Total 1