Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Human-object interation (HOI) plays a crucial role in understanding complex scenes. Recently, contrastive language image pre-training has shown great potential in providing prior knowledge about interactions in HOI detectors through knowledge extraction. However, this method typically relies on large-scale training data, and most methods directly map parameter interaction queries to a set of HOI predictions in a one-stage manner. This leads to a lack of sufficient exploration and utilization of the rich interaction structures. The use of multimodal data allows more dimensional information to be extracted and offers a more comprehensive understanding of the interaction behavior between human and object. In this study we designed a Transformer style HOI detector. This process scheme first retrieves comparative language image pre-training (CLIP) knowledge based on queries, then performs interactive suggestion generation, and finally converts non-parametric interactive suggestions into HOI predictions through a structure aware network. Structural awareness networks improve the accuracy of prediction results by encoding the overall semantic structure and local spatial structure with additional encoding. The accuracy of this model on the public dataset V-COCO reached 64.83%, and the accuracy on HICO-DET reached 28.78%. Compared with existing HOI detection algorithms, this algorithm exhibits superior performance, demonstrating its effectiveness.
This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Comments on this article