AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.6 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese | Open Access

Multimodal human-object interaction detection based on knowledge retrieval

Yan CHENYongBin GAO( )
School of Electronic and Electrical Engineering, Shanghai University of Engineering Sciences, Shanghai 201620, China
Show Author Information

Abstract

Human-object interation (HOI) plays a crucial role in understanding complex scenes. Recently, contrastive language image pre-training has shown great potential in providing prior knowledge about interactions in HOI detectors through knowledge extraction. However, this method typically relies on large-scale training data, and most methods directly map parameter interaction queries to a set of HOI predictions in a one-stage manner. This leads to a lack of sufficient exploration and utilization of the rich interaction structures. The use of multimodal data allows more dimensional information to be extracted and offers a more comprehensive understanding of the interaction behavior between human and object. In this study we designed a Transformer style HOI detector. This process scheme first retrieves comparative language image pre-training (CLIP) knowledge based on queries, then performs interactive suggestion generation, and finally converts non-parametric interactive suggestions into HOI predictions through a structure aware network. Structural awareness networks improve the accuracy of prediction results by encoding the overall semantic structure and local spatial structure with additional encoding. The accuracy of this model on the public dataset V-COCO reached 64.83%, and the accuracy on HICO-DET reached 28.78%. Compared with existing HOI detection algorithms, this algorithm exhibits superior performance, demonstrating its effectiveness.

CLC number: TP391.41

References

【1】
【1】
 
 
Journal of Beijing University of Chemical Technology (Natural Science Edition)
Pages 113-121

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
CHEN Y, GAO Y. Multimodal human-object interaction detection based on knowledge retrieval. Journal of Beijing University of Chemical Technology (Natural Science Edition), 2025, 52(1): 113-121. https://doi.org/10.13543/j.bhxbzr.2025.01.013

1

Views

0

Downloads

0

Crossref

0

Scopus

0

CSCD

Received: 10 November 2023
Published: 20 January 2025
© 2025 The Authors.

This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).