AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (3.5 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access | Just Accepted

Exoskeleton locomotion mode prediction in construction using GPT-4o: Zero-shot learning from vision and speech

Ehsan Ahmadia( )Asif Faisal Chowdhurya( )Chao Wanga( )Mahmood Jasimb

a Bert S. Turner Department of Construction Management, Louisiana State University, Baton Rouge, LA 70803, USA

b Division of Computer Science and Engineering, Louisiana State University, Baton Rouge, LA 70803, USA

Show Author Information

Abstract

Wearable exoskeletons enhance mobility and support in demanding tasks but face challenges in adapting to dynamic construction environments, particularly in predicting locomotion modes for tasks such as ladder climbing, stair navigation, low-space movement, and obstacle navigation. This study investigates the effectiveness of integrating speech and vision data for locomotion prediction while evaluating the generalization capability of large language models, specifically GPT-4o, through zero-shot learning compared to supervised fine-tuning. Using a multimodal framework with field-of-view frames and speech commands captured by smart glasses, we tested contrastive language-image pre-training (CLIP), ImageBind, and GPT-4o. Fine-tuned CLIP achieved an F1-score of 90.05% , yet GPT-4o’s zero-shot of 87.87% closely rivaled it, demonstrating strong adaptability to construction’s complex demands without task-specific training, while fine-tuned ImageBind trailed at 78.84%. This comparison underscores GPT-4o’s substantial potential to enable scalable exoskeleton control by leveraging multimodal comprehension of vision and speech data in dynamic construction settings.

References

【1】
【1】
 
 
Journal of Intelligent Construction

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ahmadi E, Chowdhury AF, Wang C, et al. Exoskeleton locomotion mode prediction in construction using GPT-4o: Zero-shot learning from vision and speech. Journal of Intelligent Construction, 2026, https://doi.org/10.26599/JIC.2026.9180128

452

Views

25

Downloads

0

Crossref

0

Scopus

Received: 24 November 2025
Revised: 14 March 2026
Accepted: 23 March 2026
Available online: 21 April 2026

© The Author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits use, distribution and reproduction in any medium, provided the original work is properly cited.