AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (3.4 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Research Article | Open Access

Can large language models capture human risk preferences? A cross-cultural study

Jianing Liu1,JBing Song2,JVinayak Dixit3Chenyang Wu4,5( )Sisi Jian2( )
School of Economics and Management, Southwest Jiaotong University, Chengdu 611756, China
Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology, Hong Kong 999077, China
School of Civil and Environmental Engineering, The University of New South Wales, Sydney 2052, Australia
School of Aeronautics, Northwestern Polytechnical University, Xi’an 710129, China
National Key Laboratory of Aircraft Configuration Design, Xi’an 710129, China

Jianing Liu and Bing Song contributed equally to this work.

Show Author Information

Highlights

• LLMs do not faithfully reproduce human risk preferences.

• Model families show systematic and distinct biases in risky choice.

• GPT models are more risk-averse, whereas Gemini is more risk-seeking.

• Prompt language shifts simulated risk attitudes across cultures.

Abstract

Large language models (LLMs) are increasingly used as agents to simulate human behavior, yet their fidelity in complex decision-making under uncertainty remains insufficiently understood. To address this gap, we developed a comparative framework that benchmarked LLM-simulated risk preferences against empirical human behavior. Using demographic profiles from surveys conducted in Sydney, Hong Kong, and Nanjing, we constructed role-playing prompts and evaluated three LLMs on abstract lottery-choice tasks. We adopted the classical constant relative risk aversion (CRRA) framework as a domain-neutral “standard ruler” to compare risk attitudes. The analysis yielded three main findings. First, off-the-shelf LLMs do not exhibit a universal risk profile: The two GPT models are more risk-averse than human benchmarks, whereas Gemini is more risk-seeking. Second, prompt language systematically affects simulated risk attitudes, with English-to-Chinese switching inducing a more conservative shift in most cases. Third, LLMs do not reliably reproduce the empirical heterogeneity of human risk preferences, tending either to generate overly concentrated distributions or unrealistically large dispersion. Taken together, these findings show that off-the-shelf LLMs remain vulnerable to model-family-specific miscalibration, language-sensitive distortions, and failures in distributional fidelity. Rigorous empirical calibration is therefore necessary before off-the-shelf LLMs can be reliably deployed in computational social science and choice modeling.

Graphical Abstract

References

【1】
【1】
 
 
Communications in Transportation Research
Article number: 9640025

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Liu J, Song B, Dixit V, et al. Can large language models capture human risk preferences? A cross-cultural study. Communications in Transportation Research, 2026, 6(2): 9640025. https://doi.org/10.26599/COMMTR.2026.9640025

1152

Views

178

Downloads

0

Crossref

0

Web of Science

0

Scopus

Received: 15 January 2026
Revised: 24 January 2026
Accepted: 12 May 2026
Published: 30 June 2026
© The Author(s) 2026.

This is an open access article under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0 http://creativecommons.org/licenses/by/4.0/).