Large language models (LLMs) are increasingly used as agents to simulate human behavior, yet their fidelity in complex decision-making under uncertainty remains insufficiently understood. To address this gap, we developed a comparative framework that benchmarked LLM-simulated risk preferences against empirical human behavior. Using demographic profiles from surveys conducted in Sydney, Hong Kong, and Nanjing, we constructed role-playing prompts and evaluated three LLMs on abstract lottery-choice tasks. We adopted the classical constant relative risk aversion (CRRA) framework as a domain-neutral “standard ruler” to compare risk attitudes. The analysis yielded three main findings. First, off-the-shelf LLMs do not exhibit a universal risk profile: The two GPT models are more risk-averse than human benchmarks, whereas Gemini is more risk-seeking. Second, prompt language systematically affects simulated risk attitudes, with English-to-Chinese switching inducing a more conservative shift in most cases. Third, LLMs do not reliably reproduce the empirical heterogeneity of human risk preferences, tending either to generate overly concentrated distributions or unrealistically large dispersion. Taken together, these findings show that off-the-shelf LLMs remain vulnerable to model-family-specific miscalibration, language-sensitive distortions, and failures in distributional fidelity. Rigorous empirical calibration is therefore necessary before off-the-shelf LLMs can be reliably deployed in computational social science and choice modeling.
Publications
- Article type
- Year
- Co-author
Article type
Year
Open Access
Research Article
Issue
Communications in Transportation Research 2026, 6(2): 9640025
Published: 30 June 2026
Downloads:178
Total 1
京公网安备11010802044758号