AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Beyond words: evaluating large language models in transportation planning

Shaowei Yinga,bZhenlong Lia ( )Manzhu Yua
Geoinformation and Big Data Research Laboratory, Department of Geography, The Pennsylvania State University, University Park, PA, USA
Office of the Chief Scientist, NCS Group, Singapore
Show Author Information

Abstract

The rapid advancement of Generative Artificial Intelligence (GenAI) in 2023 has catalyzed transformative shifts across various industries, including urban transportation planning. This study evaluates the applicability of Large Language Models (LLMs) in transportation decision-making, focusing on two hypotheses: (H1) out-of-the-box LLMs exhibit basic transportation knowledge and reasoning capabilities, enabling them to design and execute analytical workflows; and (H2) larger parameter models and fine-tuned models demonstrate superior accuracy and contextual understanding, outperforming smaller and general-purpose models. Using a three-level evaluation framework, we assessed GPT-4 and Phi-3-mini across (1) geospatial skills, (2) domain-specific transportation knowledge, and (3) real-world transport problem-solving in congestion pricing scenarios. Results confirm that while LLMs possess baseline geospatial and transportation reasoning abilities, their effectiveness varies by task complexity. GPT-4 outperformed Phi-3-mini across all evaluation levels, achieving 86% accuracy in GIS tasks, 81% in MATSim comprehension, and 91% in real-world transport decision support, while Phi-3-mini scored 43–72%. These findings highlight the advantages of larger models in structured decision-making tasks and their potential as analytical copilots for transportation planners. The study contributes to the ongoing scientific debate on the role of GenAI in transportation governance, reinforcing the need for fine-tuning and retrieval-augmented generation (RAG) to enhance LLM performance in structured analytics. Future research should explore newer LLMs, transport-specific fine-tuning, and hybrid AI architectures to improve AI-driven transportation planning and decision support.

References

【1】
【1】
 
 
Geo-Spatial Information Science
Pages 451-473

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Ying S, Li Z, Yu M. Beyond words: evaluating large language models in transportation planning. Geo-Spatial Information Science, 2026, 29(1): 451-473. https://doi.org/10.1080/10095020.2025.2493073

4

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 22 September 2024
Accepted: 09 April 2025
Published: 30 April 2025
© 2025 Wuhan University.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.