Humans use natural language to describe places, relying on mental processes to perceive, interpret, and communicate information about spatial relationships, environments, and navigation. Computational location-based systems strive to replicate this capability by enabling the retrieval of geographic locations from textual descriptions or queries. However, progress in this domain remains constrained by the limited availability of extensive, linguistically diverse textual datasets, which are essential for developing and evaluating robust geographic information retrieval methodologies. In this study, we conducted a review of existing geocoding datasets used for textual geolocation. Our objectives were to systematically compare these datasets, characterize their attributes, and assess their impact on retrieval performance as reported in the literature. A critical challenge we identified was the inconsistency in evaluation practices across studies, which complicates direct comparisons and underscores the need for standardized benchmarks. This review synthesizes the current landscape of geocoding datasets, offering insights for informed dataset development and fosters more consistent evaluation practices. By addressing the imperative of dataset standardization and availability, we aim to support the creation of more effective geographic information retrieval systems and establish a solid foundation for future research in this field.
Publications
- Article type
- Year
Article type
Year
Open Access
Article
Issue
Geo-Spatial Information Science 2026, 29(4): 2529-2546
Published: 03 September 2025
Total 1
京公网安备11010802044758号