Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Humans use natural language to describe places, relying on mental processes to perceive, interpret, and communicate information about spatial relationships, environments, and navigation. Computational location-based systems strive to replicate this capability by enabling the retrieval of geographic locations from textual descriptions or queries. However, progress in this domain remains constrained by the limited availability of extensive, linguistically diverse textual datasets, which are essential for developing and evaluating robust geographic information retrieval methodologies. In this study, we conducted a review of existing geocoding datasets used for textual geolocation. Our objectives were to systematically compare these datasets, characterize their attributes, and assess their impact on retrieval performance as reported in the literature. A critical challenge we identified was the inconsistency in evaluation practices across studies, which complicates direct comparisons and underscores the need for standardized benchmarks. This review synthesizes the current landscape of geocoding datasets, offering insights for informed dataset development and fosters more consistent evaluation practices. By addressing the imperative of dataset standardization and availability, we aim to support the creation of more effective geographic information retrieval systems and establish a solid foundation for future research in this field.
This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.
Comments on this article