Discover the SciOpen Platform and Achieve Your Research Goals with Ease.
Search articles, authors, keywords, DOl and etc.
Crowd profiling of massive transit data is valuable for analyzing the travel characteristics and traffic trends of urban groups, but the processing of the data is time-consuming, low-quality and difficult to interpret. A systematic solution for crowd profiling of massive public transport data was proposed. Based on the PageRank algorithm, the trajectories of people passing through important stations were filtered out, which greatly reduced the trajectory data of the target population. A textual analysis method for trajectories was proposed to improve the interpretability of crowd profiling. And the K-means algorithm based on cosine distance as the clustering algorithm for crowd profiling was analysed and determined. The experiments on 30 million passengers′ transit data show that the proposed algorithm can solve the problem of crowd profiling in massive transit data in a more systematic way, while the K-means algorithm based on cosine distance has the best clustering effect and the accuracy rate is about 80%. The crowd profiling and its trajectory were visually displayed by using Flow Map, and the results are consistent with real-world crowd behavioural characteristics.
This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Comments on this article