AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Article | Open Access

Automatic construction of global cloud sample database based on Landsat imagery

Tao HeaGuihua HuangbLei Zhanga( )Daiqiang WucYichuan Maa
Hubei Key Laboratory of Quantitative Remote Sensing of Land and Atmosphere, School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China
Beijing Dajia Internet Information Technology Co. Ltd., Shenzhen Branch, Shenzhen, China
Beijing DiDi Infinity Technology and Development Co, Ltd, Beijing, China
Show Author Information

Abstract

Accurately identifying clouds in imagery is a key preprocessing step in satellite remote sensing applications. In order to address problems of cloud samples required for cloud detection, such as reliance on manual interpretation, high collection costs, and limited quantity, this paper proposed a universal and automatic construction algorithm for cloud sample database (UAC-CSD) to create a large database of accurate and representative global cloud samples. The UAC-CSD sample database was applied and validated against Landsat 8 biome (L8_Biome) and spatial procedures for automated removal of cloud and shadow (SPARCS) datasets. The overall accuracies of random forest (RF) based on UAC-CSD sample database for cloud masking reached 0.921 on L8_Biome and 0.900 on SPARCS, which showed higher accuracies than RF based on Fmask4.0 sample database (0.869 and 0.866). Moreover, the UAC-CSD sample database was freely shared and could support the training and verification of various machine learning models. In addition to RF, light gradient boosting machine (LightGBM), multilayer perceptron (MLP), and support vector machine (SVM) were also trained based on UAC-CSD sample database and verified using L8_Biome. The results showed that the overall accuracies of models trained on UAC-CSD sample database were basically higher than models trained on Fmask4.0, and computational efficiencies were also greatly improved. This study provides global cloud samples with different cloud characteristics and covering different land covers for cloud detection models, which increases generalization ability of models and improves the accuracy and efficiency of processing massive medium to fine resolution satellite data in the big data era.

References

【1】
【1】
 
 
Geo-Spatial Information Science
Pages 59-77

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
He T, Huang G, Zhang L, et al. Automatic construction of global cloud sample database based on Landsat imagery. Geo-Spatial Information Science, 2026, 29(1): 59-77. https://doi.org/10.1080/10095020.2025.2514180

0

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 14 January 2025
Accepted: 28 May 2025
Published: 19 June 2025
© 2025 Wuhan University.

This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. The terms on which this article has been published allow the posting of the Accepted Manuscript in a repository by the author(s) or with their consent.