AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

A Heuristic Sampling Method for Maintaining the Probability Distribution

Jiao-Yun Yang1,2,3Jun-Da Wang1,2,4( )Yi-Fang Zhang1,2,3Wen-Juan Cheng1,2,3Lian Li1,2,3
Key Laboratory of Knowledge Engineering with Big Data of Ministry of Education, Hefei University of Technology Hefei 230601, China
National Smart Eldercare International Science and Technology Cooperation Base, Hefei University of Technology Hefei 230601, China
School of Computer Science and Information Engineering, Hefei University of Technology, Hefei 230601, China
School of Mathematics, Hefei University of Technology, Hefei 230601, China

Recommended by NCTCS 2019

Show Author Information

Abstract

Sampling is a fundamental method for generating data subsets. As many data analysis methods are developed based on probability distributions, maintaining distributions when sampling can help to ensure good data analysis performance. However, sampling a minimum subset while maintaining probability distributions is still a problem. In this paper, we decompose a joint probability distribution into a product of conditional probabilities based on Bayesian networks and use the chi-square test to formulate a sampling problem that requires that the sampled subset pass the distribution test to ensure the distribution. Furthermore, a heuristic sampling algorithm is proposed to generate the required subset by designing two scoring functions: one based on the chi-square test and the other based on likelihood functions. Experiments on four types of datasets with a size of 60000 show that when the significant difference level, α, is set to 0:05, the algorithm can exclude 99:9%, 99:0%, 93:1% and 96:7% of the samples based on their Bayesian networks—ASIA, ALARM, HEPAR2, and ANDES, respectively. When subsets of the same size are sampled, the subset generated by our algorithm passes all the distribution tests and the average distribution difference is approximately 0:03; by contrast, the subsets generated by random sampling pass only 83:8% of the tests, and the average distribution difference is approximately 0:24.

Electronic Supplementary Material

Download File(s)
jcst-36-4-896-Highlights.pdf (235.2 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 896-909

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Yang J-Y, Wang J-D, Zhang Y-F, et al. A Heuristic Sampling Method for Maintaining the Probability Distribution. Journal of Computer Science and Technology, 2021, 36(4): 896-909. https://doi.org/10.1007/s11390-020-0065-6

1163

Views

10

Crossref

7

Web of Science

11

Scopus

0

CSCD

Received: 05 October 2019
Accepted: 15 August 2020
Published: 05 July 2021
©Institute of Computing Technology, Chinese Academy of Sciences 2021