AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (970.9 KB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Open Access

TabMAGD: Combining Masked Autoencoders with Guided Diffusion for Mixed-Type Tabular Data Synthesis

College of Computer Science and Artificial Intelligence, Fudan University, Shanghai 200438, China
Shandong Artificial Intelligence Institute, Qilu University of Technology (Shandong Academy of Science), Jinan 250353, China
Show Author Information

Abstract

To address the scarcity of real data and data privacy concerns, significant research efforts have been directed toward developing tabular data synthesizers. Tabular data exhibit intrinsic heterogeneity with inter-column correlations, where features may represent diverse information types and contain mixed data formats—including both categorical and numerical values. To overcome challenges in modeling heterogeneous features and their complex dependencies, we propose a novel Tabular data synthesis framework combining Masked Autoencoder with Guided Diffusion (TabMAGD). Unlike existing tabular data generators, our method focuses on the correlation between categorical and numerical features, which learns the impact of numerical features on categorical features through masked autoencoders and captures the influence of categorical features on numerical features through guided diffusion. In our experiments, we conduct a comprehensive evaluation of TabMAGD, demonstrating its state-of-the-art performance against existing generative models. We evaluate TabMAGD’s performance in privacy-sensitive applications and observe that it consistently generates high-quality synthetic data while maintaining an outstanding trade-off between data utility and privacy preservation. Notably, tabular data synthesis can be regarded as a tool for data governance, assisting organizations in generating high-quality, compliant, and secure synthetic data.

References

【1】
【1】
 
 
Big Data Mining and Analytics
Pages 1264-1275

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Qian Y, Jing Y, He Z, et al. TabMAGD: Combining Masked Autoencoders with Guided Diffusion for Mixed-Type Tabular Data Synthesis. Big Data Mining and Analytics, 2026, 9(5): 1264-1275. https://doi.org/10.26599/BDMA.2025.9020105

685

Views

78

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 16 January 2025
Revised: 01 July 2025
Accepted: 24 September 2025
Published: 20 August 2026
© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).