AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (2 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Open Access

Empowering Large Language Models for Relational Data Cleaning and Integration

School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China
Show Author Information

Abstract

Relational data, as one of the most common data forms, are widely applied in fields such as recommendation systems, medical diagnosis, and traffic prediction. Relational data formats often face issues pertaining to incompleteness and deficiencies in quality, necessitating rigorous cleansing and integration processes to unleash their potential. However, different downstream tasks and data characteristics impose various requirements on the cleansing and integration processes. Currently, there is an absence of unified and effective methods. Large Language Models (LLMs), with their extensive real-world knowledge and strong natural language processing abilities, are emerging as robust tools for cleansing and integrating relational data. This study investigates techniques for tackling relational data cleaning and integration challenges, and proposes a comprehensive framework based on LLMs. The framework uses LLMs for problem-solving, integrating multi-granularity table representation models and alignment networks to improve task-related table understanding. Moreover, the framework facilitates straightforward deployment across various LLMs. To assess the efficacy of the LLMs and the proposed framework, this study performs experiments on the task of imputing missing values utilizing four real-world datasets. The experiments evaluate the performance of a proprietary LLM, an open-source LLM that has been fine-tuned, and the framework itself, comparing and analyzing their results with existing techniques. The experiments demonstrate that LLMs significantly outperform traditional methods in these tasks. Furthermore, ablation experiments are conducted to verify the effectiveness of the framework.

References

【1】
【1】
 
 
Big Data Mining and Analytics
Pages 1308-1327

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Shi X, Chung GJ, Julian D, et al. Empowering Large Language Models for Relational Data Cleaning and Integration. Big Data Mining and Analytics, 2026, 9(5): 1308-1327. https://doi.org/10.26599/BDMA.2025.9020073

8

Views

0

Downloads

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 27 January 2025
Revised: 16 April 2025
Accepted: 13 June 2025
Published: 20 August 2026
© The author(s) 2026.

The articles published in this open access journal are distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/).