To address the challenges posed by multi-source, heterogeneous, and high-noise data in real-world scenarios, traditional data engineering approaches, which are primarily driven by manual effort and rule-based mechanisms, face significant bottlenecks in large-scale settings, including high costs, low efficiency, and limited adaptability. The continuous advancement of artificial intelligence, particularly large-scale models, provides new technical pathways and opportunities to overcome these limitations. By adopting a learning-driven paradigm to automate and enhance core tasks such as data cleaning, integration, and discovery, it is possible to substantially improve the accuracy and robustness of data processing, reduce the reliance on human intervention, and strengthen cross-scenario generalization capabilities. Consequently, this paradigm expands the boundaries of data applications and provides critical technical support for constructing high-quality datasets and enabling data-driven systems.
Existing studies have predominantly conducted surveys from the perspective of a single task or technique. One line of work focuses on data discovery, summarizing methods for data search, navigation, annotation, and schema inference. Another line centers on data cleaning, providing systematic reviews of error types, problem formulations, and representative approaches. Additionally, some studies investigate the potential of large models in data processing pipelines, examining their roles in data processing, storage, and analysis. However, from an overall perspective, existing surveys still lack comprehensive analyses that span multiple key stages of the data engineering lifecycle. Motivated by this gap, current research adopts a process-oriented perspective on the core data engineering pipeline, focusing on three key stages: data cleaning, data linking, and data discovery. These tasks directly address the challenges of quality improvement and semantic integration for multi-source heterogeneous data, representing the most fundamental data understanding and processing problems in data engineering. They have also emerged as some of the most active areas for the application of artificial intelligence techniques in recent years. Accordingly, this paper adopts a unified survey perspective centered on these three tasks, aiming to provide a cross-stage, comprehensive analysis. Furthermore, it systematically reviews the evolution of methods ranging from traditional machine learning approaches to large-model-based techniques across these tasks, establishes a technical taxonomy for intelligent data engineering methods, and summarizes the key characteristics of different approaches along with their empirical performance on commonly used datasets. This work seeks to promote the systematic development of intelligent methodologies in data engineering.
Traditional data engineering approaches primarily rely on rule-based systems and statistical learning techniques, which struggle to effectively capture complex semantics and correlations across multi-source data. Data wrangling has advanced the end-to-end automation of tasks such as data cleaning, integration, and discovery; however, it still exhibits notable limitations in complex and dynamic scenarios. In recent years, the emergence of large-scale models has introduced a new paradigm for intelligent data engineering, driving a shift from rule-driven to learning-driven data processing. Future research should target complex and diverse scenarios, with a focus on enhancing autonomous decision-making capabilities, end-to-end collaborative processing, and seamless integration with data infrastructures. Key directions include: 1) Autonomous data agents: Existing methods are typically based on fixed pipelines and lack adaptability to diverse task scenarios. Future work should incorporate data agents with capabilities for task decomposition and dynamic decision-making, enabling the automatic selection of appropriate methods and strategies based on data characteristics and task objectives. Through continuous feedback and optimization, such agents can facilitate the transition from static workflows to dynamic, adaptive decision processes. 2) Communication protocols for data agents: Data engineering involves multi-stage collaboration, necessitating standardized communication mechanisms among agents. Building upon existing general-purpose protocols, future research should further formalize task representations, data exchange formats, and access control policies, thereby supporting efficient collaboration in multi-agent systems. 3) Current research often overlooks the integration with underlying data infrastructures such as data lakes and databases. Future efforts should emphasize the design of interface protocols that enable agents to automatically translate high-level strategies into executable low-level operations (e.g., user-defined functions, UDFs), thereby establishing a closed-loop system from strategy generation to execution.
京公网安备11010802044758号