With the prevalence of large language models (LLMs), AI-synthesized data has attracted extensive attention as an innovative tool reshaping the evidence base of educational research. However, this emerging practice, expanded from statistics to educational research, has sparked profound controversies over the changing nature of scientific evidence, with its application boundaries and potential risks remaining unclear. This paper reviewed the evolutionary trajectory of the synthetic data from statistical disclosure to LLM generation, analyzing how LLMs reshaped the generation logic of synthetic data through world models, theory-of-mind simulation and other mechanisms, and systematically explored its application forms across quantitative, qualitative, experimental simulation, evaluative research and other scenarios. Furthermore, it identified core challenges including representational distortion, cognitive mechanism discrepancies, inadequate ethical norms, and difficulties in quality assessment. This study highlighted the context dependency of the effective application of synthetic data, and called for constructing a new epistemological system adapted to human-machine collaborative research to promote the prudent and responsible application of this emerging tool.
Publications
- Article type
- Year
Year
Issue
Modern Educational Technology 2026, 36(5): 16-26
Published: 01 May 2026
Downloads:0
Total 1
京公网安备11010802044758号