AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
PDF (1.8 MB)
Collect
Submit Manuscript AI Chat Paper
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Publishing Language: Chinese

Research on the Proxy Validity of Large Language Model in the Development of Critical Thinking Disposition Scale——Comparative Analysis Based on Simulated Samples and Student Samples

Lu-Yue LIWen-Jing LINJun-Lei DUQin-Hua ZHENG
Research Center of Distance Education, Beijing Normal University, Beijing, China 100875
Show Author Information

Abstract

The development of critical thinking disposition scale has long relied on large-scale samples and high-cost data. However, large language model (LLM) offers a new pathway for simulation, whose proxy validity urgently requires validation. Based on this, the paper generated simulated samples by invoking GLM-4-Plus, GPT-4o, and o1-preview under two modes of basic instruction and role constraint, and compared them with the real samples collected from 1,477 Chinese eighth-grade students. By analyzing distributional differences, reliability and validity differences, and algorithmic biases, the results showed that the proxy validity of the model under the basic instruction mode was overall better than that under the role constraint mode. GLM-4-Plus excelled in response quantity and gender balance, while o1-preview generated simulated samples with favorable distribution and internal consistency under the basic instruction mode; GLM-4-Plus and GPT-4o performed better in model fitting, yet all three models presented gender biases. All simulated samples were distributed deviated from the “openness and inclusivity” dimension, accompanied by poor reliability in this dimension, and the “self-reflection” dimension showed unsatisfactory validity. Based on this conclusion, the paper discussed the proxy validity and simulation limitations of LLM, pointing out that LLM can effectively simulate human responses and assist in scale development, but cannot completely replace human samples, and algorithm biases needed to be guarded against. Finally, this paper proposed pathways for improving the proxy validity of LLM from two aspects of technical optimization and scale development. The research in this paper helped significantly reduce the cost of educational measurement research, improve the overall research efficiency and accessibility, promote the intelligence of critical thinking assessment tools, and provide a new methodological paradigm for artificial intelligence-assisted educational empirical research.

CLC number: G40-057 Document code: A Article ID: 1009-8097(2026)04-0033-10

References

【1】
【1】
 
 
Modern Educational Technology
Pages 33-42

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
LI L-Y, LIN W-J, DU J-L, et al. Research on the Proxy Validity of Large Language Model in the Development of Critical Thinking Disposition Scale——Comparative Analysis Based on Simulated Samples and Student Samples. Modern Educational Technology, 2026, 36(4): 33-42. https://doi.org/10.3969/j.issn.1009-8097.2026.04.004

5

Views

0

Downloads

0

Crossref

Received: 02 September 2025
Published: 01 April 2026
© The journal of Modern Educational Technology