AI Chat Paper
Note: Please note that the following content is generated by AMiner AI. SciOpen does not take any responsibility related to this content.
{{lang === 'zh_CN' ? '文章概述' : 'Summary'}}
{{lang === 'en_US' ? '中' : 'Eng'}}
Chat more with AI
Article Link
Collect
Submit Manuscript
Show Outline
Outline
Show full outline
Hide outline
Outline
Show full outline
Hide outline
Regular Paper

Exploring LLM-Based Data Synthesis Strategies for Aligning Medical Consultation Preferences

School of Computer Science, Peking University, Beijing 100190, China
Key Laboratory of High Confidence Software Technologies (Peking University), Ministry of Education, Beijing 100190, China
Beijing Key Laboratory of Traffic Data Analysis and Mining, Beijing Jiaotong University, Beijing 100190, China
Show Author Information

Abstract

This research explores the application of reinforcement learning from artificial intelligence feedback (RLAIF) techniques to enhance healthcare consultation models, with the aim of addressing the challenges associated with preference-aligned data synthesis while reducing the dependence on medical experts. Specifically, we investigate the use of RLAIF in the generation of medical dialogues, focusing on two primary challenges: accurately reflecting doctors’ preferences and the unreliability of existing automated assessment systems. To address these issues, we propose a two-stage approach for synthesizing preference-aligned datasets. In the first stage, we leverage the dialogue continuation capabilities of a large language model to sample diverse, contextually aligned dialogue branches, employing one-shot learning for intervention. The second stage involves modeling doctors’ preferences through both outcome and process feedback. For outcome feedback, a rule-based reward system is utilized, whereas a planning-based reward strategy is employed for process feedback. To validate our approach, we develop the Chinese Standardized Patient Test (CSPT) dataset that emphasizes user guiding, instruction following, and synthesis ability, and construct an objective assessment system based on standardized patient testing. Experimental results demonstrate that our data synthesis approach performs well across five datasets, achieving a 17.6% improvement in diagnostic accuracy with outcome feedback and a 23.3% improvement with process feedback.

Electronic Supplementary Material

Download File(s)
JCST-2410-14929-Highlights.pdf (977.9 KB)

References

【1】
【1】
 
 
Journal of Computer Science and Technology
Pages 1485-1498

{{item.num}}

Comments on this article

Go to comment

< Back to all reports

Review Status: {{reviewData.commendedNum}} Commended , {{reviewData.revisionRequiredNum}} Revision Required , {{reviewData.notCommendedNum}} Not Commended Under Peer Review

Review Comment

Close
Close
Cite this article:
Dou C-F, Zhang Y, Jin Z, et al. Exploring LLM-Based Data Synthesis Strategies for Aligning Medical Consultation Preferences. Journal of Computer Science and Technology, 2025, 40(6): 1485-1498. https://doi.org/10.1007/s11390-025-4929-7

385

Views

0

Crossref

0

Web of Science

0

Scopus

0

CSCD

Received: 17 October 2024
Accepted: 27 October 2025
Published: 01 November 2025
© Institute of Computing Technology, Chinese Academy of Sciences 2025