An interventional study of Socratic Agent for Guided Epidemiology and General LLM in an Intelligent Teaching Agent for the "Epidemiology" Course of the MBBS Program and Students' Critical Thinking, Practical Ability and Learning Outcomes, sponsored by The Fourth Affiliated Hospital of Zhejiang University School of Medicine. Recruiting at 1 site in China. Open to participants aged 16 Years to 25 Years, including healthy volunteers. Per ClinicalTrials.gov, last updated 2026-08-07.
Sponsored by The Fourth Affiliated Hospital of Zhejiang University School of Medicine · Not applicable, Interventional, and Basic science
Prior to the quasi-experiment, this study performed semi-structured interviews among Bachelor of Medicine and Bachelor of Surgery(MBBS) undergraduate students from the Belt and Road International Medical College, Zhejiang University. Based on learning science, cognitive load theory and causal inference frameworks, the interview outline focused on system usability, feedback quality, learning reasoning processes, cognitive burden and instructional optimization advice. Each 20-30 minute individual interview was audio-recorded with participants' informed consent and verbatim transcribed for standardized qualitative analysis.
The quasi-experiment recruited no fewer than 60 MBBS students. The sample size was determined by intergroup statistical power analysis to guarantee around 30 participants per group for valid intergroup comparison. All eligible students were randomly assigned into two groups with balanced demographic and academic baseline characteristics. The experimental group (n=30) received structured epidemiological learning assisted by the (Socratic Agent for Guided Epidemiology) SAGE (Artificial Intelligence)AI agent with professional Socratic cognitive guidance. The control group (n=30) adopted general large language model (LLM)-based learning without systematic thinking intervention. Participants with clinically diagnosed severe mental disorders, cognitive dysfunction or inability to finish the complete research process were excluded. A baseline epidemiological knowledge pre-test confirmed no significant academic differences between the two groups via independent samples t-test, and all participants had no prior experience of AI-assisted medical learning.
The semi-structured interviews were conducted to collect students' authentic learning experiences, interactive perceptions and cognitive characteristics during the use of SAGE agent and general LLMs, providing qualitative evidence for the iterative optimization of AI teaching tools. All interviewees completed baseline assessments and preliminary AI learning trials, ensuring qualified professional foundation and genuine interactive experience. Conducted by trained researchers, the standardized interviews centered on three core themes: system usability and feedback clarity; AI-induced changes in information extraction, hypothesis formulation and causal inference; and common learning barriers including interactive obstacles, comprehension difficulties, cognitive overload and potential AI over-reliance. Transcribed interview data were analyzed through thematic analysis to summarize typical user experience patterns. Qualitative outcomes were triangulated with quantitative experimental results to revise the SAGE teaching protocol, optimize agent prompt chains and improve the interpretation of experimental findings.
The quasi-experiment consisted of three standardized stages. In the pre-test stage, all participants signed informed consent, completed a 25-item clinical epidemiology knowledge scale, an 11-item reasoning ability test and a demographic questionnaire to establish consistent baseline levels. In the intervention stage, the experimental group received standardized training in confounder identification and causal inference construction in strict accordance with the SAGE teaching protocol. The SAGE agent improved students' advanced epidemiological reasoning ability through continuous multi-round Socratic questioning and targeted cognitive guidance. The control group received equal-duration learning in the same experimental environment, only using conventional search engines and unguided LLMs for basic information retrieval without any cognitive and thinking intervention. All participants submitted screenshots to record their accurate AI tool usage duration after completing learning tasks.
In the post-test stage, all participants finished parallel-version epidemiological knowledge assessments and unified reasoning ability tests. Validated scales were adopted to evaluate students' cognitive load, system usability, learning satisfaction and academic self-confidence. Students' final scores of the Epidemiology course were collected as supplementary indicators of long-term learning effectiveness. Upon the completion of data collection, backend AI interaction logs were summarized and strictly screened. Invalid samples with insufficient interaction rounds or incomplete responses were excluded to ensure high data quality and reliable experimental conclusions.
Exclusion Criteria:
Socratic Agent for Guided Epidemiology
Behavioral: Socratic Agent for Guided Epidemiology
General LLM
Behavioral: General LLM
A total of 30 participants are assigned to the experimental group receiving the SAGE teaching model, namely the Socratic Agent for Guided Epidemiology.
30 participants are allocated to the control group with teaching assistance from a general large language model (General LLM).
Clinical Epidemiology Knowledge Assessment Scale (Utrecht questionnaire on knowledge on clinical epidemiology for evidence-based practice)
The test consists of 19 multiple-choice questions and 2 calculation questions, each worth 1 point. There are also 4 essay questions, each worth 3 points. The total score is 33 points. The higher the score, the better the mastery of clinical epidemiology knowledge.
Time frame: Before the intervention (One to seven days before starting the epidemiology course) and After the intervention(Within half a month after completing the epidemiology course)
Epidemiological Reasoning Test
This section consists of eleven questions and is designed to test students' epidemiological reasoning skills. Scores range from 0 to 11, with higher scores indicating higher levels of epidemiological reasoning ability.
Time frame: Before the intervention (One to seven days before starting the epidemiology course ) and After the intervention(Within half a month after completing the epidemiology course)
System Usability Scale(SUS)
It consists of 10 items and is scored using the Likert scale (1 = strongly disagree, 5 = strongly agree). The total score is converted to a range of 0-100 to evaluate the overall usability of the system. Additionally, based on a standardized scoring system, the SUS score is assigned letter grades, ranging from "F" (0-60 points) to "A" (91-100 points).
Time frame: Immediately after intervention(One to seven days after completing the epidemiology course)
Needs and perceptions questionnaire (AI-powered simulation-based teaching agent)
This questionnaire employs the 5-point Likert scale (ranging from 1 to 5), with a total score range of 18 to 90. The higher the score, the greater the subject's acceptance of the AI-assisted teaching system, the effectiveness of teaching support, and the evaluation of its application value.
Time frame: Qualitative research stage (one to five months before starting the epidemiology course
Student's final exam score in the Epidemiology course
The score ranges from 0 to 100 points. A score of 60 or above is considered passing.
Time frame: Immediately after intervention(One to seven days after completing the epidemiology course)
Demographic information questionnaire
The title of this section is used to analyze demographic data, and it mainly includes Name,Tel,Birthday,Gender,Grade,Country,GPA or credit grade,Monthly Disposable income,your father's education level,your mother's education level, Previous AI Experience.
Time frame: Before the intervention (One to seven days before starting the epidemiology course)
Learning Satisfaction and Self-confidence Scale
Participants rated each item using a Likert-type response scale indicating their level of agreement with the statements (1 = strongly disagree to 5 = strongly agree). the total score ranges from 13 to 65.Higher scores indicated greater satisfaction with the learning experience and higher perceived self-confidence in learning.
Time frame: Immediately after intervention(One to seven days after completing the epidemiology course)
Cognitive Load Scale
The 10 items are scored on a scale of 0 ("completely not") to 10 ("completely"), with higher scores indicating greater cognitive load. Total score ranges from 0 to 100 points.
Time frame: After the intervention(Within half a month after completing the epidemiology course)
Plan to share: No
No publications or documents are linked to this record.
Eligibility is decided by the study team. Share this record with your doctor or contact the team directly.
Contact study teamGet an email when the registry record changes — status, dates, results — or when someone posts here.
Sign in to followQuestions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.
Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.
Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.
The Fourth Affiliated Hospital of Zhejiang University School of Medicine