CClinicalTrials.gg
CompletedNCT06911645Updated Apr 4, 2025

Diagnostic Reasoning With Customized GPT-4 Model

An interventional study of Immediate access to customized version of GPT-4 and Access to customized version of GPT-4 following use of conventional resources in Pathologic Processes and Disease, sponsored by Stanford University. Completed at 1 site in United States. Per ClinicalTrials.gov, last updated 2025-04-04.

Sponsored by Stanford University · Not applicable, Interventional, and Diagnostic

Phase
Not applicable
Study type
Interventional
Enrollment
70
Allocation
Randomized
Sex
All
01

Study summary

This study will assess the impact of immediate access to a customized version of GPT-4, a large language model, on performance in case-based diagnostic reasoning tasks. Specifically, it will compare this approach to a two-step process where participants first use traditional diagnostic decision support tools to support their diagnostic reasoning before gaining access to the customized GPT-4 model.

Read the detailed description

Artificial intelligence (AI) technologies, particularly advanced large language models like OpenAI's ChatGPT, have the potential to enhance medical decision-making. While ChatGPT-4 was not specifically designed for medical applications, it has demonstrated promise in various healthcare contexts, including medical note-writing, addressing patient inquiries, and facilitating medical consultations. However, its impact on clinicians' diagnostic reasoning remains largely unknown.

Clinical reasoning is a complex process that involves pattern recognition, knowledge application, and probabilistic reasoning. Integrating AI tools like ChatGPT-4 into physician workflows could help reduce clinician workload and decrease the likelihood of missed diagnoses. However, ChatGPT-4 was neither developed nor validated for diagnostic reasoning, and it may produce misleading information, including plausible but incorrect conclusions that could misguide clinicians. If not used appropriately, it may fail to improve-and could even hinder-clinical decision-making. Therefore, it is essential to study how clinicians use large language models to support clinical reasoning before integrating them into routine patient care.

This study will examine how immediate access to a customized version of ChatGPT-4 impacts performance on case-based diagnostic reasoning tasks, compared to a stepwise approach. In the stepwise approach, participants will first use traditional diagnostic decision support tools to support their case reasoning before interacting with a customized ChatGPT-4 model, at which point they will have the opportunity to revise their initial answers.

Participants will be randomized into different study arms and will respond to diagnostic cases by providing three differential diagnoses, along with supporting and opposing findings for each. They will also identify their top diagnosis and propose next diagnostic steps. Independent reviewers, blinded to treatment assignment, will evaluate their responses.

02

Conditions studied

  • Pathologic Processes
  • Disease

Browse trials for

Keywords

  • Computer-assisted diagnosis
  • Large language models
  • Clinical reasoning
03

In context

Pathologic Processes

120 studies on the registry are indexed under Pathologic Processes; 28 are open to participants now.

This study's enrollment of 70 is below the median of 81 across 91 interventional studies indexed under Pathologic Processes.

Browse Pathologic Processes studies →

Lead sponsor

Stanford University is the lead sponsor of 2,117 studies on the registry; 425 are open to participants now.

Of its 259 completed or terminated interventional studies of FDA-regulated products, 197 (76%) have results posted.

Counted across the registry records on this site, refreshed daily.

04

Who can participate

Ages eligible
Child (0–17), Adult (18–64), Older adult (65+)
Sexes eligible
All
Accepts healthy volunteers
Yes

Inclusion criteria

  • Participants must be licensed physicians and have completed at least post-graduate year 1 (PGY1) of medical training.
  • Training in Internal medicine, family medicine, or emergency medicine.

Exclusion criteria

Exclusion Criteria:

  • Not currently practicing clinically.
  • Participated in one of our previous studies that used the same six diagnostic cases.
05

Study design

Phase
Not applicable
Primary purpose
Diagnostic
Allocation
Randomized
Intervention model
Parallel assignment
Masking
Single (Outcomes assessor)
Enrollment
70 participants (actual)

Study arms

  • Active comparator
    Immediate access to customized version of GPT-4

    Group will be encouraged to immediately use a customized version of GPT-4.

    Other: Immediate access to customized version of GPT-4

  • Active comparator
    Conventional resources first, then granted access to customized version of GPT-4.

    Group will be encouraged to first use any resources they wish besides large language models (UpToDate, Pubmed, google, etc) and then will be granted access to a customized version of GPT-4.

    Other: Access to customized version of GPT-4 following use of conventional resources

Interventions

  • OtherImmediate access to customized version of GPT-4

    Group is given immediate access to a customized version of GPT-4 to support their diagnostic reasoning for each case.

  • OtherAccess to customized version of GPT-4 following use of conventional resources

    Group is first encouraged to reason through diagnostic cases with the support of conventional resources. After they submit a case's answers they are then given access to a customized version of GPT-4 and have the opportunity to change their initial answers.

06

What researchers measure

Primary outcomes

  1. Diagnostic reasoning

    The primary outcome will be the percentage of correct responses per case (range: 0 to 100). For each case, participants will be asked to provide their top three differential diagnoses, along with supporting and opposing findings for each. They will receive 1 point for each plausible diagnosis. Supporting and opposing findings will be graded based on correctness, with 1 point for a partially correct response and 2 points for a completely correct response. Participants will then select their top diagnosis, earning 1 point for a reasonable choice and 2 points for the most accurate diagnosis. Finally, they will list up to three next steps for further patient evaluation, with 1 point awarded for a partially correct response and 2 points for a completely correct response. The primary outcome will be analyzed at the case level, comparing performance between the randomized study groups.

    Time frame: Through study completion, an average of 6 months

Secondary outcomes

  1. Time Spent Per Case

    The investigators will compare the average time (in minutes) participants spend on each case across the two study arms.

    Time frame: Through study completion, an average of 6 months

  2. Prompt frequency

    The investigators will compare the frequency of participant prompts to the customized GPT-4 model between the two study groups.

    Time frame: Through study completion, an average of 6 months

  3. Sentiment

    The investigators will compare the tone and sentiment of participant prompts to the customized GPT-4 model across the two study groups. The investigators will create a qualitative coding system to categorize the nature of the participants' prompts.

    Time frame: Through study completion, an average of 6 months

  4. Participant Perceptions of AI in Clinical Reasoning

    This outcome would be assessed in both study arms and would encompass changes in attitudes, confidence, and willingness to use AI diagnostic tools before and after being exposed to the customized tool. We will assess the number of participants who were open to using AI to help with complex clinical reasoning (pre- and post-quiz), if they enjoyed working with the AI diagnostic tool, if they felt like the tool provided a valuable collaborative experience for clinical reasoning, if seeing the AI diagnostic tool's recommendations increased their confidence in their differential diagnoses, and if they would use an AI diagnostic tool like the one in this study in their daily job. These will be evaluated on a Likert scale ranking from strongly disagree to strongly agree.

    Time frame: Through study completion, an average of 6 months

  5. Customized GPT-4's diagnostic reasoning

    The customized GPT-4's 'independent' diagnoses will be assessed for accuracy. The outcome will be the percentage of correct responses per case (range: 0 to 100). For each case, the meta-prompt directs the customized GPT-4 to provide its top three differential diagnoses, along with supporting and opposing findings for each, a final diagnosis, and next steps. The customized GPT-4 will receive 1 point for each plausible diagnosis. Supporting and opposing findings will be graded based on correctness, with 1 point for a partially correct response and 2 points for a completely correct response. Its top diagnosis will earn 1 point for a reasonable choice and 2 points for the most accurate diagnosis. Finally, it will list up to three next steps for further patient evaluation, with 1 point awarded for a partially correct response and 2 points for a completely correct response. The outcome will be analyzed at the case level, comparing performance with the randomized study groups' scores.

    Time frame: Through study completion, an average of 6 months

07

Study locations

1 site
  • Stanford University
    Palo Alto, California 94305, United States
08

References and documents

Individual participant data

Plan to share: No

No publications or documents are linked to this record.

09

Updates

Tracking since Sep 25, 2026
No changes since tracking began. The registry record was last updated on Apr 4, 2025, before this site started recording changes on Sep 25, 2026. Its history is on ClinicalTrials.gov ↗
10

Registry details

Key details

Study ID
NCT06911645
Lead sponsor
Stanford University
Collaborators
Beth Israel Deaconess Medical Center
Responsible party
Jonathan Chen (Assistant Professor of Medicine (Biomedical Informatics) and of Biomedical Data Science, Stanford University) — Principal investigator
First posted
Apr 4, 2025
Start date
Dec 16, 2024
Primary completion
Jan 24, 2025
Completion
Jan 24, 2025
Last update
Apr 4, 2025

Study contacts

Jonathan H Chen, MD, PhD
principal investigator · Stanford University

Oversight

Data monitoring committee
No
FDA-regulated drug
No
FDA-regulated device
No
View the source record on ClinicalTrials.gov ↗

Not currently enrolling

This study is completed, as verified in Mar 2025. You cannot join it, but the record below documents what was studied.

Follow this study

Get an email when the registry record changes — status, dates, results — or when someone posts here.

Sign in to follow

Discussion

Questions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.

Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.

Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.

Start the discussion