CClinicalTrials.gg
CompletedNCT06208423Updated Sep 27, 2024

Physician Reasoning on Management Cases With Large Language Models

An interventional study of GPT-4 in Clinical Decision-making, sponsored by Stanford University. Completed at 1 site in United States. Per ClinicalTrials.gov, last updated 2024-09-27.

Sponsored by Stanford University · Not applicable, Interventional, and Treatment

Phase
Not applicable
Study type
Interventional
Enrollment
92
Allocation
Randomized
Sex
All
01

Study summary

This study will evaluate the effect of providing access to GPT-4, a large language model, compared to traditional management decision support tools on performance on case-based management reasoning tasks.

Read the detailed description

Artificial intelligence (AI) technologies, specifically advanced large language models like OpenAI's ChatGPT, have the potential to improve medical decision-making. Although ChatGPT-4 was not developed for its use in medical-specific applications, it has demonstrated promise in various healthcare contexts, including medical note-writing, addressing patient inquiries, and facilitating medical consultation. However, little is known about how ChatGPT augments the clinical reasoning abilities of clinicians.

Clinical reasoning is a complex process involving pattern recognition, knowledge application, and probabilistic reasoning. Integrating AI tools like ChatGPT-4 into physician workflows could potentially help reduce clinician workload and decrease the likelihood of mismanagement. However, ChatGPT-4 was not developed for clinical reasoning nor has it been validated for this purpose. Further, it may be subject to disinformation, including convincing confabulations that may mislead clinicians. If clinicians misuse this tool, it may not improve reasoning and could even cause harm. Therefore, it is important to study how clinicians use large language models to augment clinical reasoning prior to routine incorporation into patient care.

In this study, participants will be randomized to answer clinical management cases with or without access to ChatGPT-4. Each case has multiple components, and the participants will be asked to discuss their reasoning for each component. Answers will be graded by independent reviewers blinded to treatment assignment. A grading rubric was developed for each case by a panel of 4-7 expert discussants. Discussants independently developed a rubric for each case, and then any discrepancies were resolved through multiple rounds of discussions.

02

Conditions studied

  • Clinical Decision-making

Keywords

  • clinical decision support
03

In context

Lead sponsor

Stanford University is the lead sponsor of 2,117 studies on the registry; 425 are open to participants now.

Of its 259 completed or terminated interventional studies of FDA-regulated products, 197 (76%) have results posted.

Counted across the registry records on this site, refreshed daily.

04

Who can participate

Ages eligible
Child (0–17), Adult (18–64), Older adult (65+)
Sexes eligible
All
Accepts healthy volunteers
Yes

Inclusion criteria

  • Participants must be licensed physicians and have completed at least post-graduate year 2 (PGY2) of medical training.
  • Training in Internal medicine, family medicine, or emergency medicine.

Exclusion criteria

Exclusion Criteria:

  • Not currently practicing clinically.
05

Study design

Phase
Not applicable
Primary purpose
Treatment
Allocation
Randomized
Intervention model
Parallel assignment
Masking
Single (Outcomes assessor)
Enrollment
92 participants (actual)

Study arms

  • Active comparator
    GPT-4

    Group will be given access to GPT-4

    Other: GPT-4

  • No intervention
    Usual Resources

    Group will not be given access to GPT-4 but will be encouraged to use any resources they wish besides large language models (UpToDate, Dynamed, google, etc).

Interventions

  • OtherGPT-4

    OpenAI's GPT-4 large language model with chat interface.

06

What researchers measure

Primary outcomes

  1. Management Reasoning

    Percent correct (range: 0 to 100) for each case.

    Time frame: Within one-hour study

Secondary outcomes

  1. Time Spent on Management

    Time (in minutes) participants spend per case between the two study arms.

    Time frame: Within one-hour study

07

Study locations

1 site
  • Stanford University
    Palo Alto, California 94304, United States
08

References and documents

Individual participant data

Plan to share: No

No publications or documents are linked to this record.

09

Updates

Tracking since Sep 25, 2026
No changes since tracking began. The registry record was last updated on Sep 27, 2024, before this site started recording changes on Sep 25, 2026. Its history is on ClinicalTrials.gov ↗
10

Registry details

Key details

Study ID
NCT06208423
Lead sponsor
Stanford University
Collaborators
Beth Israel Deaconess Medical Center, University of Minnesota
Responsible party
Jonathan Chen (Assistant Professor of Medicine, Stanford University) — Principal investigator
First posted
Jan 17, 2024
Start date
Dec 28, 2023
Primary completion
Apr 19, 2024
Completion
Apr 19, 2024
Last update
Sep 27, 2024

Study contacts

Jonathan H Chen, MD, PhD
principal investigator · Stanford University
Adam Rodman, MD
principal investigator · Beth Israel Deaconess Medical Center
Andrew Olson, MD
principal investigator · University of Minnesota

Oversight

Data monitoring committee
No
FDA-regulated drug
No
FDA-regulated device
No
View the source record on ClinicalTrials.gov ↗

Not currently enrolling

This study is completed, as verified in Sep 2024. You cannot join it, but the record below documents what was studied.

Follow this study

Get an email when the registry record changes — status, dates, results — or when someone posts here.

Sign in to follow

Discussion

Questions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.

Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.

Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.

Start the discussion