CClinicalTrials.gg
CompletedNCT06774612Updated Jul 17, 2025

The Impact of Large Language Models on Diagnostic Reasoning Among LLM-Trained Medical Doctors

An interventional study of ChatGPT-4o in Diagnosis, sponsored by Lahore University of Management Sciences. Completed at 1 site in Pakistan. Per ClinicalTrials.gov, last updated 2025-07-17.

Sponsored by Lahore University of Management Sciences · Not applicable, Interventional, and Diagnostic

Phase
Not applicable
Study type
Interventional
Enrollment
60
Allocation
Randomized
Sex
All
01

Study summary

This study aims to evaluate whether large language model-trained medical doctors demonstrate enhanced diagnostic reasoning performance when utilizing ChatGPT-4o alongside conventional resources compared to using conventional resources alone.

Read the detailed description

Diagnostic errors are a major source of preventable patient harm. Recent advances in Large Language Models (LLM), particularly ChatGPT-4o, have shown promise in enhancing medical decision-making. However, little is known about their impact on medical doctors' (e.g., physicians' and surgeons') diagnostic reasoning.

Diagnostic accuracy relies on complex clinical reasoning and careful evaluation of patient data. While AI assistance could potentially reduce errors and improve efficiency, ChatGPT-4o lacks medical validation and could introduce new risks through incorrect information generation (also known as hallucinations). To mitigate these risks, doctors need adequate training in understanding ChatGPT-4o's capabilities, limitations, and proper usage. Given these uncertainties and the importance of proper AI training, systematic evaluation is essential before clinical implementation.

This randomized study will assess whether ChatGPT-4o access improves LLM-trained medical doctors' diagnostic performance compared to conventional resources (e.g., textbooks, online medical databases) alone. All participating doctors will have completed at least a 10-hour training program covering ChatGPT-4o usage, prompt engineering techniques, and output evaluation strategies. Participants will provide differential diagnoses with supporting evidence and recommended next steps for clinical cases, with responses evaluated by blinded reviewers.

02

Conditions studied

  • Diagnosis

Browse trials for

Keywords

  • clinical reasoning
  • large language models
  • computer-assisted diagnosis
03

In context

Disease

1,327 studies on the registry are indexed under Disease; 596 are open to participants now.

This study's enrollment of 60 is below the median of 80 across 732 interventional studies indexed under Disease.

Browse Disease studies →

Lead sponsor

Lahore University of Management Sciences is the lead sponsor of 5 studies on the registry; 2 are open to participants now.

Counted across the registry records on this site, refreshed daily.

04

Who can participate

Ages eligible
Child (0–17), Adult (18–64), Older adult (65+)
Sexes eligible
All
Accepts healthy volunteers
Yes

Inclusion criteria

  • Full or Provisionally Registered Medical Practitioners with the Pakistan Medical and Dental Council (PMDC).
  • Completed Bachelor of Medicine, Bachelor of Surgery (MBBS) Exam. The equivalent degree of MBBS in US and Canada is called Doctor of Medicine (MD).
  • Participants must have completed a structured training program on the use of ChatGPT (or a comparable large language model), totaling at least 10 hours of instruction. The program must include hands-on practice related to LLM's aspects, specifically prompt engineering and content evaluation.

Exclusion criteria

Exclusion Criteria:

  • Any other Registered Medical Practitioners (Full or Provisional) with PMDC (e.g., Professionals with Bachelor of Dental Surgery or BDS).
05

Study design

Phase
Not applicable
Primary purpose
Diagnostic
Allocation
Randomized
Intervention model
Parallel assignment
Masking
None (open label)
Enrollment
60 participants (actual)

Study arms

  • Active comparator
    ChatGPT-4o

    Group will be given access to ChatGPT-4o.

    Other: ChatGPT-4o

  • No intervention
    Conventional resources

    Group will not be given access to ChatGPT-4o but will be encouraged to use any resources they wish besides large language models (PubMed, Google without AI Overviews, etc).

Interventions

  • OtherChatGPT-4o

    OpenAI's ChatGPT-4o large language model with chat interface.

06

What researchers measure

Primary outcomes

  1. Diagnostic reasoning

    The primary outcome will be the percent correct for each case (range: 0 to 100). For each case, participants will be asked for three top diagnoses, findings from the case that support that diagnosis, and findings from the case that oppose that diagnosis. For each plausible diagnosis, participants will receive 1 point. Findings supporting the diagnosis and findings opposing the diagnosis will also be graded based on correctness, with 1 point for partially correct and 2 points for completely correct responses. Participants will then be asked to name their top diagnosis, earning one point for a reasonable response and two points for the most correct response. Finally participants will be asked to name up to 3 next steps to further evaluate the patient with one point awarded for a partially correct response and two points for a completely correct response. The primary outcome will be compared on the case-level by the randomized groups.

    Time frame: Assessed at a single time point for each case, during the scheduled diagnostic reasoning evaluation session, which takes place between 0-4 days after participant enrollment.

Secondary outcomes

  1. Time Spent on Diagnosis

    We will compare how much time (in seconds) participants spend per case between the two study arms.

    Time frame: Assessed at a single time point for each case, during the scheduled diagnostic reasoning evaluation session, which takes place between 0-4 days after participant enrollment.

07

Study locations

1 site
  • Lahore University of Management Sciences
    Lahore, Punjab Province 54792, Pakistan
08

References and documents

Individual participant data

Plan to share: No

No publications or documents are linked to this record.

09

Updates

Tracking since Sep 25, 2026
No changes since tracking began. The registry record was last updated on Jul 17, 2025, before this site started recording changes on Sep 25, 2026. Its history is on ClinicalTrials.gov ↗
10

Registry details

Key details

Study ID
NCT06774612
Lead sponsor
Lahore University of Management Sciences
Collaborators
King Edward Medical University
Responsible party
Ihsan Ayyub Qazi, PhD (Associate Professor, PhD, Lahore University of Management Sciences) — Principal investigator
First posted
Jan 14, 2025
Start date
Jan 10, 2025
Primary completion
May 17, 2025
Completion
May 17, 2025
Last update
Jul 17, 2025

Study contacts

Ihsan Ayyub Qazi, PhD
principal investigator · Lahore University of Management Sciences
Muhammad Asadullah Khawaja, MBBS
principal investigator · King Edward Medical University
Ayesha Ali, PhD
principal investigator · Lahore University of Management Sciences

Oversight

Data monitoring committee
No
FDA-regulated drug
No
FDA-regulated device
No
View the source record on ClinicalTrials.gov ↗

Not currently enrolling

This study is completed, as verified in Jul 2025. You cannot join it, but the record below documents what was studied.

Follow this study

Get an email when the registry record changes — status, dates, results — or when someone posts here.

Sign in to follow

Discussion

Questions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.

Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.

Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.

Start the discussion