An interventional study of GAPS-Agent and LLM in Large Language Models and Lung Cancer (NSCLC), sponsored by Peking University People's Hospital. Completed at 1 site in China. Open to participants aged 18 Years to 65 Years. Per ClinicalTrials.gov, last updated 2026-07-30.
Sponsored by Peking University People's Hospital · Not applicable, Interventional, and Other
This study is an exploratory effect-size estimation study, with the following specific objectives: ① to estimate the point estimate and 95% confidence interval of the Win Ratio for the experimental group (GAPS-Agent) versus the control group (large language model) in blinded pairwise preference judgments by thoracic surgery expert adjudicators, to serve as a sample size planning parameter for subsequent multicenter confirmatory clinical trials; ② to preliminarily evaluate the value of GAPS-Agent within clinical workflows.The hypothesis of this study is as follows: compared with a general-purpose large language model without medical enhancement (control group), a structured agentic workflow optimized on the basis of the GAPS evaluation framework (GAPS-Agent, experimental group) can help junior resident physicians generate clinical decision plans for complex lung cancer cases that are more strongly preferred by senior thoracic surgery expert adjudicators.
7,243 studies on the registry are indexed under Lung Neoplasms; 1,557 are open to participants now.
This study's enrollment of 8 is below the median of 60 across 5,295 interventional studies indexed under Lung Neoplasms.
Browse Lung Neoplasms studies →Peking University People's Hospital is the lead sponsor of 584 studies on the registry; 233 are open to participants now.
Counted across the registry records on this site, refreshed daily.
Resident Physician Subjects:
Study Cases:
Adjudication Expert Panel:
Exclusion Criteria:
Resident Physician Subjects:
Study Cases:
Adjudication Expert Panel:
GAPS-Agent
Other: GAPS-Agent
LLM
Other: LLM
The research group has previously developed the GAPS evaluation framework for complex clinical decision-making in lung cancer. In this framework, G (Grounding) characterizes the cognitive depth of decision-making (ranging from knowledge retrieval to decisions that go beyond clinical guidelines), A (Authority) corresponds to the grading of evidence strength, P (Perturbation) describes the identification and management of real-world clinical confounding factors, and S (Strength) corresponds to the calibration of recommendation strength. Within this framework, the research group has completed the construction of a 100-item complex lung cancer decision-making evaluation set along with its corresponding rubrics, and has invited multiple thoracic oncology experts to complete content validity validation. Based on this, the research group developed GAPS-Agent, which uses an open-source large language model as its foundation and integrates functional modules such as guideline and evidence retri
Open source large language model that is not specifically enhanced in medical field.
Overall plan Win Ratio
A total of 10 blinded expert judges made Win/Tie/Loss ternary preference judgments on 192 paired scheme comparisons in terms of overall scheme quality. The win ratio was calculated as Wins ÷ Losses, and the 95% confidence interval was estimated using a two-level (physician × case) cluster bootstrap resampling method (B = 10,000, quantile method on the log scale).
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
Inter-rater agreement
For the ternary preference judgment results of 10 expert judges across 192 paired comparisons and 6 evaluation domains, Fleiss' kappa was used to assess inter-rater agreement. The kappa value and its 95% confidence interval are reported for each evaluation domain.
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
Redundancy Win Ratio
A total of 10 blinded expert judges made Win/Tie/Loss ternary preference judgments on 192 paired scheme comparisons in terms of overall scheme quality. The win ratio was calculated as Wins ÷ Losses, and the 95% confidence interval was estimated using a two-level (physician × case) cluster bootstrap resampling method (B = 10,000, quantile method on the log scale).
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
Evidence-based medicine adherence Win Ratio
A total of 10 blinded expert judges made Win/Tie/Loss ternary preference judgments on 192 paired scheme comparisons in terms of overall scheme quality. The win ratio was calculated as Wins ÷ Losses, and the 95% confidence interval was estimated using a two-level (physician × case) cluster bootstrap resampling method (B = 10,000, quantile method on the log scale).
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
Actionability Win Ratio
A total of 10 blinded expert judges made Win/Tie/Loss ternary preference judgments on 192 paired scheme comparisons in terms of overall scheme quality. The win ratio was calculated as Wins ÷ Losses, and the 95% confidence interval was estimated using a two-level (physician × case) cluster bootstrap resampling method (B = 10,000, quantile method on the log scale).
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
Completeness Win Ratio
A total of 10 blinded expert judges made Win/Tie/Loss ternary preference judgments on 192 paired scheme comparisons in terms of overall scheme quality. The win ratio was calculated as Wins ÷ Losses, and the 95% confidence interval was estimated using a two-level (physician × case) cluster bootstrap resampling method (B = 10,000, quantile method on the log scale).
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
Safety Win Ratio
A total of 10 blinded expert judges made Win/Tie/Loss ternary preference judgments on 192 paired scheme comparisons in terms of overall scheme quality. The win ratio was calculated as Wins ÷ Losses, and the 95% confidence interval was estimated using a two-level (physician × case) cluster bootstrap resampling method (B = 10,000, quantile method on the log scale).
Time frame: Measured at the time when experts completed their preference judgements. Calculated up to 3 weeks after the preference judgements.
GAPS automated rubric score
A third-party large language model, independent of the two study arms' base models, served as the judge model and automatically scored all 96 plans according to the GAPS rubric.
Time frame: Generated up to 3 weeks after residents finished their plan generation.
Subject physician's self-confidence score
After submitting each case plan, the participating physicians self-rated their confidence in their own plan using a 1-5 point Likert scale.
Time frame: Completed at the time when residents submitted their plans. Calculated up to 3 weeks after the submission.
Tool satisfaction score
After submitting each case plan, the participating physicians rated their satisfaction with the tool using a 1-5 point Likert scale.
Time frame: Completed at the time when residents submitted their plans. Calculated up to 3 weeks after the submission.
Tool trustworthiness score
After submitting each case plan, the participating physicians rated the tool's credibility using a 1-5 point Likert scale.
Time frame: Completed at the time when residents submitted their plans. Calculated up to 3 weeks after the submission.
Decision-making time
The time taken (in minutes) by each participating physician to complete the production of each case plan was automatically recorded by the evaluation platform. Differences between groups were analyzed using a linear mixed-effects model.
Time frame: Completed at the time when residents submitted their plans. Calculated up to 3 weeks after the submission.
Plan to share: No
No publications or documents are linked to this record.
This study is completed, as verified in Jul 2026. You cannot join it, but the record below documents what was studied.
Get an email when the registry record changes — status, dates, results — or when someone posts here.
Sign in to followQuestions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.
Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.
Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.
Peking University People's Hospital