CClinicalTrials.gg
Not yet recruitingNCT07696221ASA-LLMUpdated Jul 10, 2026

Large Language Models Versus Anesthesiologists for ASA Physical Status Classification

An observational study in Anesthesia, Preoperative Risk Prediction and Preoperative Risk Assessment, sponsored by Marmara University Pendik Training and Research Hospital. Not yet recruiting. Open to participants aged 18 Years and older. Per ClinicalTrials.gov, last updated 2026-07-10.

Sponsored by Marmara University Pendik Training and Research Hospital · Observational

Study type
Observational
Model
Cohort
Time perspective
Retrospective
Enrollment
350
Ages
18 Years and older
Sex
All
01

Study summary

The American Society of Anesthesiologists Physical Status (ASA-PS) classification is a cornerstone of preoperative risk assessment, yet interrater variability among clinicians is well documented. Large language models (LLMs) have recently demonstrated expert-level performance in several clinical classification tasks, including ASA-PS assignment.

This retrospective observational study evaluates whether four widely used LLMs - ChatGPT, DeepSeek, Gemini, and Claude - can accurately and consistently assign ASA-PS classes from structured, fully anonymized clinical vignettes derived from real preoperative anesthesia evaluations, using a consensus of senior anesthesiologists as the reference standard.

No patient data will be transmitted to third-party platforms. Clinical information will be converted by the investigators into de-identified structured vignettes containing only age range, sex, body mass index range, presence or absence of systemic diseases, functional capacity, and the major/minor nature of the planned surgery, in full compliance with national data protection legislation (KVKK).

Read the detailed description

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at Marmara University Pendik Training and Research Hospital will be included retrospectively. For each patient, demographic data (age, sex, body mass index), systemic comorbidities (hypertension, diabetes mellitus, coronary artery disease, chronic obstructive pulmonary disease, and others), functional capacity (metabolic equivalents, MET), type of planned surgery (major/minor), and the ASA-PS class assigned by the attending anesthesiologist will be recorded.

Clinical data will be anonymized and converted into structured clinical vignettes by the investigators. Vignettes will contain no identifiers, dates, protocol numbers, or rare diagnostic combinations that could directly or indirectly identify a patient.

Standardization of the LLM assessment process: To ensure independence between assessments, each vignette will be evaluated in a separate, history-free session. A new conversation will be initiated in the relevant model for every patient vignette, thereby eliminating the possibility that the model is influenced by its responses to previous vignettes (context anchoring). The ASA-PS class assigned to one vignette will not be carried over as context into the evaluation of any subsequent vignette. Each vignette will be presented to all four models using an identical, standardized prompt requesting only an ASA-PS class (I-VI) with a brief rationale, in a strictly defined output format. Model outputs will play no role in clinical decision-making. External information retrieval by the models will be disabled, and all queries will be completed within a narrow time window to minimize variability in model versions.

Each vignette will be submitted to each model once (single querying). Consequently, the intra-model test-retest reliability of the LLMs will not be assessed; this is acknowledged as a study limitation, consistent with the probabilistic nature of large language models, which may produce between-session variability in their outputs.

Model versions: The current version of each model available at the time of data collection will be used - ChatGPT (GPT-5.5, OpenAI), Gemini (Gemini 3.5, Google DeepMind), DeepSeek (DeepSeek V4, DeepSeek AI), and Claude (Claude Opus 4.8, Anthropic). These versions reflect the versions current at the time of protocol submission; the most recent stable version of each model accessible during data collection will be used, and the exact version and access date will be recorded. Because publicly available chat interfaces may perform automatic background routing to different model tiers, this is acknowledged as a reproducibility limitation.

The reference standard ASA-PS class will be determined by an independent, blinded panel of at least three senior anesthesiologists; consensus or majority vote will define the reference classification.

Statistical analysis: The primary (confirmatory) analysis will quantify the agreement between each LLM and the reference standard using quadratic weighted Cohen's kappa, respecting the ordinal structure of ASA-PS. Multi-rater agreement across the four models and the human raters will be assessed with Fleiss' kappa. Pairwise accuracy comparisons among the four models (six pairwise contrasts) will be treated as secondary/exploratory analyses and compared with McNemar or permutation tests for paired data, applying correction for multiple comparisons (e.g., Bonferroni or Holm); 95% confidence intervals will be estimated by bootstrap methods. Prespecified subgroup analyses include ASA III-IV boundary cases, multimorbidity burden, major versus minor surgery, and rater experience.

Primary hypothesis: The ASA-PS assignments of the LLMs (ChatGPT, DeepSeek, Gemini, and Claude) will show at least good agreement with the reference standard (weighted kappa ≥ 0.60). Secondary hypothesis: LLM errors will cluster in specific subgroups (e.g., the ASA III-IV boundary, multimorbid patients).

02

Conditions studied

  • Anesthesia
  • Preoperative Risk Prediction
  • Preoperative Risk Assessment

Browse trials for

Keywords

  • asa physical status
  • large language models
  • artificial intelligence
  • chatgpt
  • gemini
  • deepseek
  • Claude
  • risk classification
03

Who can participate

Ages eligible
18 Years and older
Sexes eligible
All
Accepts healthy volunteers
No
Sampling method
Probability sample

Study population

Adult patients who underwent preoperative anesthesia evaluation before elective surgery at a tertiary university hospital in Istanbul, Turkey.

Inclusion criteria

  • Age 18 years or older
  • Planned elective surgery
  • Completed preoperative anesthesia evaluation

Exclusion criteria

Exclusion Criteria:

  • Emergency surgical procedures
  • ASA VI (brain death)
  • Incomplete clinical records
04

Study design

Observational model
Cohort
Time perspective
Retrospective
Enrollment
350 participants (estimated)
Patient registry
No

Groups and cohorts

  • elective surgery patients

    Adult patients (≥18 years) who underwent preoperative anesthesia evaluation before elective surgery. Anonymized structured vignettes derived from their records will be classified by four LLMs (ChatGPT, DeepSeek, Gemini, Claude) and by a blinded senior anesthesiologist panel serving as the reference standard.

05

What researchers measure

Primary outcomes

  1. Agreement between LLM-assigned and reference-standard ASA-PS class

    Quadratic weighted Cohen's kappa between each large language model's ASA-PS assignment (ChatGPT, DeepSeek, Gemini, Claude) and the reference standard defined by consensus of a blinded panel of at least three senior anesthesiologists. Agreement of at least "good" level (weighted kappa ≥ 0.60) is hypothesized.

    Time frame: Through study completion, an average of 3 months

Secondary outcomes

  1. Overall classification accuracy of each LLM

    Proportion of vignettes in which the LLM-assigned ASA-PS class exactly matches the reference standard, with exploratory pairwise comparisons among the four models (McNemar/permutation tests, corrected for multiple comparisons)

    Time frame: Through study completion, an average of 3 months

  2. Subgroup error patterns

    Frequency and direction (over- vs. under-classification) of LLM misclassifications in prespecified subgroups: ASA III-IV boundary, multimorbidity, major vs. minor surgery

    Time frame: Through study completion, an average of 3 months

06

Study locations

No study locations are listed for this record.

07

References and documents

Publications

  • Chen YH, Ruan SJ, Chen PF. Predicting 30-Day Postoperative Mortality and American Society of Anesthesiologists Physical Status Using Retrieval-Augmented Large Language Models: Development and Validation Study. J Med Internet Res. 2025 Jun 3;27:e75052. doi: 10.2196/75052. PubMed 40460423 ↗
  • Cheng T, Li Y, Gu J, He Y, He G, Zhou P, Li S, Xu H, Bao Y, Wang X. The performance of ChatGPT in day surgery and pre-anesthesia risk assessment: a case-control study of 150 simulated patient presentations. Perioper Med (Lond). 2024 Nov 21;13(1):111. doi: 10.1186/s13741-024-00469-6. PubMed 39574189 ↗
  • Yoon SB, Lee J, Lee HC, Jung CW, Lee H. Comparison of NLP machine learning models with human physicians for ASA Physical Status classification. NPJ Digit Med. 2024 Sep 28;7(1):259. doi: 10.1038/s41746-024-01259-6. PubMed 39341936 ↗
  • Chung P, Fong CT, Walters AM, Aghaeepour N, Yetisgen M, O'Reilly-Shah VN. Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication. JAMA Surg. 2024 Aug 1;159(8):928-937. doi: 10.1001/jamasurg.2024.1621. PubMed 38837145 ↗
  • Turan EI, Baydemir AE, Ozcan FG, Sahin AS. Evaluating the accuracy of ChatGPT-4 in predicting ASA scores: A prospective multicentric study ChatGPT-4 in ASA score prediction. J Clin Anesth. 2024 Sep;96:111475. doi: 10.1016/j.jclinane.2024.111475. Epub 2024 Apr 23. PubMed 38657530 ↗

Individual participant data

Plan to share: No

08

Registry details

Key details

Study ID
NCT07696221
Lead sponsor
Marmara University Pendik Training and Research Hospital
Responsible party
dilara gocmen (asistan prof, Marmara University Pendik Training and Research Hospital) — Principal investigator
First posted
Jul 10, 2026
Start date
Jul 21, 2026 (estimated)
Primary completion
Aug 21, 2026 (estimated)
Completion
Oct 21, 2026 (estimated)
Last update
Jul 10, 2026

Study contacts

Dilara Göçmen, Assistant Prof
Contact
dilara.gocmen@marmara.edu.tr
+905413439438

Oversight

Data monitoring committee
No
FDA-regulated drug
No
FDA-regulated device
No
View the source record on ClinicalTrials.gov ↗

Not currently enrolling

This study is not yet recruiting, as verified in Jul 2026. You cannot join it, but the record below documents what was studied.

Follow this study

Get an email when the registry record changes — status, dates, results — or when someone posts here.

Sign in to follow

Discussion

Questions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.

Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.

Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.

Start the discussion