CClinicalTrials.gg
Active, not recruitingNCT06814327AI4ICU-ObsUpdated Aug 28, 2025

Assessment of Organ Failure Risk Predictions in ICU

An observational study in Circulatory Failure, Respiratory Failure and Renal Failure, sponsored by ETH Zurich. Active, not recruiting at 1 site in Switzerland. Open to participants aged 18 Years and older. Per ClinicalTrials.gov, last updated 2025-08-28.

Sponsored by ETH Zurich · Observational

Study type
Observational
Model
Cohort
Time perspective
Prospective
Enrollment
499
Ages
18 Years and older
Sex
All
01

Study summary

During this observational study, the investigators aim to assess the ability of ICU clinicians to predict the risk of impending organ failure and retrospectively compare it to the performance of previously published machine learning models. The central hypothesis of this study is that the treating physician can predict impending organ failure in adult ICU patients with similar accuracy as the best previously publishes machine learning models.

Read the detailed description

In this observational study, clinician's (physicians and nurses) assessment of the estimated imminent organ failure risk in an ICU setting are prospectively collected. Circulatory failure is investigated in the primary objective, and respiratory failure, renal failure, and mortality are investigated in secondary objectives. These assessments investigate the predictive performance and influencing factors for clinician prediction. The assessments will be collected in questionnaires and be performed by the clinicians directly involved in the patient treatment and by clinicians who are not actively responsible for the patient treatment. Furthermore, this study aims to benchmark these risk assessments made by healthcare professionals against retrospectively generated AI risk scores for the same patients and timepoints. The AI risk scores will be calculated retrospectively from a set of models from a systematic search of the current literature. The AI models that will be employed for this analysis will be identified as indicated by a systematic review protocol and must satisfy the following two criteria: they do not require any data beyond what is routinely collected during an ICU stay and may be accessed as open source. Such a comparison is vital for the understanding of the relative accuracy and reliability of AI-based predictions in the context of organ failure risk compared to human performance. The data and findings from this study are anticipated to provide evidence for the clinical utility of AI-based risk scores and pave the way for future research into the optimization of AI systems for healthcare applications.

02

Conditions studied

  • Circulatory Failure
  • Respiratory Failure
  • Renal Failure

Keywords

  • Machine learning
  • Intensive care
  • artificial intelligence
03

In context

Shock

916 studies on the registry are indexed under Shock; 176 are open to participants now.

This study's enrollment of 499 is above the median of 91 across 346 observational studies indexed under Shock.

Browse Shock studies →

Lead sponsor

ETH Zurich is the lead sponsor of 19 studies on the registry; 12 are open to participants now.

Counted across the registry records on this site, refreshed daily.

04

Who can participate

Ages eligible
18 Years and older
Sexes eligible
All
Accepts healthy volunteers
No
Sampling method
Probability sample

Study population

Adult Intensive Care Unit at the University Hospital Bern

Inclusion criteria

  • patient minimum age of 18 years
  • emergency admission to the ICU
  • arterial line in place

Exclusion criteria

Exclusion Criteria:

  • documented refusal (on the general consent form) to participate to clinical research
  • patients with neurologic conditions that impair the patient's level of consciousness (including, but not limited to stroke, traumatic brain injury, intracranial hemorrhage, CNS infections; except polytrauma)
  • patients on mechanical circulatory support systems (IABP, VA-ECMO, Impella, VAD) or extracorporeal membrane oxygenation (VV-ECMO) at any time during their ICU stay;
  • patients receiving end-of-life care or are admitted for the sole purpose of evaluating organ donation
05

Study design

Observational model
Cohort
Time perspective
Prospective
Enrollment
499 participants (actual)
Patient registry
No

Groups and cohorts

  • Adult ICU patients

    Adult ICU patients

06

What researchers measure

Primary outcomes

  1. Clinician prediction of circulatory failure within 8 hours compared to published ML models

    This outcome compares the area under the receiver operating characteristic curve (auROC) for two methods of predicting circulatory failure within 8 hours of each assessment time point: (1) ICU clinicians' risk estimates, and (2) previously published machine learning (ML) models applied retrospectively. For each assessment, we compute the auROC separately for clinicians and for the ML model for the same time points and patients. The difference in auROC (clinician minus ML) is the main measure of interest, evaluated under a non-inferiority framework with a margin of 0.025.

    Time frame: Assessments are collected within the first 72 hours following admission.

Secondary outcomes

  1. Clinician prediction of respiratory failure within 24 hours compared to published ML models

    This outcome compares the area under the receiver operating characteristic curve (auROC) for two methods of predicting respiratory failure within 24 hours of each assessment time point: (1) ICU clinicians' risk estimates, and (2) previously published machine learning (ML) models applied retrospectively. For each assessment, we compute the auROC separately for clinicians and for the ML model. The difference in auROCs (clinician minus ML) is the main measure of interest, evaluated using the same methodological framework as the primary outcome.

    Time frame: Assessments are collected within the first 72 hours following admission.

Other outcomes

  1. Exploratory analysis of clinician prediction of renal failure within 48 hours compared to published ML models

    This exploratory outcome compares the area under the receiver operating characteristic curve (auROC) for two methods of predicting renal failure within 48 hours of each assessment time point: (1) ICU clinicians' risk estimates, and (2) previously published machine learning (ML) models applied retrospectively. For each assessment, we compute the auROC separately for clinicians and for the ML model. The difference in auROCs is evaluated using the same methodological framework as the primary outcome.

    Time frame: Assessments are collected within the first 72 hours following admission.

  2. Exploratory analysis of clinician mortality prediction compared to published ML models

    All-cause mortality prediction by McNemar's test. Evaluated and tested using the performance of respective assessments by clinicians (binary response) and corresponding (paired) predictions by a machine learning model (probability prediction reduced to a binary response) trained on historical data for the three individual binary outcomes of all-cause mortality \<28-day, \<6-months, \<12-months. For each horizon separately the machine-based probabilities are thresholded to match the sensitivity of the clinicians, a McNemar's test is performed to test for a significant difference in predictive capabilities and p-values will be adjusted to account for multiple testing if necessary.

    Time frame: Assessments are collected within the first 72 hours following admission.

  3. Exploratory analysis of predictive accuracy of treating physicians versus treating nurses

    Evaluated and tested for the same as the primary, secondary and first exploratory outcomes (i.e., circulatory, respiratory, and renal failure risks) using the respective assessments by treating physicians and treating nurses with paired time-point assessments.

    Time frame: Assessments are collected within the first 72 hours following admission.

  4. Comparison of predictive accuracy of treating versus non-treating physicians

    Evaluated and tested for the same as the primary, secondary and first exploratory outcomes (i.e., circulatory, respiratory, and renal failure risks) using the respective assessments by clinicians (treating) and solely relying on EHR data (non-treating) with paired time-point assessments.

    Time frame: Assessments are collected within the first 72 hours following admission.

  5. Calibration analysis of clinician prediction scores

    Calibration analysis of clinician prediction scores. The study assesses the prediction capabilities of a collective of clinicians. However, humans might be ill-calibrated amongst each other with respect to providing probability estimates. Assessing the collective's performance without calibration of the individuals amongst each other might underestimate the actual predictive capabilities of the clinicians if they were well-calibrated. We propose to re-assess the primary and secondary outcomes but additionally perform a risk score calibration amongst physicians.

    Time frame: Assessments are collected within the first 72 hours following admission.

  6. Exploratory analysis of patterns (e.g., systematic errors/biases) of predictive performance in human and machine assessors

    1. Influencing factors for the prediction capabilities of clinicians such as clinician seniority, clinician role (incl. nurse vs physician), time of day, weekday, workload, order of patients seen and other collected information on participating clinicians. 2. Systematic errors and biases of both human expert assessments and machine learning based predictions on patient age, diagnostic group, or any of the collected patient-specific or clinician-related collected variables.

    Time frame: Assessments are collected within the first 72 hours following admission.

07

Study locations

1 site
  • University Hospital Inselspital, Berne
    Bern, Canton of Bern 3010, Switzerland
08

References and documents

Individual participant data

Plan to share: Undecided

No publications or documents are linked to this record.

09

Updates

Tracking since Sep 25, 2026
No changes since tracking began. The registry record was last updated on Aug 28, 2025, before this site started recording changes on Sep 25, 2026. Its history is on ClinicalTrials.gov ↗
10

Registry details

Key details

Study ID
NCT06814327
Lead sponsor
ETH Zurich
Collaborators
Insel Gruppe AG, University Hospital Bern
Responsible party
Sponsor
First posted
Feb 7, 2025
Start date
Nov 18, 2024
Primary completion
May 15, 2025
Completion
May 15, 2026 (estimated)
Last update
Aug 28, 2025

Study contacts

Martin Faltys, Dr. med.
principal investigator · Insel Gruppe AG, University Hospital Bern

Oversight

Data monitoring committee
No
FDA-regulated drug
No
FDA-regulated device
No
View the source record on ClinicalTrials.gov ↗

Not currently enrolling

This study is active, not recruiting, as verified in Nov 2024. You cannot join it, but the record below documents what was studied.

Follow this study

Get an email when the registry record changes — status, dates, results — or when someone posts here.

Sign in to follow

Discussion

Questions and observations about this study, from anyone following it. Not medical advice, and not a channel to the study team — their contact details are on the registry record.

Sign in to join the discussion. Reading takes no account; posting does. You choose a display name, and a pseudonym is the default.

Nothing here yet. If you are running this trial, taking part in it, or weighing whether to, this is the place to say so.

Start the discussion