Hospital prediction and decision benchmark analysis

Two views of hospital evidence.

Independent EHRSHOT and MIMIC-CDM analysis: longitudinal prediction, few-shot denominators, diagnostic simulation and the limits of hospital benchmark transfer.

Hospital evidence / two experimental designs

Prediction and diagnosis follow different paths

01

EHRSHOT

6,739 patients [1]

  1. 01Fix patient partitions
  2. 02Define prediction time
  3. 03Sample training and validation labels
  4. 04Evaluate held-out labels

AUROC and AUPRC [1]

02

MIMIC-CDM

2,400 patient cases [5]

  1. 01Select the eligible cohort
  2. 02Reveal information
  3. 03Collect the conclusion
  4. 04Score by pathology

Per-class diagnostic accuracy [4]

2,295 train; 2,232 validation; 2,212 test. EHRSHOT partitions refer to patients. MIMIC-CDM covers four selected abdominal pathologies.

01 /

Inspect the task

Trace the actual input, output and evaluation setting for each named benchmark.

02 /

Read the evidence

Explore cited cohort counts and selected historical results with their measurement boundaries.

03 /

Make the inference explicit

Separate the authors’ observations from our analytical interpretation and proposed evaluation questions.

Our analytical question

EHRSHOT tests prediction from longitudinal coded records. MIMIC-CDM tests diagnostic decisions from selected retrospective cases. This independent analysis puts their inputs, denominators and information settings side by side. Explore what a prediction label represents, why a selected disease cohort cannot establish general emergency-care performance, and where published evidence stops. We report aggregate metadata and historical study results, with no patient material, new model runs or implied affiliation with the benchmark authors.

Read the original evidence closely

The benchmark, unpacked.

Primary source library ↗
Dossier01

2023 paper, arXiv v2

EHRSHOT ↗

The denominator is a prediction time, not just a patient.

UnitA labeled prediction event; patients can contribute multiple labels.MeasureAUROC and AUPRC
Dossier02

2024 paper; dataset metadata v1.1

MIMIC-IV-Ext Clinical Decision Making ↗

Requesting evidence is harder than receiving the finished case.

UnitOne patient case within a selected abdominal pathology.MeasurePer-class diagnostic accuracy

An original analytical tool

Prediction or diagnostic action?

Evidence explorer

Compare EHRSHOT’s prediction targets with MIMIC-CDM’s information settings. Each row states the actual unit being scored.

8 of 8 evidence entries shown

Operational prediction

Length of stay

Read dossier ↗
Input
Coded record up to prediction time
Output
Binary outcome
Measure
AUROC / AUPRC
Interpretation boundary

Repeated labels can come from the same patient.

[1]
Operational prediction

Thirty-day readmission

Read dossier ↗
Input
Coded history before discharge
Output
Binary outcome
Measure
AUROC / AUPRC
Interpretation boundary

Only events observed in the source system can form labels.

[1]
Operational prediction

ICU transfer

Read dossier ↗
Input
Coded record during admission
Output
Binary outcome
Measure
AUROC / AUPRC
Interpretation boundary

Rare positive labels change precision–recall context.

[1]
Clinical prediction

Future lab result

Read dossier ↗
Input
Prior structured record
Output
Four-way class
Measure
AUROC / AUPRC aggregation
Interpretation boundary

Appendix Table 7 and the release describe four classes; paper §3.2 prose says five. This atlas follows the task table and release.

[1][2]
Clinical prediction

New diagnosis

Read dossier ↗
Input
Prior coded timeline
Output
Future diagnosis label
Measure
AUROC / AUPRC
Interpretation boundary

Coding and prediction horizon shape the target.

[1]
Clinical prediction

Chest X-ray findings

Read dossier ↗
Input
Prior coded record
Output
Fourteen-label finding prediction
Measure
AUROC / AUPRC
Interpretation boundary

The model does not receive the image.

[1]
Diagnostic simulation

Interactive abdominal diagnosis

Read dossier ↗
Input
History, then requested findings
Output
Diagnosis and treatment proposal
Measure
Per-class diagnostic accuracy
Interpretation boundary

Four selected diseases do not establish general specificity.

[4]
Diagnostic simulation

Full-information diagnosis

Read dossier ↗
Input
Prepared case evidence up front
Output
Diagnosis
Measure
Per-class diagnostic accuracy
Interpretation boundary

Removes the need to choose the next evidence request.

[4]

Different cohorts, outputs and metrics make a combined hospital-accuracy number inappropriate. [1][4]

Original analysis / methods and interpretation

What the score leaves unsaid.

All analyses →

Questions, answered

Read the result in context.

Specific tasks. Stated conditions.
Inspect every source.

Does EHRSHOT evaluate clinical note understanding?

The released benchmark uses structured coded records and excludes free-text notes. Its results do not directly establish note reading or summarization performance.

Is an EHRSHOT label the same as a patient?

No. One patient can contribute multiple labeled prediction events. The paper reports patient and label counts separately.

What is MIMIC-CDM-FI?

It is the full-information comparison setting, in which a prepared case is supplied up front. The interactive setting requires the model to request findings.

Does MIMIC-CDM represent every abdominal-pain presentation?

No. It selects four target pathologies and requires specified modalities. Its reported accuracy cannot establish general emergency-care specificity or positive predictive value.

Working tool / saved on this device

Prepare a reproducible benchmark reading

Interactive worksheet

Use this secondary worksheet to record the version, task and evidence boundary of the benchmark you are considering. Checked items indicate documented decisions, not measured performance.

Identify the experiment

Evidence you can inspect. Benchmark dossiers distinguish published facts from our interpretation, with source versions and access notes attached.

Download the evidence ↗