CorpX Analytics · How the models are measured
EVIDENCE · METHODS

The models are ours, and they are scored.

WHAT WE TRAINED, ON WHAT, AND HOW IT DOES
THE SHORT VERSION

We trained the segmentation models ourselves, on a public dataset that permits it. On cases the models never saw, the four chambers and the myocardium reach a median Dice of 0.956 or better across 119 scans. Everything below says how that number was produced, and what it does not mean.

01

The models

Two models do the segmentation. A localiser finds the heart and the aorta in the whole volume. A chamber model then separates the myocardium and the four chambers inside it. Both are nnU-Net v2 networks that we trained. The architecture is open and widely used; the trained weights are our own work, and they are what the measurements depend on.

We did not license a third-party chamber model, and we do not run one. That matters for what we can tell you about it: we hold the training log, the case lists and the held-out scores, so every figure on this page can be traced rather than taken on trust.

02

What they were trained on

The chamber model was trained on 484 cases of the TotalSegmentator v1 dataset, across all acquisition phases rather than contrast alone. That dataset is published under Creative Commons Attribution 4.0, which permits derived work provided the source is credited. It is credited here, and this is the only cohort in our work whose derived output we may release.

TRAINING DATA ATTRIBUTION

Wasserthal, J. et al. TotalSegmentator: Robust Segmentation of 104 Anatomic Structures in CT Images. Radiology: Artificial Intelligence (2023). Dataset licensed CC BY 4.0.

03

Why the holdout is genuinely held out

A score only means something if the model never saw the cases it was scored on. We checked that from the training artefacts rather than assuming it. The candidate list held 604 cases. Training used 387, validation another 97. The remaining 120 were never seen, and their overlap with the validation set is zero.

The arithmetic closes exactly. That is the point: it is checkable, and you can ask us for the case lists.

04

What it scores

Dice measures how far two outlines agree, where 1.000 is exact. Volume error is the difference between the measured volume and the reference, as a percentage of the reference.

StructureDice, median10–90Volume error %
Myocardium0.9560.904–0.970+0.1
Left atrium0.9790.960–0.986−0.1
Left ventricle0.9750.951–0.983−0.1
Right atrium0.9710.931–0.981−0.1
Right ventricle0.9730.944–0.981−0.2

119 HELD-OUT CASES · ONE REFUSED BY THE BORDER GATE, NOT SCORED

The localiser is scored separately, on 122 held-out cases: a median Dice of 0.970 for the heart and 0.979 for the aorta.

05

What this does not establish

Three limits, stated because a careful reader would find them anyway.

  • It is not a comparison. No other chamber model has been scored on these cases. The claim is that ours is measured, not that it is better.
  • It is resampled data. The public dataset ships at 1.5 mm, so performance on native-resolution clinical scans is not established by this run.
  • It is a single reconstructed phase. These are not end-diastolic volumes and must not be read against end-diastolic reference ranges.
06

Every value carries its record

Each delivery ships the model version, the parameters it ran with, the hash of the volume it read, and the de-identification record. A measurement you cannot trace back to the scan and the method that produced it is not evidence, and we do not ship one.

07

Questions

Ask for the case lists, the training log or the scoring artefacts: anamaria.chioran@corpxanalytics.com

← BACK TO HOME REGULATORY STATUS →