An evaluation of 308,774 records compared cardiovascular-risk models. Histogram gradient boosting achieved PR-AUC 0.3177, AUROC 0.8407 and Brier score 0.0633, alongside subgroup calibration and decision-utility audits.
Key findings
- Histogram gradient boosting achieved PR-AUC 0.3177, AUROC 0.8407, Brier 0.0633 and ECE 0.0045. Discrimination was stable across subgroups, calibration varied by age and self-rated health, and net benefit was positive at thresholds 0.05–0.15.
Why this matters globally
The study reinforces that clinical AI evaluation must go beyond discrimination to include calibration, subgroup equity and decision consequences at clinically relevant thresholds.
Thai researcher contribution
Mahasarakham University researchers contributed a translational evaluation framework linking explainability, calibration decomposition and decision utility.
Limitations to consider
This was secondary data, with population source and label collection not fully described in the abstract. External and prospective validation are absent, and subgroup calibration may shift with setting and prevalence.