TY - JOUR
T1 - Large language models approach clinician performance in ESC cardiovascular risk stratification
T2 - a vignette-based benchmark study
AU - Ferreira Santos, José
AU - de Brito Duarte, Regina
AU - Mota, Inês
AU - Santos, Rita Carvalheira
AU - Moreira, José Maria
AU - Campos, Joana
AU - Silva, Nuno André
AU - Neves, Bernardo
AU - Ladeiras-Lopes, Ricardo
AU - Leite, Francisca
AU - Dores, Hélder
N1 - Publisher Copyright:
© The Author(s) 2026. Published by Oxford University Press on behalf of the European Society of Cardiology. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
PY - 2026/6
Y1 - 2026/6
N2 - Aims: Guideline-based cardiovascular risk stratification requires three distinct competencies: extracting risk factor data from clinical text, computing a validated risk score, and applying guideline-defined thresholds to assign a final risk category. We evaluated contemporary large language models (LLMs) on each of these tasks within the European Society of Cardiology (ESC) SCORE2 framework and compared LLM performance against a pooled individual clinician benchmark to contextualize findings against real-world human reproducibility. Methods and results: Eleven LLMs were evaluated using 30 simulated outpatient clinical vignettes presented in both Portuguese and English. For each vignette, models extracted cardiovascular risk factors, determined SCORE2 applicability, generated 10-year risk estimates where appropriate, and assigned a final three-class ESC risk category. A committee of three cardiologists established the reference standard; eight independent clinicians provided an individual-level human benchmark. Traditional risk-factor extraction was near-perfect across all models (micro-F1 0.97–0.99). Agreement with expert-assigned final risk categories was moderate and variable (best: GPT-4o, quadratic-weighted κw 0.69, 95% CI 0.44–0.84), with 10 of 11 models more often underestimating than overestimating risk. To isolate the source of classification error, post hoc deterministic recalculation of SCORE2 was performed using model-extracted variables in eligible vignettes; this markedly improved agreement across all models (κw 0.85–0.90), demonstrating that extraction was largely intact and computational execution was the primary failure mode. The pooled individual clinician benchmark showed moderate agreement with the reference standard (κw 0.52, 95% CI 0.28–0.67), indicating that the best-performing LLMs matched or exceeded the average individual clinician on this guideline-based task. Performance was broadly consistent across Portuguese and English. Conclusion: Contemporary LLMs reliably extract cardiovascular risk information from clinical text, and the best-performing systems achieved agreement within the range of average individual clinicians on this structured task. Their principal limitation lies in downstream computation and rule application.
AB - Aims: Guideline-based cardiovascular risk stratification requires three distinct competencies: extracting risk factor data from clinical text, computing a validated risk score, and applying guideline-defined thresholds to assign a final risk category. We evaluated contemporary large language models (LLMs) on each of these tasks within the European Society of Cardiology (ESC) SCORE2 framework and compared LLM performance against a pooled individual clinician benchmark to contextualize findings against real-world human reproducibility. Methods and results: Eleven LLMs were evaluated using 30 simulated outpatient clinical vignettes presented in both Portuguese and English. For each vignette, models extracted cardiovascular risk factors, determined SCORE2 applicability, generated 10-year risk estimates where appropriate, and assigned a final three-class ESC risk category. A committee of three cardiologists established the reference standard; eight independent clinicians provided an individual-level human benchmark. Traditional risk-factor extraction was near-perfect across all models (micro-F1 0.97–0.99). Agreement with expert-assigned final risk categories was moderate and variable (best: GPT-4o, quadratic-weighted κw 0.69, 95% CI 0.44–0.84), with 10 of 11 models more often underestimating than overestimating risk. To isolate the source of classification error, post hoc deterministic recalculation of SCORE2 was performed using model-extracted variables in eligible vignettes; this markedly improved agreement across all models (κw 0.85–0.90), demonstrating that extraction was largely intact and computational execution was the primary failure mode. The pooled individual clinician benchmark showed moderate agreement with the reference standard (κw 0.52, 95% CI 0.28–0.67), indicating that the best-performing LLMs matched or exceeded the average individual clinician on this guideline-based task. Performance was broadly consistent across Portuguese and English. Conclusion: Contemporary LLMs reliably extract cardiovascular risk information from clinical text, and the best-performing systems achieved agreement within the range of average individual clinicians on this structured task. Their principal limitation lies in downstream computation and rule application.
KW - Artificial intelligence
KW - Cardiovascular prevention
KW - Clinical decision support
KW - Large language models
KW - Risk stratification
KW - SCORE2
UR - https://www.scopus.com/pages/publications/105041854433
U2 - 10.1093/ehjdh/ztag073
DO - 10.1093/ehjdh/ztag073
M3 - Article
AN - SCOPUS:105041854433
SN - 2634-3916
VL - 7
JO - European Heart Journal - Digital Health
JF - European Heart Journal - Digital Health
IS - 5
M1 - ztag073
ER -